๐ AI Interview Questions with Answers โ Part 4
31. What is a classification problem in Machine Learning?
A classification problem is a type of supervised learning where the model predicts categories or labels instead of numerical values.
Examples:
- Spam or Not Spam
- Fraud or Not Fraud
- Disease Positive or Negative
Goal: Assign input data to the correct class.
Example: An email spam filter classifies emails into:
- Spam
- Not Spam
32. What is the difference between Logistic Regression and Linear Regression?
Linear Regression
- Predicts continuous values
- Used for regression tasks
- Output can be any number
- Straight-line relationship
Logistic Regression
- Predicts categories
- Used for classification tasks
- Output ranges between 0 and 1
- Uses sigmoid function
Linear Regression Example: Predicting house prices.
Logistic Regression Example: Predicting whether a customer will buy a product or not.
Sigmoid Function Used in Logistic Regression:
sigma(x) = 1 / (1 + e^(-x))
This converts output into probabilities.
33. How does a Decision Tree work?
A Decision Tree splits data into branches based on conditions.
It works like a flowchart:
- Root node โ Starting point
- Decision nodes โ Conditions
- Leaf nodes โ Final prediction
How It Works:
1. Select best feature
2. Split the dataset
3. Repeat recursively
Advantages:
- Easy to understand
- Works for classification and regression
- Handles nonlinear data
Example: Loan approval system:
Income > โน50,000?
Credit score good?
Approve or reject loan
34. What are the advantages of Random Forest?
Random Forest is an ensemble learning algorithm that combines multiple Decision Trees.
Advantages:
- Higher accuracy
- Reduces overfitting
- Handles large datasets
- Works with missing values
- Robust to noise
How It Works: Many trees vote for the final prediction.
Example: If 100 trees predict:
80 say โSpamโ
20 say โNot Spamโ
Final output = Spam
35. What is Support Vector Machine (SVM)?
Support Vector Machine (SVM) is a supervised learning algorithm mainly used for classification.
It finds the best boundary (hyperplane) that separates classes.
Goal: Maximize the distance between classes.
Advantages:
- Effective in high-dimensional data
- Works well with smaller datasets
- Powerful for complex classification tasks
Example: Separating:
Cats vs Dogs
Fraud vs Non-Fraud
using the best possible boundary.
36. Why is Naive Bayes called โnaiveโ?
Naive Bayes is called โnaiveโ because it assumes all features are independent of each other.
In real life, this assumption is often unrealistic.
Example: While predicting spam emails:
Words may actually be related
But Naive Bayes assumes independence
Despite this โnaiveโ assumption, the algorithm performs surprisingly well in:
- Text classification
- Spam detection
- Sentiment analysis
37. How does the KNN algorithm work?
K-Nearest Neighbors (KNN) classifies data based on the closest neighboring data points.
How It Works:
1. Choose value of K
2. Find nearest neighbors
3. Majority vote determines class
Example:
If K = 5
Among 5 nearest neighbors:
4 are โRedโ
1 is โBlueโ
Prediction = Red
Advantages:
- Simple and intuitive
- No training phase
Disadvantages:
- Slow for large datasets
- Sensitive to irrelevant features
38. What is a confusion matrix?
A confusion matrix is a table used to evaluate classification models.
It compares actual values and predicted values.
Main Components:
- Actual Positive, Predicted Positive โ True Positive (TP)
- Actual Positive, Predicted Negative โ False Negative (FN)
- Actual Negative, Predicted Positive โ False Positive (FP)
- Actual Negative, Predicted Negative โ True Negative (TN)
Why Itโs Important:
It helps calculate accuracy, precision, recall, and F1-score.
39. What is the difference between precision and recall?
Post #2183
899