✅ Top AI Interview Questions with Answers: Part-3 🧠
21. What are generative vs discriminative models?
- Generative Models learn the joint probability (P(x, y)) and can generate new data.
- Examples: Naive Bayes, GANs, HMM
- Discriminative Models learn the conditional probability (P(y|x)) and focus on classification.
- Examples: Logistic Regression, SVM, Neural Networks
22. Explain PCA (Principal Component Analysis)
PCA is a dimensionality reduction technique.
It transforms features into a new coordinate system (principal components), keeping only the most important ones that explain the variance in data.
Helps reduce overfitting, improves visualization, and speeds up training.
23. What is feature selection and why is it important?
Feature selection involves choosing the most relevant features for your model.
Benefits:
- Reduces overfitting
- Improves model accuracy
- Speeds up training
Methods: Filter (correlation), Wrapper (RFE), Embedded (Lasso)
24. What is one-hot encoding?
A method to convert categorical data into numerical format.
Each category becomes a binary column (0 or 1).
Example:
Color = Red, Green, Blue →
Red = [1,0,0], Green = [0,1,0]
25. What is dimensionality reduction?
Reducing the number of input variables in your dataset while retaining important information.
Techniques:
- PCA (unsupervised)
- LDA (supervised)
Used to simplify models and avoid the curse of dimensionality.
26. What is regularization? (L1 vs L2)
Regularization prevents overfitting by penalizing large weights.
- L1 (Lasso): Adds absolute values → can shrink some weights to zero (feature selection).
- L2 (Ridge): Adds squared values → reduces weight magnitudes but keeps all features.
27. What is the curse of dimensionality?
As dimensions increase, the data becomes sparse and harder to model.
Distance-based algorithms (like KNN) become less effective.
Solution: Use dimensionality reduction or feature selection.
28. How does K-Means clustering work?
An unsupervised algorithm that groups data into k clusters.
Steps:
1. Choose k centroids
2. Assign each point to the nearest centroid
3. Recalculate centroids
4. Repeat until convergence
Used in market segmentation, image compression.
29. Difference between KNN and K-Means
- KNN (K-Nearest Neighbors): Supervised, used for classification/regression
- K-Means: Unsupervised, used for clustering
KNN uses labeled data; K-Means does not.
30. What is Naive Bayes classifier?
A probabilistic classifier based on Bayes' Theorem.
It assumes features are independent (naive assumption).
Very fast and works well for text classification like spam detection.
💬 Double Tap ♥️ For More
Post #1713
3.11K
- ❤ 12