❔Explain the Bias-Variance Tradeoff in simple terms.
✅ Answer:
Bias is error from overly simple assumptions (Underfitting). The model misses patterns. Variance is error from over-sensitivity to noise in training data (Overfitting). The model fails on new data.
Strategy: I aim for the "sweet spot" where both errors are minimized. I use Cross-Validation to monitor this balance. If bias is high, I increase model complexity. If variance is high, I use regularization or more data.
❔How do you choose between Precision and Recall?
✅ Answer:
It depends on the cost of a mistake. Precision matters when "False Alarms" are expensive (e.g., marking a safe email as Spam). Recall matters when "Missing a Case" is dangerous (e.g., failing to detect Cancer).
Strategy: I use the F1-Score if I need a balance between the two. For business stakeholders, I always translate these into "Lost Revenue" or "Customer Trust" to make the trade-off clear.
❔What is the difference between Random Forest and XGBoost?
✅ Answer:
Random Forest uses "Bagging", it builds many independent trees in parallel and averages them. It’s hard to overfit and works great out of the box.
XGBoost uses "Boosting", it builds trees sequentially, where each new tree tries to fix the errors of the previous one.
Strategy: I start with Random Forest as a robust baseline. I move to XGBoost (or LightGBM) when I need maximum accuracy and have the time to fine-tune hyperparameters.
❔ How do you handle a dataset where 99% of labels are 'Class A' and only 1% are 'Class B'?
✅ Answer:
First, I stop using Accuracy as a metric, as a "dumb" model would be 99% accurate by just guessing 'Class A'. I switch to AUPRC or Confusion Matrices.
Techniques: I use Resampling (Oversampling the minority or Undersampling the majority). I also use "Class Weights" in the model settings to penalize mistakes on the 1% more heavily.
❔What is the purpose of PCA (Principal Component Analysis)?
✅ Answer:
PCA is a dimensionality reduction tool. It transforms many correlated features into a smaller set of uncorrelated variables called "Principal Components" while keeping as much variance (information) as possible.
Usage: I use it to speed up training and reduce "Noise." Caution: It makes features hard to interpret. If the business needs to know why a prediction was made, I avoid PCA and use Feature Selection instead.