🚀 Top 100 Data Science Interview Questions
🧠 Data Science Fundamentals
1. What is data science?
2. What is the difference between data science, data analytics, and data engineering?
3. What are the main stages of a data science lifecycle?
4. What is a problem statement in data science?
5. What is the difference between descriptive, predictive, and prescriptive analytics?
6. What is feature engineering?
7. What is a data pipeline for data science?
8. What is exploratory data analysis (EDA)?
9. How do you approach a new dataset for the first time?
10. What is the difference between a model and a prototype?
📊 Statistics & Probability
11. What is the difference between population and sample?
12. What are mean, median, mode, variance, and standard deviation?
13. What is skewness and kurtosis?
14. What is a normal distribution?
15. What is central limit theorem (CLT)?
16. What is p‑value and how do you interpret it?
17. What are Type I and Type II errors?
18. What is confidence interval?
19. What is hypothesis testing?
20. What is correlation vs causation?
📉 Machine Learning Basics
21. What is machine learning?
22. What is the difference between supervised, unsupervised, and reinforcement learning?
23. What is overfitting and how do you prevent it?
24. What is underfitting and how do you detect it?
25. What is the bias‑variance tradeoff?
26. What is train/validation/test split?
27. What is cross‑validation?
28. What is regularization?
29. What is feature selection vs feature extraction?
30. What is the difference between bagging and boosting?
📊 Regression & Classification
31. What is linear regression and its assumptions?
32. What is logistic regression and where is it used?
33. What is multicollinearity and why is it a problem?
34. What is RMSE, MAE, and R²?
35. What is a confusion matrix?
36. What is precision, recall, and F1‑score?
37. What is ROC curve and AUC?
38. What is the difference between decision tree and random forest?
39. What is Gradient Boosting (e.g., XGBoost, LightGBM)?
40. When would you choose regression over classification?
🧩 Unsupervised Learning & Dimensionality Reduction
41. What is clustering?
42. How does K‑Means work?
43. What is hierarchical clustering?
44. What is DBSCAN?
45. What is dimensionality reduction?
46. What is PCA and why is it used?
47. What is SVD?
48. What is an elbow plot and silhouette score?
49. What is anomaly detection?
50. What is association rule learning?
🐍 Python for Data Science
51. How do you load and inspect data in pandas?
52. How do you handle missing values in pandas?
53. How do you perform group‑by and aggregation in pandas?
54. How do you merge or join DataFrames?
55. How do you handle categorical variables?
56. How do you write a custom function for data transformation?
57. How do you optimize a slow pandas script?
58. What are vectorized operations in pandas?
59. How do you plot basic charts with Matplotlib/Seaborn?
60. How do you unit‑test a data‑science pipeline?
📊 SQL & Data Wrangling
61. What is the difference between INNER, LEFT, RIGHT, and FULL JOIN?
62. What is GROUP BY and HAVING?
63. What is a subquery and CTE?
64. What is window function (e.g., ROW_NUMBER, RANK)?
65. How do you deduplicate records in SQL?
66. How do you handle time‑based aggregations?
67. How do you calculate month‑over‑month or day‑over‑day metrics?
68. How do you join a user table with a purchase table?
69. How do you optimize a slow SQL query?
70. What is indexing and when should you use it?
📊 Model Evaluation & Experimentation
71. How do you evaluate a classification model?
72. How do you evaluate a regression model?
Post #2285
6.69K
- ❤ 11
- 👍 1