"After identifying duplicates, I investigate whether they are genuine duplicate records or legitimate repeated transactions before removing anything."
8️⃣ What is an outlier? How would you handle it?
Sample Answer:
"An outlier is a value that is significantly different from the typical observations in a dataset.
I wouldn't automatically remove an outlier. First, I would investigate whether it represents a data-quality issue or a genuine business event.
For example, a transaction worth ₹10 million might initially look like an outlier, but it could be a legitimate high-value transaction. If it is a data-entry error, I would correct or exclude it according to the business rules."
9️⃣ What is the difference between a dimension and a measure?
Sample Answer:
"A dimension is generally used to categorize or describe data, while a measure is a numerical value that can usually be aggregated.
For example, in a sales dataset:
Dimensions: Customer, Product, Region, Date
Measures: Sales Amount, Quantity, Profit, Discount
In a dashboard, dimensions are commonly used to slice or group the data, while measures are used to calculate KPIs and metrics."
🔟 What steps do you follow when solving a data analysis problem?
Sample Answer:
"I generally follow a structured approach:
1. Understand the business problem.
2. Define the required metrics and success criteria.
3. Identify the relevant data sources.
4. Extract and validate the data.
5. Clean and transform the data.
6. Perform exploratory analysis.
7. Identify trends, patterns, and anomalies.
8. Validate the results.
9. Communicate the insights using appropriate visualizations.
10. Recommend actions based on the findings.
The most important step is understanding the business question first, because technically correct analysis can still be useless if it doesn't answer the actual business problem."
📌 Double Tap ❤️ For Part-2
Post #3160
3.04K
- ❤ 12