🚀 Data Science Roadmap 2026📘 Phase 2: Mathematics for Data Science 📖 Topic 7: Descriptive Statistics — Range, Percentiles, Quartiles, IQR & Five-Number SummaryWelcome back! 👋
In the previous lesson you covered Probability Distributions.
Now we’re moving to
Descriptive Statistics — how we summarize data without predicting the population.
Today we’ll cover:
Range, Percentiles, Quartiles, IQR, Five-number summary, Outlier detection These are core for EDA.
🔹 1. What is Descriptive Statistics?Summarizes key characteristics of a dataset.
Example: Salaries: 30000, 35000, 40000, 45000, 50000
Instead of checking each value, use: Min, Max, Mean, Median, Quartiles, Percentiles, Std Dev
🔹 2. RangeFormula: Range = Maximum − Minimum
Example: 10, 20, 30, 40, 50 → Range = 50 − 10 = 40
Note: Very sensitive to outliers. 50 → 500 makes range jump to 490.
🔹 3. Percentiles ⭐Value below which X% of observations fall.
50th Percentile = Median
25th Percentile = 25% at or below
90th Percentile = 90% at or below
🔹 4. Real-World Example90th percentile score ≠ 90% marks. It means you did better than ∼90% of people.
🔹 5. QuartilesDivide data into 4 equal parts:
Q1 = 25th percentile
Q2 = 50th percentile = Median
Q3 = 75th percentile
🔹 6. Visualizing Quartiles0% ---- Q1 ---- Q2 ---- Q3 ---- 100%
25% 50% 75%
🔹 7. Interquartile Range (IQR) ⭐Formula: IQR = Q3 − Q1
Example: Q1=20, Q3=60 → IQR = 40. Middle 50% spans 40 units.
🔹 8. Why IQR MattersLess affected by outliers than Range.
Data: 10,20,30,40,50,1000 → Range=990 but IQR ignores the 1000.
🔹 9. Detecting Outliers Using IQR ⭐Lower Bound = Q1 − 1.5 × IQR
Upper Bound = Q3 + 1.5 × IQR
Values outside = potential outliers
🔹 10. Outlier ExampleQ1=20, Q3=60 → IQR=40
Lower = 20-60 = -40
Upper = 60+60 = 120
So < -40 or > 120 are outliers
🔹 11. Five-Number Summary ⭐ 1. Minimum 2. Q1 3. Median 4. Q3 5. Maximum
Ex: 10, 20, 30, 40, 50
🔹 12. Box PlotVisualizes the 5-number summary.
Box = Q1 to Q3. Line inside = Median. Whiskers = range without outliers.
🔹 13. Python Exampleimport numpy as np
data = [10, 20, 30, 40, 50, 60, 70]
q1 = np.percentile(data, 25)
median = np.percentile(data, 50)
q3 = np.percentile(data, 75)
iqr = q3 - q1
print("Q1:", q1, "Median:", median, "Q3:", q3, "IQR:", iqr)
🔹 14. Descriptive Statistics in Pandasimport pandas as pd
df = pd.DataFrame({"Salary": [30000, 35000, 40000, 45000, 50000]})
print(df["Salary"].describe())
describe() gives Count, Mean, Std, Min, 25%, 50%, 75%, Max
🔹 15. Real-World ExampleTransactions: Q1=₹500, Median=₹1000, Q3=₹2000 → IQR=₹1500
Use IQR to flag fraud, bulk orders, errors, or VIP customers. Investigate before deleting.
🔹 16. Range vs IQRRange: Easy but outlier-sensitive
IQR: Middle 50% only, robust to outliers
🔹 17. Percentile vs PercentagePercentage = out of 100.
Ex: 80% marks
Percentile = relative position.
Ex: 90th percentile
🔹 18. Common Mistakes❌ 90th percentile = 90% score
❌ Deleting all outliers blindly
❌ Thinking IQR covers all data
🎯 Practice Questions 1. Range of 10, 20, 30, 40, 50 = ?
2. Median = which percentile?
3. Q1=25, Q3=75 → IQR = ?
4. Upper outlier boundary formula?
5. 5 components of five-number summary?
🎯 Key Takeaways✅ Range = Max - Min✅ Q1=25th, Q2=50th=Median, Q3=75th✅ IQR = Q3 - Q1✅ 5-number summary = Min, Q1, Median, Q3, Max✅ Percentile ≠ Percentage👉 Double Tap ❤️ For More