🚀 Data Science Roadmap 2026📘 Phase 2: Mathematics for Data Science 📖 Topic 3: Variance & Standard DeviationWelcome back! 👋
In the previous lesson, you learned about Mean, Median, and Mode, which help us find the center of a dataset.
But knowing the average alone is not enough.
Imagine these two datasets:
Dataset A40, 45, 50, 55, 60
Dataset B10, 20, 50, 80, 90
Both datasets have the same mean (50), but they are very different.
• Dataset A has values close to the mean.
• Dataset B has values spread far away from the mean.
To measure this spread, we use Variance and Standard Deviation.
These are among the most important statistical concepts in Data Science and Machine Learning.
🔹 1. What is Variance?Variance measures how far each value is from the mean.
• Small variance → Data points are close together.
• Large variance → Data points are widely spread.
Formula (Population Variance)Variance = Σ(x − Mean)² / N
Where:
• Σ = Sum
• x = Each data point
• Mean = Average
• N = Total number of observations
🔹 2. Example of VarianceDataset: 10, 20, 30
Step 1: Find the Mean(10 + 20 + 30) / 3 = 20
Step 2: Find the Difference from the Mean10 − 20 = -10
20 − 20 = 0
30 − 20 = 10
Step 3: Square the Differences100, 0, 100
Step 4: Calculate Variance(100 + 0 + 100) / 3 = 66.67
🔹 3. What is Standard Deviation? ⭐Standard Deviation (SD) is simply the square root of the variance.
FormulaStandard Deviation = √Variance
Using the previous example:
Variance = 66.67
SD = √66.67 ≈ 8.16
🔹 4. Why Standard Deviation is Preferred?Variance is measured in squared units, making it harder to interpret.
Standard Deviation is measured in the same units as the original data, making it easier to understand.
Example:If salaries are measured in rupees:
• Variance → Rupees² ❌
• Standard Deviation → Rupees ✅
🔹 5. Python ExampleUsing the "statistics" module:
import statistics
numbers = [10, 20, 30]
print(statistics.pvariance(numbers))
print(statistics.pstdev(numbers))