๐ Phase 2: Mathematics for Data Science
๐ Topic 3: Variance & Standard Deviation
Welcome back! ๐
In the previous lesson, you learned about Mean, Median, and Mode, which help us find the center of a dataset.
But knowing the average alone is not enough.
Imagine these two datasets:
Dataset A
40, 45, 50, 55, 60
Dataset B
10, 20, 50, 80, 90
Both datasets have the same mean (50), but they are very different.
โข Dataset A has values close to the mean.
โข Dataset B has values spread far away from the mean.
To measure this spread, we use Variance and Standard Deviation.
These are among the most important statistical concepts in Data Science and Machine Learning.
๐น 1. What is Variance?
Variance measures how far each value is from the mean.
โข Small variance โ Data points are close together.
โข Large variance โ Data points are widely spread.
Formula (Population Variance)
Variance = ฮฃ(x โ Mean)ยฒ / N
Where:
โข ฮฃ = Sum
โข x = Each data point
โข Mean = Average
โข N = Total number of observations
๐น 2. Example of Variance
Dataset: 10, 20, 30
Step 1: Find the Mean
(10 + 20 + 30) / 3 = 20
Step 2: Find the Difference from the Mean
10 โ 20 = -10
20 โ 20 = 0
30 โ 20 = 10
Step 3: Square the Differences
100, 0, 100
Step 4: Calculate Variance
(100 + 0 + 100) / 3 = 66.67
๐น 3. What is Standard Deviation? โญ
Standard Deviation (SD) is simply the square root of the variance.
Formula
Standard Deviation = โVariance
Using the previous example:
Variance = 66.67
SD = โ66.67 โ 8.16
๐น 4. Why Standard Deviation is Preferred?
Variance is measured in squared units, making it harder to interpret.
Standard Deviation is measured in the same units as the original data, making it easier to understand.
Example:
If salaries are measured in rupees:
โข Variance โ Rupeesยฒ โ
โข Standard Deviation โ Rupees โ
๐น 5. Python Example
Using the "statistics" module:
import statistics
numbers = [10, 20, 30]
print(statistics.pvariance(numbers))
print(statistics.pstdev(numbers))