๐น 8. PMF vs PDF vs CDF
PMF: Used for Discrete data. Represents Probability of an exact outcome
PDF: Used for Continuous data. Represents Probability density
CDF: Used for Discrete & continuous. Represents Probability up to a value
A simple way to remember:
PMF โ Exact probability for discrete outcomes
PDF โ Density across continuous values
CDF โ Cumulative probability up to a value
๐น 9. Example: Discrete Distribution
Suppose a machine produces defective products.
Let: X = Number of defective products
Possible values: 0, 1, 2, 3
Suppose:
P(X=0) = 0.50
P(X=1) = 0.30
P(X=2) = 0.15
P(X=3) = 0.05
Check: 0.50 + 0.30 + 0.15 + 0.05 = 1.00
Therefore, this is a valid probability distribution.
๐น 10. Example: Continuous Distribution
Suppose: X = Customer waiting time
Waiting time could be: 2.1 minutes, 2.15 minutes, 2.157 minutes, 2.1578 minutes...
Because there are infinitely many possible values, we treat it as a continuous random variable.
A PDF can describe how densely the waiting times are distributed.
๐น 11. Normal Distribution โญ
One of the most important probability distributions in Data Science is the Normal Distribution.
It is often called the bell curve because of its shape.
A normal distribution is characterized by: Mean, Standard deviation
Many natural and measurement-related variables can be approximately normally distributed under suitable conditions.
Examples: Measurement errors, Certain biological measurements, Standardized test scores
๐น 12. Properties of Normal Distribution
For a perfectly symmetric normal distribution: Mean = Median = Mode
The distribution is symmetric around its mean.
A common rule of thumb is the 68โ95โ99.7 rule:
Within 1 Standard Deviation: Approximately 68%
Within 2 Standard Deviations: Approximately 95%
Within 3 Standard Deviations: Approximately 99.7%
๐น 13. Binomial Distribution
The Binomial Distribution is a discrete probability distribution used when:
There are a fixed number of trials, Each trial has two possible outcomes, The probability of success is constant, Trials are independent.
Examples: Number of successful predictions, Number of heads in coin tosses, Number of defective products in a fixed sample
Example: 10 coin tosses. X = Number of Heads. Possible values: 0, 1, 2, ..., 10
๐น 14. Poisson Distribution
The Poisson Distribution is commonly used to model the number of events occurring within a fixed interval when events occur at a certain average rate under appropriate assumptions.
Examples: Number of customer calls per hour, Number of website visits per minute, Number of machine failures per month, Number of support tickets per day
๐น 15. Why Probability Distributions Matter in Data Science?
Probability distributions help Data Scientists:
โ Understand data patterns
โ Detect unusual observations
โ Model uncertainty
โ Perform statistical tests
โ Build predictive models
โ Simulate data
โ Estimate probabilities
๐น 16. Python Example
import numpy as np
data = np.random.normal(
loc=50,
scale=10,
size=1000
)
print(data[:5])