So: P(Disease) = 0.01.
A medical test is positive for someone who has the disease 99% of the time. But the test can also be positive for healthy people.
Suppose: P(Positive | No Disease) = 5%
Now someone receives a positive test. The important question is:
What is the probability that this person actually has the disease?
This is not simply 99%. We need to consider: The prior probability of the disease, The probability of a positive test among people with the disease, The probability of a positive test among people without the disease. Bayes' theorem combines these pieces of information.
🔹 8. Solving the Example
Let's assume:
P(Disease) = 0.01
P(Positive | Disease) = 0.99
P(No Disease) = 0.99
P(Positive | No Disease) = 0.05
First calculate the overall probability of a positive test:
P(Positive) = (0.99 × 0.01) + (0.05 × 0.99) = 0.0099 + 0.0495 = 0.0594
Now: P(Disease | Positive) = (0.99 × 0.01) / 0.0594 ≈ 0.167
So the probability is approximately 16.7%. This is much lower than 99%.
Because the disease is relatively rare and false positives occur. This demonstrates why base rates matter.
🔹 9. Base Rate
The base rate is the underlying frequency of an event in the population. In the previous example: Disease prevalence = 1%. That's the base rate.
Ignoring the base rate can lead to incorrect conclusions. This is known as the Base Rate Fallacy. A test can be highly accurate while the probability that a randomly selected person with a positive result actually has the disease can still be considerably lower than expected if the condition is rare.
🔹 10. Bayesian Updating
One of the most useful ideas in Bayesian Statistics is updating.
Suppose we initially believe: Probability of an event = 20%. Then we observe strong evidence supporting the event. Our posterior might become: 45%. Then we receive additional evidence. The probability might update again: 65%.
The process continues as new evidence arrives. So Bayesian inference is naturally suited to situations where:
New information arrives continuously.
🔹 11. Prior, Likelihood and Posterior
A simple way to remember the three:
🟦 Prior - What did I believe before seeing the data?
🟨 Likelihood - How strongly does the observed data support different possibilities?
🟩 Posterior - What do I believe after considering the data?
Remember: Posterior ∝ Prior × Likelihood
🔹 12. Bayesian vs Frequentist Statistics
Frequentist Approach: Generally treats unknown parameters as fixed but unknown. Probability is associated with the behavior of random data and procedures. Examples include: p-values, Confidence intervals, Hypothesis testing
Bayesian Approach: Treats uncertainty about parameters using probability distributions. It combines: Prior information + Data → Posterior. Examples include: Posterior distributions, Credible intervals, Bayesian parameter estimation
🔹 13. Confidence Interval vs Credible Interval
Confidence Interval: A frequentist concept. A 95% confidence interval is interpreted through the long-run behavior of the procedure that generates the interval.
Credible Interval: A Bayesian concept.