• +0.9 → Very strong positive
• +0.5 → Moderate positive
• +0.1 → Weak positive
• 0 → No linear relationship
• -0.1 → Weak negative
• -0.5 → Moderate negative
• -0.9 → Very strong negative
The exact interpretation depends on the domain and context.
🔹 9. Pearson Correlation Coefficient ⭐The most commonly used correlation measure is the Pearson correlation coefficient.
It is calculated as:
• r = Cov(X,Y) / (StdDev X ** StdDev Y)
Where:
• Cov(X,Y) = Covariance between X and Y
• StdDev X = Standard deviation of X
• StdDev Y = Standard deviation of Y
Because covariance is divided by the standard deviations, the result is standardized between -1 and +1.
🔹 10. Covariance vs Correlation• Covariance: Measures direction of joint variation, Can have any numerical value, Depends on units, Harder to interpret, Useful mathematically
• Correlation: Measures direction and strength, Always between -1 and +1, Unitless, Easier to interpret, Very useful for EDA
🔹 11. Positive Correlation ExampleSuppose: Advertising Spend ↑ → Sales ↑
If higher advertising spending generally corresponds to higher sales, the correlation may be positive.
• For example: r = 0.85 → This indicates a strong positive linear relationship.
🔹 12. Negative Correlation ExampleSuppose: Price ↑ → Demand ↓
You might observe: r = -0.80 → This indicates a strong negative linear relationship.
🔹 13. Correlation Does NOT Mean Causation ⭐This is one of the most important concepts in Data Science.
Suppose we observe: Ice Cream Sales ↑ ↔ Swimming Pool Accidents ↑
There may be a positive correlation. But eating ice cream doesn't necessarily cause swimming accidents.
A third variable — hot weather — could influence both:
• Hot Weather → Ice Cream Sales
• Hot Weather → Swimming Activity → Accidents
Therefore: Correlation does not prove causation.
🔹 14. Correlation and Machine LearningCorrelation is frequently used during Exploratory Data Analysis.
For example, suppose you're predicting house prices. You might examine correlations between:
• House size
• Number of bedrooms
• Location-related variables
• Age of property
• Price
A strong correlation between house size and price may indicate that house size could be a useful predictive feature.
However, correlation alone does not determine whether a feature should be included in a model.
🔹 15. Correlation Matrix ⭐When a dataset contains many numerical variables, we can calculate correlations between every pair of variables. This produces a correlation matrix.
Example:
• Age | Income | Spending
• Age: 1.00, 0.65, -0.10
• Income: 0.65, 1.00, 0.72
• Spending: -0.10, 0.72, 1.00
The diagonal is always 1.00 because every variable has a perfect correlation with itself.
🔹 16. Detecting Multicollinearity• Correlation can help identify multicollinearity.
• Multicollinearity occurs when two or more predictor variables are highly correlated with each other.
• For example: Annual Income ↔ Monthly Income — These variables contain very similar information.
• Including highly correlated predictors can create problems for some models, particularly linear regression, because it can make coefficient estimates unstable and harder to interpret.
🔹 17. Python ExampleUsing Pandas:
import pandas as pd
data = {
"Hours": [2, 4, 6, 8, 10],
"Score": [50, 60, 70, 80, 90]
}
df = pd.DataFrame(data)
print(df["Hours"].corr(df["Score"]))