β
Model Evaluation in Machine Learning ππ
Once you've trained a model, how do you know if it's any good? Thatβs where
model evaluation comes in.
1οΈβ£ For Supervised LearningYou compare the modelβs predictions to the actual labels using metrics like:
πΉ Confusion Matrix A confusion matrix shows how many predictions were correct vs. incorrect, broken down by class.
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
y_pred = model.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot()
This helps you compute:
β’
True Positives (TP): Correctly predicted positives
β’
True Negatives (TN): Correctly predicted negatives
β’
False Positives (FP): Incorrectly predicted as positive
β’
False Negatives (FN): Incorrectly predicted as negative
πΉ Accuracy from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
Measures overall correctness:
Accuracy = (TP + TN) / (TP + TN + FP + FN)
Best when classes are balanced.
πΉ Precision Recall from sklearn.metrics import precision_score, recall_score
precision = precision_score(y_test, y_pred, average='macro')
recall = recall_score(y_test, y_pred, average='macro')
β’
Precision: Of all predicted positives, how many were correct?
Precision = TP / (TP + FP)
β’
Recall: Of all actual positives, how many did we catch?
Recall = TP / (TP + FN)
Use average='macro' for multiclass problems.
πΉ F1 Score from sklearn.metrics import f1_score
f1 = f1_score(y_test, y_pred, average='macro')
Balances precision and recall:
F1 = 2 * (Precision * Recall) / (Precision + Recall)
Great when you need a single score that considers both false positives and false negatives.
πΉ Mean Squared Error (MSE) β For Regression
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)
Measures average squared difference between predicted and actual values.
Lower is better.
2οΈβ£ For Unsupervised LearningSince there are no labels, we use different strategies:
πΉ Silhouette Score from sklearn.metrics import silhouette_score
score = silhouette_score(X, kmeans.labels_)
Measures how similar a point is to its own cluster vs. others.
Ranges from -1 (bad) to +1 (good separation).
πΉ Inertia print("Inertia:", kmeans.inertia_)
Sum of squared distances from each point to its cluster center.
Lower inertia = tighter clusters.
πΉ Visual Inspection import matplotlib.pyplot as plt
plt.scatter(X[:, 0], X[:, 1], c=kmeans.labels_)
plt.title("KMeans Clustering")
plt.show()
Plotting clusters often reveals structure or overlap.
π§ Pro Tip: Always split your data into
training and
testing sets to avoid overfitting. For more robust evaluation, try:
from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=5)
print("Cross-Validation Scores:", scores)
π¬
Double Tap β€οΈ for more!