TGViewer
Artificial Intelligence Artificial Intelligence @machinelearning_deeplearning · 56.3K subscribers
Post #1738 3.81K
✅ Model Evaluation in Machine Learning 📊🔍

Once you've trained a model, how do you know if it's any good? That’s where model evaluation comes in.

1️⃣ For Supervised Learning

You compare the model’s predictions to the actual labels using metrics like:

🔹 Confusion Matrix
A confusion matrix shows how many predictions were correct vs. incorrect, broken down by class.

from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay

y_pred = model.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot()


This helps you compute:
• True Positives (TP): Correctly predicted positives
• True Negatives (TN): Correctly predicted negatives
• False Positives (FP): Incorrectly predicted as positive
• False Negatives (FN): Incorrectly predicted as negative

🔹 Accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)

Measures overall correctness:
Accuracy = (TP + TN) / (TP + TN + FP + FN)
Best when classes are balanced.

🔹 Precision Recall
from sklearn.metrics import precision_score, recall_score
precision = precision_score(y_test, y_pred, average='macro')
recall = recall_score(y_test, y_pred, average='macro')

• Precision: Of all predicted positives, how many were correct?
Precision = TP / (TP + FP)
• Recall: Of all actual positives, how many did we catch?
Recall = TP / (TP + FN)
Use average='macro' for multiclass problems.

🔹 F1 Score
from sklearn.metrics import f1_score
f1 = f1_score(y_test, y_pred, average='macro')

Balances precision and recall:
F1 = 2 * (Precision * Recall) / (Precision + Recall)
Great when you need a single score that considers both false positives and false negatives.

🔹 Mean Squared Error (MSE) – For Regression
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)

Measures average squared difference between predicted and actual values.
Lower is better.

2️⃣ For Unsupervised Learning

Since there are no labels, we use different strategies:

🔹 Silhouette Score
from sklearn.metrics import silhouette_score
score = silhouette_score(X, kmeans.labels_)

Measures how similar a point is to its own cluster vs. others.
Ranges from -1 (bad) to +1 (good separation).

🔹 Inertia
print("Inertia:", kmeans.inertia_)

Sum of squared distances from each point to its cluster center.
Lower inertia = tighter clusters.

🔹 Visual Inspection
import matplotlib.pyplot as plt
plt.scatter(X[:, 0], X[:, 1], c=kmeans.labels_)
plt.title("KMeans Clustering")
plt.show()

Plotting clusters often reveals structure or overlap.

🧠 Pro Tip:
Always split your data into training and testing sets to avoid overfitting. For more robust evaluation, try:

from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=5)
print("Cross-Validation Scores:", scores)


💬 Double Tap ❤️ for more!
  • ❤ 9
More from @machinelearning_deeplearning
  1. Sep 22, 2026🚀 Welcome back to our AI Engineer Roadmap! ❤️ In the previous posts, we learned about fun…
  2. Sep 21, 2026🚀 Welcome back to our AI Engineer Roadmap! ❤️ In the previous post, we explored functions…
  3. Sep 20, 2026In the previous post, we learned how conditional statements allow Python programs to make…
  4. Sep 19, 2026In the previous post, we explored Python operators and even tested ourselves with some tri…
  5. Sep 18, 2026In the previous post, we learned how to convert one data type into another using Type Cast…
  6. Sep 17, 2026In the previous post, we learned how to take input from users and display output. One impo…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →