π§ AI Project #2: Spam Email Detector with NLPπ― Project Goal Build an AI app that reads any email text and tells you if itβs
Spam or
Ham in 1 second.
Spam = unwanted/promotional mail. Ham = legit mail like βTeam meeting at 4 PMβ.
π§ Skills Youβll LearnPython: Strings, Functions, Lists
Data: Pandas, NumPy for dataset handling
NLP: Text cleaning, Tokenization, Stopwords, Stemming, TF-IDF
ML: Naive Bayes, Logistic Regression, Random Forest
Deployment: Streamlit for web app
π Step 1: Dataset Typical format: 2 columns
Email Text | Label
"Win βΉ1 Lakh now, click here" | 1 β Spam
"Project report attached" | 0 β Ham
Label: 0 = Ham, 1 = Spam
π Step 2: Load & Explore Data import pandas as pd
df = pd.read_csv("spam.csv")
print(df.head())
print(df['label'].value_counts())
Check: total emails, spam %, missing values.
π§Ή Step 3: Text PreprocessingRaw: "Congratulations!!! You WON βΉ50,000... CLICK NOW!!!"
Clean: "congratulation won click"
Convert to lowercase text = text.lower()
Remove punctuation import string
text = text.translate(str.maketrans('', '', string.punctuation))
Step 4: Tokenization Split sentence into words
"you won prize" β ["you", "won", "prize"]
from nltk.tokenize import word_tokenize
tokens = word_tokenize(text)
Step 5: Remove StopwordsRemove common words:
the, is, a, an
["you", "won", "the", "prize"] β ["won", "prize"]
from nltk.corpus import stopwords
words = [w for w in tokens if w not in stopwords.words('english')]
Step 6: StemmingReduce words to root form
running, runs, ran β run
playing, played β play
from nltk.stem import PorterStemmer
ps = PorterStemmer()
words = [ps.stem(w) for w in words]
Step 7: Convert Text to Numbers with TF-IDFML models canβt read text. TF-IDF gives importance score to words.
Spam words like βwin, free, offerβ get high scores.
from sklearn.feature_extraction.text import TfidfVectorizer
tfidf = TfidfVectorizer()
X = tfidf.fit_transform(df['text'])
Step 8: Train Model from sklearn.linear_model import LogisticRegression
model = LogisticRegression()
model.fit(X_train, y_train)
Other models to try: Multinomial Naive Bayes, Random Forest, XGBoost
Step 9: Predict New Emailsample = ["Congratulations you won free iPhone"]
sample_vec = tfidf.transform(sample)
pred = model.predict(sample_vec)
print("Spam" if pred[0]==1 else "Ham")
Step 10: Check Performance from sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix
print("Accuracy:", accuracy_score(y_test, y_pred))
Also check Precision, Recall, F1-Score, Confusion Matrix.
π¨ Step 11: Build Streamlit App import streamlit as st
st.title("Spam Email Detector")
email = st.text_area("Paste email text here")
if st.button("Check"):
email_vec = tfidf.transform([email])
result = model.predict(email_vec)
if result[0]==1:
st.error("π¨ Spam Email Detected")
else:
st.success("β
Legit Email")
Run:
streamlit run app.pyβ Features to AddBeginner: Show accuracy, simple UI
Intermediate: Add spam probability %, save email history
Advanced: Multi-language support, Phishing detection, Gmail API integration
π Project Folder Structure spam-detector/
βββ data/spam.csv
βββ models/model.pkl
βββ notebooks/training.ipynb
βββ
app.py βββ
train.py βββ requirements.txt
βββ
README.md πΌ Resume Bullet Spam Email Classifier using NLP Built end-to-end text classification pipeline with Python, NLTK, TF-IDF, and Logistic Regression. Achieved 97%+ accuracy. Deployed interactive Streamlit app for real-time spam detection.
π Mini Challenge for You1. Add phishing email detection
2. Show spam probability percentage
3. Deploy on Hugging Face Spaces for free
π₯ Double Tap β€οΈ For Part-3