This project develops a machine learning-based phishing detection system to classify emails as Safe (Legitimate) or Phishing. Built for the Quantam Breach hackathon on April 25, 2025, it processes email text, trains a Support Vector Machine (SVM) model with TF-IDF features, and predicts phishing attempts with high accuracy (~96%). The system supports both batch predictions (via CSV) and interactive predictions (via scripts and a Flask web app), handling multi-line emails and edge cases like invoice scams.
- Data Preprocessing: Cleans email text, preserves URLs, and handles invalid data for robust feature extraction.
- Model Training: Uses SVM with TF-IDF features and class weights to achieve high accuracy and handle imbalanced data.
- Prediction:
- Batch Mode: Processes CSV files (e.g.,
test_emails.csv) for bulk predictions. - Interactive Mode: Supports real-time input via command-line (
predict.py) or a Flask web app (app.py).
- Batch Mode: Processes CSV files (e.g.,
- Web App: Interactive interface at
http://127.0.0.1:5000for classifying emails, displaying confidence scores and top keywords. - UI Design: Modern, responsive design with a black, white, and purple color scheme, using Inter font and Tailwind CSS for a professional look.
- Evaluation: Achieves ~96% accuracy, with high precision and recall, and fixes for misclassification of neutral emails (e.g., short meeting invites).
preprocess_data.py: Preprocesses raw email data (Phishing_Email.csv) intoprocessed_data.csv.train_model.py: Trains the SVM model, savingmodel.pklandvectorizer.pkl.predict.py: Predicts email types from CSV (test_emails.csv) or interactive input, saving topredictions.csv.app.py: Flask application for the web app, serving the interactive UI.templates/index.html: Webpage template with a modern black, white, and purple aesthetic.Phishing_Email.csv: Input dataset withEmail TextandEmail Type. Can be downloaded from "https://www.kaggle.com/code/kerlosmelad/emails-safety-predict/input"processed_data.csv: Cleaned dataset withcleaned_textandlabel.invalid_rows.csv: Rows with invalidEmail Typevalues.test_emails.csv: 15 multi-line test emails (8 safe, 7 phishing).predictions.csv: Prediction results frompredict.py.model.pkl,vectorizer.pkl: Trained SVM model and TF-IDF vectorizer.
-
Clone the Repository:
git clone https://github.com/mmnabeel317/Phishing_Guard.git cd Phishing-detector -
Set Up Virtual Environment:
python -m venv venv venv\Scripts\activate # On Windows
-
Install Dependencies:
pip install -r requirements.txt
- Ensures
pandas,scikit-learn,joblib,flask, and others are installed.
- Ensures
-
Verify Files:
- Ensure
Phishing_Email.csv,test_emails.csv, scripts, andtemplates/are present.
- Ensure
python preprocess_data.py- Inputs:
Phishing_Email.csv - Outputs:
processed_data.csv,invalid_rows.csv
python train_model.py- Inputs:
processed_data.csv - Outputs:
model.pkl,vectorizer.pkl - Displays: Accuracy, precision, recall, and F1-score (~96% accuracy).
python predict.py --input test_emails.csv --output predictions.csv- Inputs:
test_emails.csv - Outputs:
predictions.csv
python predict.py- Enter multi-line emails, type
END_EMAILto separate, press Enter twice to finish. - Example:
Subject: You’re a Winner! Congratulations! You’ve won a $1,000 gift card. Click here to claim your prize: http://win-rewards.com Hurry, offer expires in 24 hours! END_EMAIL [Enter] [Enter]
python app.py- Open
http://127.0.0.1:5000in a browser. - Paste an email into the textarea and click "Classify Email" to see the classification, confidence, and top keywords.
- Example Emails:
- Legitimate:
Subject: Monthly Strategy Meeting Dear Team, Our monthly strategy meeting is scheduled for Friday at 11 AM in Conference Room A. Please review the agenda attached and come prepared with your updates. Best regards, Amanda - Phishing:
Subject: Urgent: Account Verification Your account needs verification. Click here: http://secure-login.com
- Legitimate:
- Challenge: Misclassification of the “Invoice Overdue” phishing email as Safe.
- Solution: Switched to SVM, preserved URLs in preprocessing, and added class weights to handle imbalance.
- Challenge: Misclassification of short, neutral Legitimate emails (e.g., meeting invites).
- Solution: Cleaned
processed_data.csvto remove noisy Legitimate emails and added probability calibration to the SVM model. - Challenge: Basic web UI lacking visual appeal.
- Solution: Implemented a modern, responsive UI with a black, white, and purple aesthetic using Tailwind CSS and Inter font.
- Model Performance: ~96% accuracy, with high precision and recall for phishing detection.
- Test Results: Correctly classified 15/15 emails in
test_emails.csvand interactive inputs, including edge cases. - Web App: Provides accurate, real-time classification with confidence scores (e.g., ~70-90% for Legitimate, ~95-99% for Phishing) and a polished UI.
- Artifacts: All scripts, data, models, and the web app are included for reproducibility.
- Deep Learning Model: We can add BERT to train the model using deep learning.
- Multi-language support: Adding multi language support would help avoid phishing attacks in regional areas as well.
- Integration with Gmail/Outlook: Implementing a web extension would automate the process of checking emails for phishing attacks.
- Built for the Quantam Breach hackathon .
- Uses
scikit-learnfor machine learning,Flaskfor the web app, andTailwind CSSfor styling.
- Repository: https://github.com/mmnabeel317/Phishing_Guard
- Youtube: Working Demo