A machine learning model that classifies media (image, video, or audio) as Real or Fake using metadata and signal-quality features rather than raw pixel/audio analysis.
This project trains a Logistic Regression classifier on a metadata dataset describing media samples (e.g. lip-sync consistency, visual artifacts, lighting inconsistencies, compression level, platform, and generation method) to predict whether the sample is real or AI-generated ("fake").
The model expects a CSV file named deepfake_detection_metadata_dataset.csv with the following columns:
| Column | Description |
|---|---|
media_id |
Unique identifier for the media sample |
media_type |
Type of media (Image, Video, Audio) |
content_category |
Content category (News, Interview, Social Media, etc.) |
face_count |
Number of faces detected in the media |
audio_present |
Whether audio is present (Yes/No) |
lip_sync_score |
Score indicating lip-sync consistency |
visual_artifacts_score |
Score indicating presence of visual artifacts |
compression_level |
Compression level of the media file |
lighting_inconsistency_score |
Score indicating lighting inconsistencies |
source_platform |
Platform the media was sourced from |
generation_method |
Method used to generate fake media (GAN, Diffusion, VoiceClone, etc.), if applicable |
label |
Target variable — Real or Fake |
- Load & explore data — inspect shape, head/tail, dtypes, and missing values.
- Handle missing values —
generation_method(missing for real media) is imputed with its mode. - Encode categorical features — one-hot encoding via
pd.get_dummies()onmedia_type,content_category,audio_present,source_platform, andgeneration_method. - Train/test split — 80/20 split (
random_state=42). - Train model —
sklearn.linear_model.LogisticRegression. - Evaluate — confusion matrix and accuracy score on the held-out test set.
- Visualize — box plots comparing
lip_sync_score,visual_artifacts_score, andlighting_inconsistency_scorebetween real and fake samples. - Save model — trained model is serialized to
deepfake_model.pklwithjoblib.
pandas
numpy
matplotlib
seaborn
scikit-learn
joblib
Install with:
pip install pandas numpy matplotlib seaborn scikit-learn joblib- Place
deepfake_detection_metadata_dataset.csvin the project directory. - Run the notebook (
deepfake_model.ipynb) cell by cell, or export it to a script:jupyter nbconvert --to script deepfake_model.ipynb python deepfake_model.py
- The trained model is saved as
deepfake_model.pkl, and box plot figures are saved as<feature>_boxplot.png.
import joblib
model = joblib.load("deepfake_model.pkl")
predictions = model.predict(X_new)The model achieves high accuracy on the test split (confusion matrix and accuracy score are printed in the notebook).
Note: Accuracy came out at 100% on this dataset, which is unusually high for a real-world classifier. This is very likely because
generation_methodis only missing forRealsamples (and filled in forFakesamples), so its one-hot encoding leaks the label. Before using this model on new data, consider dropping or re-derivinggeneration_methodso the model relies on genuine signal-quality features (lip_sync_score,visual_artifacts_score,lighting_inconsistency_score,compression_level) instead.
- Trained on metadata/scores rather than raw media (no image, video frame, or audio waveform analysis).
- Potential label leakage via
generation_method(see note above) should be addressed before deployment. - No hyperparameter tuning, cross-validation, or regularization scaling was applied (
LogisticRegressionalso raised a convergence warning — consider scaling features or increasingmax_iter). - Could be extended with a CNN/RNN pipeline on raw video frames or audio spectrograms for true content-based deepfake detection.
Add a license of