Deep learning · Computer vision
Visual Emotion Recognition
A complete facial emotion recognition system, from exploring the data to explainable, deployable models running in a real-time API and a mobile app.
- Role
- ML Engineer
- Year
- 2025
- Focus
- Deep learning
- Stack
- 7 technologies
- PyTorch
- timm
- Stable Diffusion
- Grad-CAM
- SHAP
- ONNX
- Hugging Face
43.7K
Images · 6 classes
4
Explainability methods
3
Deployment targets
The problem
Emotion datasets are heavily imbalanced: 'happy' alone is about 31% of this one. Models trained on them under-perform on rarer emotions and are hard to trust without explanations.
Overview
The project covers the whole ML lifecycle. EDA and quality scanning, balancing with augmentation, and synthetic data from Stable Diffusion to expand minority classes, with strict quality filters so generated faces don't skew the distribution.
Models are trained with staged transfer learning and checked with four explainability methods. They're exported to ONNX and Hugging Face, then served through a FastAPI real-time app and a Flutter mobile app.
Architecture
ML lifecycle
- 01
EDA
Class counts, formats, and quality metrics across 43,756 grayscale images.
- 02
Balance & augment
Oversample minority classes with augmentation and build stratified train/val/test splits.
- 03
Synthetic data
Stable Diffusion v1.5Generate class-specific faces with filters: single-face detection, blur threshold, perceptual-hash dedupe, and a cap on synthetic share.
- 04
Transfer learning
CNN / ViTTwo-stage fine-tuning (head warm-up, then full unfreeze), mixed precision, EMA, label smoothing, MixUp/CutMix.
- 05
Explain
Grad-CAM, Grad-CAM++, SHAP, and LIME for sanity checks, bias audits, and error analysis.
- 06
Deploy
ONNXExport to ONNX and the Hugging Face Hub, served by a FastAPI real-time app and a Flutter mobile app.
Results
Class distribution
Why balancing and synthetic generation were needed.
- Happy30.6%
- Neutral18.9%
- Sad17.2%
- Angry11.6%
- Surprised11.3%
- Fearful10.5%