Unveiling Music Patterns: Clustering Spotify Data with Bayesian Gaussian and Hierarchical Mixture Models
A Bayesian unsupervised learning study that uncovers latent structure in Spotify track features using two probabilistic clustering approaches — Bayesian Gaussian Mixture Models (B-GMM) and Hierarchical Bayesian Mixture Models (HBMM) — applied to a stratified sample of 40,000 tracks across 125 genres.
Course: STATS 551 — Bayesian Statistics · University of Michigan, Ann Arbor
Traditional clustering assigns tracks to a single genre, but music is fluid. This project models Spotify tracks probabilistically, allowing soft cluster assignments where a track can belong to multiple clusters with varying probability. Two models are compared:
- B-GMM: A flat Bayesian Gaussian Mixture Model with K clusters, using Skew Student's t-distributions to handle asymmetric and heavy-tailed feature distributions.
- HBMM: A two-layer hierarchical extension that captures both broad genre clusters (top-level) and finer sub-genre distinctions (sub-level) within a single unified model.
Both models are fit using Automatic Differentiation Variational Inference (ADVI) in PyMC for scalability, with convergence monitored via ELBO and validated through posterior predictive checks and trace plots.
| Aspect | B-GMM | HBMM |
|---|---|---|
| Training time | 3.2 hrs | 6.1 hrs (1.9× slower) |
| Clusters found | 5 broad genres | 6 meta-clusters + 3 sub-genres |
| Genre purity | 0.89 | 0.81 genre + 0.73 subgenre F1 |
| Convergence | ~6,000 iterations (ELBO −110,000) | ~10,000 iterations (ELBO −115,000) |
| Mean ESS | 969 ± 87 | 1892 ± 142 |
| Hybrid track handling | Struggles | Captures 38% more border tracks |
Key findings:
- B-GMM is faster and better suited for broad genre analysis, with strong posterior predictive alignment (KL divergence < 0.05).
- HBMM provides superior sub-genre granularity and better handles outlier tracks and hybrid styles, at the cost of higher compute and optimization complexity.
- Task complexity and LLM grounding are both dominant factors — loudness/energy and danceability/valence were the strongest correlated feature pairs.
- Popularity does not correlate strongly with musical features, suggesting non-acoustic factors drive chart performance.
spotify-clustering/
├── notebooks/
│ ├── sampling.ipynb # Stratified sampling from the full 100K dataset
│ ├── bgmm.ipynb # Bayesian Gaussian Mixture Model (B-GMM)
│ ├── hbmm.ipynb # Hierarchical Bayesian Mixture Model (HBMM)
│ └── experiments/ # Exploratory work: MCMC pilots, Dirichlet PMM, outlier handling
├── data/
│ └── spotify_sampled.csv # Stratified sample of 40,000 tracks (saved locally)
├── images/ # All figures from the report (figures 1–11)
└── report/
├── STATS551_Final_Project.pdf
└── STATS551_Presentation.pptx
The full dataset (~100,000 Spotify tracks, 20 features, 125 genres) is sourced from Hugging Face:
maharshipandya/spotify-tracks-dataset
A stratified sample of 40,000 tracks was drawn based on track_genre, popularity, and duration to preserve diversity while keeping ADVI tractable. The sampled data is saved to data/spotify_sampled.csv. To regenerate it, run notebooks/sampling.ipynb.
- Python 3.10+
- PyMC 5+
pip install pymc arviz pandas numpy scikit-learn matplotlib seaborn datasets- (Optional) Regenerate the sample: run
notebooks/sampling.ipynb - Fit B-GMM: run
notebooks/bgmm.ipynb - Fit HBMM: run
notebooks/hbmm.ipynb
⚠️ Both models are computationally intensive. B-GMM takes ~3.2 hrs and HBMM ~6.1 hrs on CPU. A GPU or cloud runtime (e.g., Google Colab) is recommended.
Mixture of K Gaussian components with Skew Student's t-distributions for robustness to heavy tails:
Priors:
Two-level nested structure capturing genre and sub-genre simultaneously:
Top-level means share a hyperprior (
Both models use ADVI (mean-field variational inference) with adaptive learning rates and ΔELBO < 0.1% over 500 iterations as the convergence criterion — a 3.2× speedup over MCMC in pilot tests.
| Figure | Description |
|---|---|
| Fig. 1 | PCA elbow plot — 4 components capture >95% variance |
| Fig. 2 | B-GMM ELBO convergence |
| Fig. 3 | B-GMM trace plots |
| Fig. 4 | B-GMM feature distributions by cluster |
| Fig. 5 | B-GMM t-SNE cluster visualization |
| Fig. 6 | B-GMM posterior predictive checks |
| Fig. 7 | HBMM ELBO convergence |
| Fig. 8 | HBMM trace plots |
| Fig. 9 | HBMM feature distributions by cluster |
| Fig. 10 | HBMM t-SNE cluster visualization |
| Fig. 11 | HBMM posterior predictive checks |
- Modeling: PyMC, ArviZ
- Data: Hugging Face
datasets, pandas, numpy - Preprocessing: scikit-learn (RobustScaler, PCA, LabelEncoder)
- Visualization: matplotlib, seaborn, t-SNE