Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Unveiling Music Patterns: Clustering Spotify Data with Bayesian Gaussian and Hierarchical Mixture Models

A Bayesian unsupervised learning study that uncovers latent structure in Spotify track features using two probabilistic clustering approaches — Bayesian Gaussian Mixture Models (B-GMM) and Hierarchical Bayesian Mixture Models (HBMM) — applied to a stratified sample of 40,000 tracks across 125 genres.

Course: STATS 551 — Bayesian Statistics · University of Michigan, Ann Arbor

📄 Read the report · 📊 View the slides

Overview

Traditional clustering assigns tracks to a single genre, but music is fluid. This project models Spotify tracks probabilistically, allowing soft cluster assignments where a track can belong to multiple clusters with varying probability. Two models are compared:

  • B-GMM: A flat Bayesian Gaussian Mixture Model with K clusters, using Skew Student's t-distributions to handle asymmetric and heavy-tailed feature distributions.
  • HBMM: A two-layer hierarchical extension that captures both broad genre clusters (top-level) and finer sub-genre distinctions (sub-level) within a single unified model.

Both models are fit using Automatic Differentiation Variational Inference (ADVI) in PyMC for scalability, with convergence monitored via ELBO and validated through posterior predictive checks and trace plots.

Results Summary

Aspect B-GMM HBMM
Training time 3.2 hrs 6.1 hrs (1.9× slower)
Clusters found 5 broad genres 6 meta-clusters + 3 sub-genres
Genre purity 0.89 0.81 genre + 0.73 subgenre F1
Convergence ~6,000 iterations (ELBO −110,000) ~10,000 iterations (ELBO −115,000)
Mean ESS 969 ± 87 1892 ± 142
Hybrid track handling Struggles Captures 38% more border tracks

Key findings:

  • B-GMM is faster and better suited for broad genre analysis, with strong posterior predictive alignment (KL divergence < 0.05).
  • HBMM provides superior sub-genre granularity and better handles outlier tracks and hybrid styles, at the cost of higher compute and optimization complexity.
  • Task complexity and LLM grounding are both dominant factors — loudness/energy and danceability/valence were the strongest correlated feature pairs.
  • Popularity does not correlate strongly with musical features, suggesting non-acoustic factors drive chart performance.

Project Structure

spotify-clustering/
├── notebooks/
│   ├── sampling.ipynb          # Stratified sampling from the full 100K dataset
│   ├── bgmm.ipynb              # Bayesian Gaussian Mixture Model (B-GMM)
│   ├── hbmm.ipynb              # Hierarchical Bayesian Mixture Model (HBMM)
│   └── experiments/            # Exploratory work: MCMC pilots, Dirichlet PMM, outlier handling
├── data/
│   └── spotify_sampled.csv     # Stratified sample of 40,000 tracks (saved locally)
├── images/                     # All figures from the report (figures 1–11)
└── report/
    ├── STATS551_Final_Project.pdf
    └── STATS551_Presentation.pptx

Dataset

The full dataset (~100,000 Spotify tracks, 20 features, 125 genres) is sourced from Hugging Face: maharshipandya/spotify-tracks-dataset

A stratified sample of 40,000 tracks was drawn based on track_genre, popularity, and duration to preserve diversity while keeping ADVI tractable. The sampled data is saved to data/spotify_sampled.csv. To regenerate it, run notebooks/sampling.ipynb.

Setup

Prerequisites

  • Python 3.10+
  • PyMC 5+

Install dependencies

pip install pymc arviz pandas numpy scikit-learn matplotlib seaborn datasets

Run the models

  1. (Optional) Regenerate the sample: run notebooks/sampling.ipynb
  2. Fit B-GMM: run notebooks/bgmm.ipynb
  3. Fit HBMM: run notebooks/hbmm.ipynb

⚠️ Both models are computationally intensive. B-GMM takes ~3.2 hrs and HBMM ~6.1 hrs on CPU. A GPU or cloud runtime (e.g., Google Colab) is recommended.

Model Details

B-GMM

Mixture of K Gaussian components with Skew Student's t-distributions for robustness to heavy tails:

$$X_i \sim \sum_{k=1}^{K} \pi_k \cdot \text{SkewStudentT}(\mu_k, \sigma_k, a=2, b=2)$$

Priors: $\mu_k \sim \mathcal{N}(0,1)$, $\Sigma_k \sim \text{LogN}(0,1)$, $\pi \sim \text{Dir}(1_K)$

HBMM

Two-level nested structure capturing genre and sub-genre simultaneously:

$$X_i \sim \sum_{j=1}^{K_1} \sum_{k=1}^{K_2} \pi_j \cdot \pi_{k|j} \cdot \mathcal{N}(X_i | \mu_{j,k}, \Sigma_{j,k})$$

Top-level means share a hyperprior ($\sigma_\text{top} = 10$); sub-level means inherit from parent clusters with local variance ($\sigma_\text{sub-hyper} = 5$).

Inference

Both models use ADVI (mean-field variational inference) with adaptive learning rates and ΔELBO < 0.1% over 500 iterations as the convergence criterion — a 3.2× speedup over MCMC in pilot tests.

Figures

Figure Description
Fig. 1 PCA elbow plot — 4 components capture >95% variance
Fig. 2 B-GMM ELBO convergence
Fig. 3 B-GMM trace plots
Fig. 4 B-GMM feature distributions by cluster
Fig. 5 B-GMM t-SNE cluster visualization
Fig. 6 B-GMM posterior predictive checks
Fig. 7 HBMM ELBO convergence
Fig. 8 HBMM trace plots
Fig. 9 HBMM feature distributions by cluster
Fig. 10 HBMM t-SNE cluster visualization
Fig. 11 HBMM posterior predictive checks

Tech Stack

  • Modeling: PyMC, ArviZ
  • Data: Hugging Face datasets, pandas, numpy
  • Preprocessing: scikit-learn (RobustScaler, PCA, LabelEncoder)
  • Visualization: matplotlib, seaborn, t-SNE

About

A Bayesian unsupervised learning study that uncovers latent structure in Spotify track features using two probabilistic clustering approaches. Bayesian Gaussian Mixture Models (B-GMM) and Hierarchical Bayesian Mixture Models (HBMM), applied to a stratified sample of 40,000 tracks across 125 genres.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages