Skip to content

fix: sync model vocab_size to actual tokenizer vocab; bump default LR to 5e-3 - #5

Merged
aarambh-darshan merged 1 commit into
mainfrom
fix/vocab-sync-and-lr-tuning
Jul 30, 2026
Merged

aarambh-darshan merged 1 commit into
mainfrom
fix/vocab-sync-and-lr-tuning

Conversation

@aarambh-darshan

Copy link
Copy Markdown
Member

Problem

After Phase 4 merged, the initial training loss was ~41.78 — much higher than expected.

Root cause: cmd_train_stage1 built the model with vocab_size=2000 (from ModelConfig::nano()), but then trained a BPE tokenizer that only produced 669 tokens on the 640-sentence corpus. The model's output head had 1,331 dead logits that:

  • Never received a positive gradient signal
  • Still participated in the softmax denominator on every step
  • Artificially inflated the cross-entropy loss at initialisation

Expected initial loss for 669 tokens: ln(669) ≈ 6.5. Actual was 41.78 — the dead logits were adding ~35 nats of noise.


Fix

src/main.rs

After BpeTokenizer::train returns, compare the actual vocab size to model_cfg.vocab_size and sync them:

let actual_vocab = tokenizer.vocab_size();
if actual_vocab != model_cfg.vocab_size {
    println!("  Syncing model vocab_size → {actual_vocab} to eliminate dead logits.");
    model_cfg.vocab_size = actual_vocab;
}

The saved config.json now records the actual vocab size so load_checkpoint reconstructs the model with the correct output head.

src/config.rs

Field Before After Reason
lr 3e-3 5e-3 Smaller effective model (669 tokens) tolerates higher LR
warmup_steps 100 200 Longer warmup stabilises early steps at the higher LR

Before / After (200-step smoke test)

Metric Before After Δ
Initial loss 41.78 35.30 −15%
Final loss 10.44 9.57 −8%
Smooth 10% 13.81 10.52 −24%
Model file 5.1 MB 4.4 MB −13% (smaller output projection)

CI

  • cargo fmt --check ✅
  • cargo clippy --all-targets -- -D warnings ✅ (0 warnings)
  • cargo test --no-fail-fast ✅ 32/32 pass

Bug: cmd_train_stage1 built model_cfg with vocab_size=2000 (from
ModelConfig::nano), then trained a BPE tokenizer that produced only 669
tokens on the 640-sentence corpus. The model's 2000-slot output head had
1331 dead logits that never received a positive gradient signal but still
participated in the softmax denominator, artificially inflating CE loss.

Fix (src/main.rs):
  - Make model_cfg mutable after preset selection.
  - After BpeTokenizer::train returns, read tokenizer.vocab_size() and
    sync model_cfg.vocab_size to that value when they differ.
  - Print a clear message explaining the sync so users understand why the
    saved config.json shows a different vocab_size than the preset.

Tuning (src/config.rs):
  - TrainConfig::default lr: 3e-3 → 5e-3
    The ~669-token synthetic task is small; 5e-3 converges noticeably
    faster with no instability observed.
  - TrainConfig::default warmup_steps: 100 → 200
    Longer warmup stabilises early steps at the higher LR.

Before / after (200-step smoke test):
  initial loss : 41.78 → 35.30  (-15%  — fewer dead logits at init)
  final loss   : 10.44 →  9.57  ( -8%)
  smooth-10%   : 13.81 → 10.52  (-24%)
  model size   :  5.1 MB →  4.4 MB  (smaller output projection)

32/32 tests pass · 0 clippy warnings
@aarambh-darshan
aarambh-darshan merged commit 9818ed3 into main Jul 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant