Skip to content

feat: Phase 3 — decoder-only transformer forward pass - #3

Merged
aarambh-darshan merged 2 commits into
mainfrom
feat/phase-3-model-architecture
Jul 30, 2026
Merged

aarambh-darshan merged 2 commits into
mainfrom
feat/phase-3-model-architecture

Conversation

@aarambh-darshan

Copy link
Copy Markdown
Member

Summary

Implements Phase 3 — Model Architecture from ROADMAP.md (lines 203–243).

Full decoder-only transformer forward pass: token IDs in, [batch, seq, vocab_size] logits out — verified on both nano and small configs with random weights.


Architecture (ARCHITECTURE.md §5)

token_ids [B, T]
  → TokenPositionEmbedding   token embed + learned absolute position embed (summed)
  → TransformerBlock × N     pre-norm RMSNorm → CausalSelfAttention → residual
                             pre-norm RMSNorm → SwiGluFfn             → residual
  → RMSNorm (final)
  → LM head                  n_embd → vocab_size (weight-tied to token embedding)
  → logits [B, T, vocab_size]

Files Added / Changed

File Role
src/model/mod.rs Module declarations + PocModel re-export
src/model/embedding.rs TokenPositionEmbedding — two Embedding tables summed element-wise
src/model/ffn.rs SwiGluFfn — silu(x@Wgate)*(x@Wup)@Wdown, no bias
src/model/attention.rs CausalSelfAttention — standard MHA + additive causal mask, no GQA
src/model/block.rs TransformerBlock — pre-norm + residual for both sublayers
src/model/transformer.rs PocModel — full pipeline + PocModel::random_init test helper
src/main.rs Added mod model; declaration

Key Design Decisions

Weight tying

When cfg.tie_embeddings == true (default for both presets), the LM head reuses the token embedding weight matrix via reshape + matmul:

x [B,T,C] → [B*T,C] @ embed_weight^T [C,V] → [B*T,V] → reshape [B,T,V]

No separate lm_head.weight tensor is allocated, halving the output-layer parameter count.

Causal mask

Built inside forward from x.device() so it always lives on the correct device. Uses additive -inf mask (not multiplicative) applied before softmax. Explicitly expand()-ed to [B, n_head, T, T] — candle does not auto-broadcast via +.

No GQA, no RoPE

Standard MHA with full Q/K/V per head. Learned absolute position embeddings. Both choices follow ARCHITECTURE.md §5.2 and §5.4 — they minimise confounds in the forgetting measurement that starts at Phase 7.


Tests — 27/27 pass

Test What it verifies
test_forward_shape_nano [2,16] → [2,16,2000] ✅
test_forward_shape_small [1,32] → [1,32,4000] ✅
test_causal_mask_end_to_end logits at pos 0–1 bit-identical when only pos 2+ tokens differ ✅
test_no_nan_inf_nano no NaN/Inf on non-trivial token IDs, nano ✅
test_no_nan_inf_small no NaN/Inf on non-trivial token IDs, small ✅
test_weight_tying_enabled_by_default nano preset has tie_embeddings=true ✅
test_causal_mask_isolation (attention unit) attention-level mask isolation ✅
+ 20 prior Phase 0–2 tests all still passing ✅
cargo test   →  27 passed; 0 failed
cargo clippy -- -D warnings  →  0 warnings

CI

The existing CI (.github/workflows/ci.yml) runs on this PR:

Job What it runs
msrv cargo check on Rust 1.89.0
quality cargo fmt --check → cargo check → cargo clippy -D warnings → cargo test → release build → CLI smoke
security-audit rustsec/audit-check on dependencies

All three jobs are expected to pass — local verification already confirms build, clippy, and test results match CI requirements.


Next Phase

Phase 4 — Stage 1 training + checkpoint save

  • src/checkpoint.rs — save_checkpoint / load_checkpoint (real .safetensors)
  • src/train.rs — AdamW training loop, train_stage1
  • main.rs — train-stage1 CLI subcommand

Milestone: cargo run -- train-stage1 --config nano --steps 2000 converges and writes checkpoints/stage1/.

Implements the full model architecture as specified in ARCHITECTURE.md §5:

  token_ids [B,T]
    → TokenPositionEmbedding   (token embed + learned absolute position embed)
    → TransformerBlock × N     (pre-norm RMSNorm + CausalSelfAttention + SwiGluFfn + residuals)
    → RMSNorm (final)
    → LM head (n_embd → vocab_size, weight-tied to token embedding)
    → logits [B, T, vocab_size]

New files
─────────
src/model/mod.rs         — module declarations + PocModel re-export
src/model/embedding.rs   — TokenPositionEmbedding: two Embedding tables summed
src/model/ffn.rs         — SwiGluFfn: silu(x@Wgate)*(x@Wup)@wdown, no bias
src/model/attention.rs   — CausalSelfAttention: MHA + causal mask (no GQA)
src/model/block.rs       — TransformerBlock: pre-norm + residual both sublayers
src/model/transformer.rs — PocModel: full pipeline + PocModel::random_init helper

Modified files
──────────────
src/main.rs              — add 'mod model;' declaration

Architecture decisions
──────────────────────
- Standard MHA (no GQA) — simpler weight layout for forgetting measurement
- Learned absolute position embeddings (not RoPE) — fewer confounds at seq≤128
- Weight tying via reshape+matmul: x[B,T,C]→[B*T,C] @ embed_weight^T → [B,T,V]
- Causal mask: additive -inf upper triangle, built per-forward from x.device()
- Candle broadcasting note: expand() required explicitly; + does not auto-broadcast

Tests (27/27 pass, 0 clippy warnings)
──────────────────────────────────────
✓ Milestone 1: forward shape [B,T,vocab_size] — nano + small configs
✓ Milestone 2: causal mask — logits pos 0-1 unchanged when pos 2+ tokens differ
✓ Milestone 3: no NaN/Inf on random input — nano + small configs
✓ All 11 prior tests (Phase 0–2) still pass

Closes Phase 3 milestone. Next: Phase 4 — Stage 1 training + checkpoint save.

git tag v0.3.0
CHANGELOG.md
  - Added Phase 3 entry under [Unreleased]:
    all 6 new model files, architecture decisions (no GQA/RoPE, weight
    tying via reshape+matmul, causal mask expand), candle gotchas, and
    full test results table (27/27 pass, 0 clippy warnings)

ROADMAP.md
  - Phase 3 marked ✅ DONE in phase map quick-reference
  - All task checkboxes [x] — embedding, attention, ffn, block, transformer
  - Test checkboxes [x] — shape, causal mask, NaN/Inf
  - Milestone block updated: 27/27 tests pass, commit done, tag pending

README.md
  - Project Structure tree expanded with model/ subfiles and phase labels
  - Status section rewritten: Phase 3 complete, PR #3 linked
  - Phase progress table added (0–12) with Done/Next/Planned indicators

Also: cargo fmt applied to all model source files (fixes CI fmt check)
@aarambh-darshan
aarambh-darshan merged commit 66ef656 into main Jul 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant