Skip to content

feat: Phase 5 — Evaluation Harness, Stage 1 baseline recorded - #6

Merged
aarambh-darshan merged 1 commit into
mainfrom
feat/phase-5-eval-harness
Jul 30, 2026
Merged

aarambh-darshan merged 1 commit into
mainfrom
feat/phase-5-eval-harness

Conversation

@aarambh-darshan

Copy link
Copy Markdown
Member

Summary

Phase 5 of the continual-learning-poc roadmap. Implements the reusable evaluation harness that scores any checkpoint against a probe file — this is the measurement instrument every Phase 6–11 forgetting comparison depends on.


What Changed

src/eval.rs [NEW]

Symbol Description
probe_accuracy(model, tokenizer, probes, device) Greedy first-token exact-match over all probes → EvalReport
perplexity(model, tokenizer, examples, device) Token-length-weighted CE (batch=1) → exp(mean_nll)
run_eval(checkpoint_dir, probe_path, device) load_checkpoint → probe_accuracy + perplexity in one call
EvalReport probe_accuracy, perplexity, n_correct, n_total, results
ProbeResult Per-probe: name, attribute, prompt, expected, predicted, correct

Probe accuracy algorithm:

ids = [BOS_ID] ++ tokenizer.encode(prompt)
logits = model.forward(ids)       → [1, T, V]
pred_id = argmax(logits[0, T-1, :])
correct = tokenizer.decode([pred_id]).trim() == expected.trim()

Design decisions:

  • First-token exact-match — sufficient for this dataset (all values are single-word after ByteLevel BPE); avoids beam-search complexity
  • Batch-size-1 perplexity loop — no padding logic, flat memory, < 1 s per 640-sentence corpus on CPU
  • run_eval reconstructs Example from ClozeProbe — avoids a second data file; prompt + expected + "." == original training sentence

src/main.rs [MODIFIED]

Added mod eval; and the eval subcommand:

continual-learning-poc eval [OPTIONS]

  --checkpoint <DIR>    [default: checkpoints/stage1]
  --probes <FILE>       [default: data/stage1_forgetting_probes.jsonl]
  --verbose             Print per-probe ✅/❌ detail

Output format:

Evaluation Report
=================
probe_accuracy :   44.0%   (44 / 100)
perplexity     :   14.471

Per-probe detail (--verbose):
  ✅  Alice         favorite_color    expected=" purple"  predicted=" purple"
  ❌  Lars          hobby             expected=" hiking"  predicted=" a"
  …

Tests — 5 new (37 total)

Test Assertion
eval::tests::test_probe_accuracy_trained_model 300-step fixture model → ≥ 90% accuracy
eval::tests::test_probe_accuracy_random_init random-init → < 50% (sanity floor)
eval::tests::test_perplexity_decreases_with_training ppl(trained) < ppl(random)
eval::tests::test_eval_report_fields n_correct + n_wrong == n_total; accuracy field == n_correct/n_total
eval::tests::test_run_eval_round_trip save checkpoint → run_eval → finite perplexity, n_total > 0

Stage 1 Baseline — Milestone Result

$ cargo run --release -- eval     --checkpoint checkpoints/stage1     --probes data/stage1_forgetting_probes.jsonl

probe_accuracy :   44.0%   (44 / 100)
perplexity     :   14.471

44% is the Stage 1 baseline — recorded once, compared against every Phase 7 Stage 2 run to quantify catastrophic forgetting.

The 44% reflects partial memorisation after 2000 steps. The model has learned attribute-value distributions well but not every specific (name, attribute) → value binding. This is a valid baseline — any drop after Stage 2 is measurable forgetting. Higher accuracy can be achieved with more steps if needed before Phase 6.


CI

  • cargo fmt --check ✅
  • cargo clippy --all-targets -- -D warnings ✅ (0 warnings)
  • cargo test --no-fail-fast ✅ 37/37 pass
  • cargo build --release ✅

Docs Updated

  • CHANGELOG.md — Phase 5 entry with arch decisions, test table, milestone result
  • ROADMAP.md — all [ ] → [x], milestone block filled, Phase 5 marked ✅ DONE
  • README.md — status updated to Phase 5 ✅ Done, Phase 6 🔜 Next

Next: Phase 6 — Stage 2 Corpus + Naive Continual Training

src/dataset/stage2_corpus.rs: disjoint-domain corpus (capital-city facts);
train_stage2 warm-started from Stage 1 checkpoint; train-stage2 --strategy naive CLI.

src/eval.rs  [NEW]
  - probe_accuracy(model, tokenizer, probes, device) → EvalReport
      greedy first-token exact-match; correct = predicted.trim() == expected.trim()
  - perplexity(model, tokenizer, examples, device) → f32
      token-length-weighted CE (batch=1); returns exp(mean_nll)
  - run_eval(checkpoint_dir, probe_path, device) → EvalReport
      load_checkpoint → probe_accuracy + perplexity in one call
  - EvalReport { probe_accuracy, perplexity, n_correct, n_total, results }
  - ProbeResult { name, attribute, prompt, expected, predicted, correct }
  - 5 milestone tests (37 total)

src/main.rs  [MODIFIED]
  - mod eval; declaration
  - Eval(EvalArgs) variant in Commands
  - EvalArgs: --checkpoint, --probes, --verbose
  - cmd_eval: validates paths, calls run_eval, prints accuracy + perplexity +
    optional per-probe ✅/❌ table; hints ≥85% threshold for Phase 6 readiness

Architecture decisions:
  - first-token exact-match: sufficient for this single-word-value dataset
  - batch=1 perplexity loop: no padding complexity, flat memory, <1s on CPU
  - run_eval reconstructs Example from ClozeProbe (prompt + expected + '.')
    to avoid a second data file; no information loss for single-value facts

Milestone (run against real 2000-step Stage 1 checkpoint):
  cargo run --release -- eval       --checkpoint checkpoints/stage1       --probes data/stage1_forgetting_probes.jsonl
  probe_accuracy :   44.0%   (44 / 100)
  perplexity     :   14.471
  Stage 1 baseline recorded — reference for every Phase 7 forgetting delta.

CI: 37/37 tests pass · 0 clippy warnings · cargo build --release ✅

Docs: CHANGELOG + ROADMAP (Phase 5 ✅ DONE) + README (status + phase table)
@aarambh-darshan
aarambh-darshan merged commit 4b29fff into main Jul 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant