Repository navigation
feat: Phase 2 — Synthetic names-facts corpus + probe generator - #2
Merged
Merged
Conversation
Implements src/dataset/names_facts.rs and src/dataset/mod.rs, the complete
Stage 1 dataset pipeline required by ARCHITECTURE.md §8.
## What's added
### Types
- Example { name, attribute, value, text } — serde Serialize+Deserialize
- ClozeProbe { name, attribute, prompt, expected } — serde Serialize+Deserialize
expected has a leading space (' blue') matching ByteLevel tokenizer encoding,
so Phase 7 token-level accuracy comparisons need zero post-processing.
### Functions
- generate_corpus(n, seed) deterministic via StdRng; each (name,attr) once
- split_train_probe(ex, frac, seed) zero-overlap on (name,attr) keys
- to_cloze_probes(examples) strips value, preserves leading space
- write_jsonl<T> / read_jsonl<T> generic .jsonl I/O with parent-dir creation
### NamesFactsCorpus builder
NamesFactsCorpus::generate(n, fraction, fp_n, seed)
.save() writes:
data/stage1_names_facts.jsonl (train)
data/stage1_probes.jsonl (held-out probe)
data/stage1_forgetting_probes.jsonl (cloze probes from train subset)
### Static data
200 fictional first names (culturally diverse, #[rustfmt::skip] tabular layout)
4 attribute categories x 10 values each = 8,000 max unique triples
### Tests — 5 new (11 total, all pass)
test_generate_corpus_deterministic same seed → identical Vec
test_split_no_overlap zero (name,attr) overlap train/probe
test_cloze_probe_format format rules + sentence reconstruction
test_write_read_jsonl_roundtrip .jsonl write/read preserves all fields
test_names_facts_corpus_generate sizes and forgetting-probe membership
## CI (all gates pass locally)
cargo fmt --check ✅
cargo check --all-targets ✅
cargo clippy -- -D warnings ✅
cargo test --no-fail-fast ✅ 11/11
cargo build --release ✅
CLI smoke ✅
## Housekeeping
- src/main.rs: add mod dataset;
- .gitignore: add /data/ and /checkpoints/ (runtime-generated, not source)
- ROADMAP.md: Phase 2 fully checked off; Phase Map updated to ✅ DONE
- CHANGELOG.md: Phase 2 entry added
Replace manual env::args dispatch with a full clap 4.6.4 derive-based CLI.
## Changes
### Cargo.toml
- Pin clap to 4.6.4 (latest stable) with derive feature
### src/main.rs — complete rewrite
Top-level binary now exposes a rich CLI:
continual-learning-poc --version
continual-learning-poc --help (shows full 12-phase roadmap)
continual-learning-poc generate-data [OPTIONS]
generate-data flags (all with typed defaults + long_help):
--n-examples <N> total examples before split [default: 800]
--probe-fraction <F> held-out probe fraction [default: 0.2]
--forgetting-probes <N> train examples → cloze probes [default: 100]
--seed <SEED> StdRng seed [default: 42]
--output-dir <DIR> output directory [default: data]
Output after generation:
Generating Stage 1 dataset (n=800, probe=20%, fp=100, seed=42) ...
Done.
train corpus: 640 examples → data/stage1_names_facts.jsonl
probe set: 160 examples → data/stage1_probes.jsonl
forgetting probes: 100 entries → data/stage1_forgetting_probes.jsonl
Tip: re-evaluate data/stage1_forgetting_probes.jsonl after Stage 2 training.
CI: all 11 tests pass, clippy -D warnings clean, release build OK.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements the complete Stage 1 dataset pipeline (ARCHITECTURE.md §8).
Three deterministic, probe-based datasets are generated and written as
.jsonl:data/stage1_names_facts.jsonlVec<Example>)data/stage1_probes.jsonldata/stage1_forgetting_probes.jsonlNew files
src/dataset/mod.rs— module entry point, re-exports public typessrc/dataset/names_facts.rs— full implementationKey design decisions
expectedhas a leading space (" blue", not"blue") — matches ByteLevel tokenizer's word-initial encoding; Phase 7 accuracy comparison needs zero post-processing.split_train_probesplits on(name, attribute)keys, not individual examples — guarantees the model is never trained on any(name, attribute)pair it will later be probed on.#[rustfmt::skip]on the 200-name static table — prevents rustfmt oscillation on large tabular string literals.Tests — 5 new (11 total)
test_generate_corpus_deterministicVec; different seed → different ordertest_split_no_overlap(name, attribute)overlap between train and probetest_cloze_probe_formattest_write_read_jsonl_roundtrip.jsonlwrite/read preserves all fieldstest_names_facts_corpus_generateCI
All gates pass locally:
cargo fmt --check✅cargo check --all-targets --locked✅cargo clippy --all-targets --locked -- -D warnings✅cargo test --no-fail-fast --locked✅ 11/11cargo build --release --locked✅--version/--help) ✅Housekeeping
src/main.rs:mod dataset;added.gitignore:/data/and/checkpoints/excluded (runtime-generated)ROADMAP.md: Phase 2 fully checked off; Phase Map → ✅ DONECHANGELOG.md: Phase 2 entry prependedNext
Phase 3 — Model Architecture (
src/model/): token + position embedding, causal multi-head attention, SwiGLU FFN, TransformerBlock, full decoder.