Conversation
…, fold the contraction families onto one setup, move timing into the launcher
… their own dispatch
…im narrative docs
…unch targeting, drop dead skip ratio
…drop unused accessors
…it tables, host bookkeeping, terse comments
…uites, skip the 400-step training test on software adapters
… level into one group after extraction, row-per-workgroup scatter, packed coop slots in groups
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add reusable GPU training programs and a configurable browser transformer trainer. A forward/backward/optimizer graph compiles once, then reuses kernels, allocations, and device state across steps. The compiler owns indexing, fusion boundaries, buffer reuse, and synchronization.
Changes:
TrainingProgramwith persistent inputs, simultaneous state feedback, bounded asynchronous submissions, and independent exported snapshots. Logical indexing and workgroup ownership checks determine legal fusion; lifetime interference and size-based packing reuse storage.compiler-testsfeature.Measured on Apple M2 Max, f32:
The Session comparison measures the two execution paths on this branch. Tiny transformer timings exclude initialization, evaluation, sampling, and rendering. These are single-device measurements.
Validation on Apple M2 Max: the retained training/program/library/cache checks pass (30 tests); 58 CPU/GPU matmul and quantized cases, 36 normalization cases, and 32 sampling cases passed. The native BERT integration test passed with differential member checks enabled. Training-state parity also passed with compiler checks disabled. Release Chrome passed 73 targeted matmul/normalization/attention cases and all eight training capability/fallback modes, with matching losses and no GPU errors. Constructor tests cover cooperative scratch combinations, promoted reduction domains, absent CPU schedules, and equivalent output shapes. All 58 targeted convolution/pooling, normalization, and elementwise CPU/GPU cases pass with complete default shape sampling. The production browser configuration test trains a nine-block model and covers generation, attention images, reset, and pause/resume. Tests compare bitmap pixels and persisted cache formats directly; duplicate assertions and print-only diagnostics are removed. Workspace strict Clippy, CPU-only strict Clippy, production WASM checking, and formatting passed.
Model checks compare every parameter and Adam moment with Session across three architectures. A full default browser training run completed at 0.912 held-out loss and 72% next-character accuracy. These local checks do not replace the Linux/macOS/Windows CI results shown on this PR. Training tests are serialized within the Fusor workspace, with a finite timeout suitable for software GPUs.
TrainingProgramrequires a GPU and static, nonempty f32/u32/i32 shapes; browser matrix acceleration depends on experimental browser/adapter support. Large schedule domains are sampled in numerical conformance; small domains and cooperative resource-domain construction are checked exhaustively.