Add reproducible MLX benchmark protocol - #33
Merged
Merged
Conversation
There was a problem hiding this comment.
Pull request overview
This PR introduces a reproducible benchmarking protocol for MLX Transformers by adding a dedicated mlx-transformers-benchmark CLI, bundling deterministic protocol v1 scenarios, and documenting how to generate/validate/share benchmark results without leaking private machine details.
Changes:
- Added
mlx-transformers-benchmarkCLI (run,validate,render) plus validation logic for scenario/result determinism and privacy constraints. - Added built-in protocol v1 scenarios and a packaged JSON schema under
benchmark_data/. - Updated docs/config to ship benchmark assets, document usage, and harden the legacy example by disabling
trust_remote_code.
Reviewed changes
Copilot reviewed 11 out of 11 changed files in this pull request and generated 6 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/test_benchmark.py | Adds protocol/CLI tests for offline run + validation + render output. |
| src/mlx_transformers/benchmark.py | Implements CLI, scenario/result validation, deterministic run loop, and log rendering. |
| src/mlx_transformers/benchmark_data/short-decode-128.json | Adds built-in protocol v1 scenario (64 prefill / 128 decode). |
| src/mlx_transformers/benchmark_data/prefill-512-decode-64.json | Adds built-in protocol v1 scenario (512 prefill / 64 decode). |
| src/mlx_transformers/benchmark_data/result-v1.schema.json | Adds packaged JSON schema for protocol v1 results. |
| README.md | Documents reproducible benchmark CLI usage and updates test baseline counts. |
| pyproject.toml | Registers the new CLI entry point and includes benchmark JSON package data. |
| MANIFEST.in | Ensures benchmark markdown docs are included in sdist. |
| examples/text_generation/benchmark_generation.py | Disables remote code execution in the legacy benchmark example. |
| CONTRIBUTING.md | Updates expected offline test baseline counts. |
| benchmarks/REPRODUCIBILITY_LOG.md | Adds initial validated reference results in the reproducibility log. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+111
to
+117
| scenario = result["scenario"]["definition"] | ||
| validate_scenario(scenario) | ||
| expected_hash = hashlib.sha256( | ||
| json.dumps(scenario, sort_keys=True, separators=(",", ":")).encode() | ||
| ).hexdigest() | ||
| if result["scenario"].get("sha256") != expected_hash: | ||
| raise ValueError("Scenario checksum does not match its embedded definition.") |
Comment on lines
+192
to
+197
| bos_token_id = getattr(tokenizer, "bos_token_id", None) | ||
| if bos_token_id is not None: | ||
| token_ids = [bos_token_id, *token_ids] | ||
| target = scenario["prompt_tokens"] | ||
| repeated = (token_ids * ((target + len(token_ids) - 1) // len(token_ids)))[:target] | ||
| input_ids = mx.array([repeated], dtype=mx.int32) |
Comment on lines
+219
to
+223
| first_token = next(tokens) | ||
| mx.eval(first_token) | ||
| token_hash = hashlib.sha256() | ||
| token_hash.update(int(first_token.item()).to_bytes(8, "little", signed=True)) | ||
| first_token_at = time.perf_counter() |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
mlx-transformers-benchmarkCLI withrun,validate, andrendercommandsWhy
MLX performance claims need a portable protocol and inspectable evidence. This gives maintainers and community members fixed workloads, machine-readable results, and validation rules so results from different Apple silicon systems can be reproduced and shared responsibly.
Local M2 Max reference
Using the cached
mlx-community/Phi-3-mini-4k-instruct-4bitsnapshot at revision5b3819ed6317784fb20eddeae9bed984f778d0d0, with one warmup and five measured runs:short-decode-128prefill-512-decode-64Both runs were executed on an Apple M2 Max with 32 GB unified memory in strict offline mode. All five generated-token checksums matched for each scenario.
Validation
ruff check .ruff format --check .twine checkon both artifacts