Skip to content

Add reproducible MLX benchmark protocol - #33

Merged
ToluClassics merged 5 commits into
mainfrom
agent/reproducible-benchmarks
Jul 29, 2026
Merged

ToluClassics merged 5 commits into
mainfrom
agent/reproducible-benchmarks

Conversation

@ToluClassics

@ToluClassics ToluClassics commented Jul 29, 2026 •

Copy link
Copy Markdown
Owner

What changed

  • add a packaged mlx-transformers-benchmark CLI with run, validate, and render commands
  • define deterministic protocol v1 scenarios for 64-token prefill / 128-token decode and 512-token prefill / 64-token decode
  • record raw runs, summary statistics, hardware and software metadata, quantization details, exact model and implementation revisions, and generated-token checksums
  • reject non-immutable revisions, nondeterministic output, and private path or credential markers
  • add a reproducibility log and packaged JSON schema
  • disable remote code in the legacy generation benchmark example

Why

MLX performance claims need a portable protocol and inspectable evidence. This gives maintainers and community members fixed workloads, machine-readable results, and validation rules so results from different Apple silicon systems can be reproduced and shared responsibly.

Local M2 Max reference

Using the cached mlx-community/Phi-3-mini-4k-instruct-4bit snapshot at revision 5b3819ed6317784fb20eddeae9bed984f778d0d0, with one warmup and five measured runs:

Scenario Median TTFT Median prefill Median decode Peak MLX memory
short-decode-128 63.6 ms 1006.2 tok/s 71.8 tok/s 2.20 GiB
prefill-512-decode-64 428.8 ms 1194.0 tok/s 62.2 tok/s 3.43 GiB

Both runs were executed on an Apple M2 Max with 32 GB unified memory in strict offline mode. All five generated-token checksums matched for each scenario.

Validation

  • current environment offline suite: 125 tests, 103 passed, 22 skipped
  • minimum compatibility environment offline suite: 125 tests, 103 passed, 22 skipped
  • ruff check .
  • ruff format --check .
  • wheel and source distribution build
  • twine check on both artifacts
  • isolated installed-wheel CLI and packaged-scenario smoke test
  • protocol validator and privacy scan on both locally generated reference JSON files

@ToluClassics
ToluClassics marked this pull request as ready for review July 29, 2026 01:11
@ToluClassics
ToluClassics requested a review from Copilot July 29, 2026 01:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR introduces a reproducible benchmarking protocol for MLX Transformers by adding a dedicated mlx-transformers-benchmark CLI, bundling deterministic protocol v1 scenarios, and documenting how to generate/validate/share benchmark results without leaking private machine details.

Changes:

  • Added mlx-transformers-benchmark CLI (run, validate, render) plus validation logic for scenario/result determinism and privacy constraints.
  • Added built-in protocol v1 scenarios and a packaged JSON schema under benchmark_data/.
  • Updated docs/config to ship benchmark assets, document usage, and harden the legacy example by disabling trust_remote_code.

Reviewed changes

Copilot reviewed 11 out of 11 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
tests/test_benchmark.py Adds protocol/CLI tests for offline run + validation + render output.
src/mlx_transformers/benchmark.py Implements CLI, scenario/result validation, deterministic run loop, and log rendering.
src/mlx_transformers/benchmark_data/short-decode-128.json Adds built-in protocol v1 scenario (64 prefill / 128 decode).
src/mlx_transformers/benchmark_data/prefill-512-decode-64.json Adds built-in protocol v1 scenario (512 prefill / 64 decode).
src/mlx_transformers/benchmark_data/result-v1.schema.json Adds packaged JSON schema for protocol v1 results.
README.md Documents reproducible benchmark CLI usage and updates test baseline counts.
pyproject.toml Registers the new CLI entry point and includes benchmark JSON package data.
MANIFEST.in Ensures benchmark markdown docs are included in sdist.
examples/text_generation/benchmark_generation.py Disables remote code execution in the legacy benchmark example.
CONTRIBUTING.md Updates expected offline test baseline counts.
benchmarks/REPRODUCIBILITY_LOG.md Adds initial validated reference results in the reproducibility log.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/mlx_transformers/benchmark.py
Comment thread src/mlx_transformers/benchmark.py
Comment thread src/mlx_transformers/benchmark.py Outdated
Comment on lines +111 to +117
scenario = result["scenario"]["definition"]
validate_scenario(scenario)
expected_hash = hashlib.sha256(
json.dumps(scenario, sort_keys=True, separators=(",", ":")).encode()
).hexdigest()
if result["scenario"].get("sha256") != expected_hash:
raise ValueError("Scenario checksum does not match its embedded definition.")
Comment thread src/mlx_transformers/benchmark.py Outdated
Comment on lines +192 to +197
bos_token_id = getattr(tokenizer, "bos_token_id", None)
if bos_token_id is not None:
token_ids = [bos_token_id, *token_ids]
target = scenario["prompt_tokens"]
repeated = (token_ids * ((target + len(token_ids) - 1) // len(token_ids)))[:target]
input_ids = mx.array([repeated], dtype=mx.int32)
Comment thread src/mlx_transformers/benchmark.py Outdated
Comment on lines +219 to +223
first_token = next(tokens)
mx.eval(first_token)
token_hash = hashlib.sha256()
token_hash.update(int(first_token.item()).to_bytes(8, "little", signed=True))
first_token_at = time.perf_counter()
Comment thread src/mlx_transformers/benchmark.py
@ToluClassics
ToluClassics merged commit 5883d57 into main Jul 29, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants