Repository navigation
feat(demo): stream_client.py warmup mode - #179
Merged
Merged
Conversation
…bers The first request after a model loads pays one-off Metal pipeline and compile costs. On a Mac mini M6, Qwen3.6-35B-A3B decoded at 39.7 tok/s cold vs 44.5 tok/s warm on the same prompt. `stream_client.py warmup` sends a short "count from 1 to 8" request so recordings can show the warm-up and the measured request as two visible, separate steps rather than quoting a cold number or hiding the warm-up. The prompt-token/prefill stat now only prints for `long`, where it's meaningful. Used for the Qwen3.6-35B recording linked from the M6 launch post. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Member
Author
|
Reviewed (code-review skill, M5 agent): no issues found. The warmup mode is fine, and |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Small demo-tooling PR: adds a
warmupmode toscripts/demo/stream_client.py.Why
The first request after a model loads pays one-off Metal pipeline and compile costs. On a base Mac mini M6 (32 GB), Qwen3.6-35B-A3B decoded the same prompt at 39.7 tok/s cold vs 44.5 tok/s warm. For terminal recordings we want neither a misleading cold number nor a warm-up done off screen. With this mode the recording shows both requests as separate, labelled steps:
This is the sequence in the Qwen3.6-35B GIF attached to the M6 launch post (https://x.com/Simba_Zhang/status/2103293610670297167), so anyone can reproduce that exact recording.
Changes
warmup: a short "Count from 1 to 8" request,max_tokens24.prompt … tok · prefill … tok/sstat now prints only forlong, where the prompt is big enough for it to mean something. (Forwarmupit printedprefill 8 tok/son an 18-token prompt.)No server changes.
🤖 Generated with Claude Code