Repository navigation
chore: bump mlx-swift (mlx core v0.32.3) and mlx-swift-lm (#77) - #215
Merged
Merged
Conversation
…to a8c85d5 mlx-swift 34bc52f -> 3764829 (SharpAI/mlx-swift#22): vendored mlx core v0.32.3 with the gather_mm_offsets kernel bundled into the SwiftPM metallib. mlx-swift-lm 6490e3f -> a8c85d5 (SharpAI/mlx-swift-lm#77): skip the per-layer MoE GPU flush on single-token decode with --stream-experts (MLX_MOE_DECODE_FLUSH=1 restores it). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bump both submodules to the current fork
main. No SwiftLM source change.mlx-swift34bc52f→3764829(Vendor mlx core v0.32.3 (with a floor_divide fix for unsigned integers) mlx-swift#22, by @CodeAndCanvas728 plus our follow-up fixes): vendors mlx core v0.32.3. Kernel files new in this release are bundled into the SwiftPM metallib (gather_mm_offsets) and excluded from the xcodeproj Cmlx target. The fork now has a CI job that runs MLXTests against the SwiftPM bundle metallib and buildsxcode/MLX.xcodeproj.mlx-swift-lm6490e3f→a8c85d5(perf(qwen35): skip the per-layer MoE GPU flush on single-token decode mlx-swift-lm#77, by @CodeAndCanvas728 plus a follow-up commit): Qwen3.5 MoE with--stream-expertsno longer syncs the GPU twice per layer on single-token decode (prefill and verify steps keep it). The author measured about +23% decode speed on a 16 GB M2 with the 35B-A3B model.MLX_MOE_DECODE_FLUSH=1restores the old behaviour.Sources/,Tests/andPackage.swiftbuild unchanged against both.Test plan
Local (M5), against mlx-swift
3764829+ mlx-swift-lma8c85d5:swift build -c releaseOK,swift build --build-testsOKmlx.metallibbuilt withbuild.sh's cmake steps (-DMLX_ENABLE_NAX=1 -DMLX_METAL_JIT=OFF), and it containsgather_mm_offsetsSwiftLMTests: 276 tests, 0 failuresSwiftBuddyTests: 127 tests, 9 skipped (environment-gated), 0 failuresQwen3.5-0.8B-4bitloads and answers a chat completion--stream-expertswith a real Qwen3.5 MoE (no such checkpoint here; the fix: memory auto-cap strategy for SSD MoE streaming + speculative decoding (Issue #72) #77 review used a small synthetic one),MLX_MOE_STACKED=1,--mtpwith a real model, the model-download integration tests,tests/test-*.sh,xcodebuild -scheme SwiftBuddyKnown and not verified: the opt-in
MLX_MOE_STACKED=1path callsgatherQuantizedMM(..., sortedIndices: true)with LRU slot ids that are not necessarily ascending (SwitchLayers.swift). The sorted kernels rely on ascending indices. I did not check whether 0.32.3 changes the result for that path.CI:
build_and_unit_testspeculative-decoding,dflash-speculative-decoding,speculative-decoding-evalintegration_matrixjobsAI disclosure
This PR was prepared with AI assistance (Claude Code: build/test runs and this description). It still needs human review before merging.
🤖 Generated with Claude Code