Skip to content

chore: bump mlx-swift (mlx core v0.32.3) and mlx-swift-lm (#77) - #215

Merged
solderzzc merged 1 commit into
mainfrom
chore/bump-mlx-core-0.32.3-and-lm-77
Oct 7, 2026
Merged

solderzzc merged 1 commit into
mainfrom
chore/bump-mlx-core-0.32.3-and-lm-77

Conversation

@solderzzc

Copy link
Copy Markdown
Member

Summary

Bump both submodules to the current fork main. No SwiftLM source change.

Sources/, Tests/ and Package.swift build unchanged against both.

Test plan

Local (M5), against mlx-swift 3764829 + mlx-swift-lm a8c85d5:

  • swift build -c release OK, swift build --build-tests OK
  • mlx.metallib built with build.sh's cmake steps (-DMLX_ENABLE_NAX=1 -DMLX_METAL_JIT=OFF), and it contains gather_mm_offsets
  • SwiftLMTests: 276 tests, 0 failures
  • SwiftBuddyTests: 127 tests, 9 skipped (environment-gated), 0 failures
  • Smoke test: Qwen3.5-0.8B-4bit loads and answers a chat completion
  • Not run locally: --stream-experts with a real Qwen3.5 MoE (no such checkpoint here; the fix: memory auto-cap strategy for SSD MoE streaming + speculative decoding (Issue #72) #77 review used a small synthetic one), MLX_MOE_STACKED=1, --mtp with a real model, the model-download integration tests, tests/test-*.sh, xcodebuild -scheme SwiftBuddy

Known and not verified: the opt-in MLX_MOE_STACKED=1 path calls gatherQuantizedMM(..., sortedIndices: true) with LRU slot ids that are not necessarily ascending (SwitchLayers.swift). The sorted kernels rely on ascending indices. I did not check whether 0.32.3 changes the result for that path.

CI:

  • build_and_unit_test
  • speculative-decoding, dflash-speculative-decoding, speculative-decoding-eval
  • Remaining integration_matrix jobs

AI disclosure

This PR was prepared with AI assistance (Claude Code: build/test runs and this description). It still needs human review before merging.

  • I have read this PR description in full and approve it as my own

🤖 Generated with Claude Code

…to a8c85d5

mlx-swift 34bc52f -> 3764829 (SharpAI/mlx-swift#22): vendored mlx core v0.32.3 with
the gather_mm_offsets kernel bundled into the SwiftPM metallib.
mlx-swift-lm 6490e3f -> a8c85d5 (SharpAI/mlx-swift-lm#77): skip the per-layer MoE
GPU flush on single-token decode with --stream-experts (MLX_MOE_DECODE_FLUSH=1
restores it).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@solderzzc
solderzzc merged commit 114fe30 into main Oct 7, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant