Skip to content

Warm 512-wide DeepSeek sparse-MLA kernels before serving - #11

Merged
Aaronontheweb merged 1 commit into
masterfrom
fix/warm-512-sparse-mla
Aug 4, 2026
Merged

Aaronontheweb merged 1 commit into
masterfrom
fix/warm-512-sparse-mla

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Summary

  • warm the observed 512-wide portable sparse-MLA Triton layout during startup
  • retain the full 576-wide NoPE plus RoPE warmup
  • cover both 32- and 64-head TP-local layouts and update locked source provenance

Evidence

A live 16K prefill compiled _accumulate_indexed_attention_chunk_multihead_kernel after startup with:

BLOCK_D=512
head_dim=512
num_heads=32
stride_q_t=16384
stride_q_h=512

The prior image only warmed 576-wide Q/K tensors. Triton treats head_dim and the contiguous strides as compile keys.

Validation

  • ./scripts/validate-repository.sh
  • Python syntax compilation for the changed warmup and test modules
  • cumulative patch applies cleanly to pinned vLLM commit 568afb3a13806beb53bb2e6bd518269357b237c0
  • pinned upstream source archive checksum matches dependency.lock.json

@Aaronontheweb
Aaronontheweb merged commit 2d22ab2 into master Aug 4, 2026
1 check passed
@Aaronontheweb
Aaronontheweb deleted the fix/warm-512-sparse-mla branch August 4, 2026 21:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant