Skip to content

openrknn: document multi-core dispatch blocker for FP16 - #91

Merged
widgetii merged 1 commit into
masterfrom
openrknn/multicore-todo
Apr 13, 2026
Merged

widgetii merged 1 commit into
masterfrom
openrknn/multicore-todo

Conversation

@widgetii

Copy link
Copy Markdown
Owner

Summary

Documents the final blocker for correct FP16 transformer output: the vendor's kernel-level multi-core dispatch.

Single-core tasks cover ~512 of 1024 spatial positions. The computation at covered positions is verified correct (768/768 channel bias match, byte-exact DMA addresses). The remaining 50% of positions are zero because multi-core dispatch isn't implemented in OWN mode.

No functional changes — just a TODO comment.

Test plan

  • CI: 9/9 pass, 6/7 byte-exact (l0_mlp BYTE-EXACT)

🤖 Generated with Claude Code

Single-core NPU tasks only cover ~512 of 1024 spatial positions for
SmolVLM transformer shards. The vendor runs all 3 NPU cores via
kernel-level multi-core dispatch configured during rknn_init.

openrknn's OWN mode skips vendor init, so core_mask=0 defaults to
single core. Add TODO comment tracking this as the final blocker
for correct FP16 output.

The DMA address patching is now byte-exact (0 diffs) and the per-
channel computation is verified correct (768/768 channel bias match).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@widgetii
widgetii merged commit bf59446 into master Apr 13, 2026
7 checks passed
@widgetii
widgetii deleted the openrknn/multicore-todo branch April 13, 2026 08:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant