You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Tracking the multi-phase work needed to run FP16 transformer models (SmolVLM / ViT-family, and secondarily mobilesam_encoder / lprnet) end-to-end under openrknn's OWN path. This issue is the umbrella for ViT-specific work; it hangs off the main openrknn roadmap at #68.
All phases below are gated by a single prerequisite that is currently unresolved (see §0 below). Without unblocking that, none of the patch-rule work lower in the list will produce a runnable model — the oracle-patch scouting mode already proves that even byte-exact vendor regcmd bytes aren't enough to get SmolVLM shards to complete a submit.
Background
Context lives in openrknn/docs/fp16-transformer-support.md (landed in #78). Read that end-to-end first — it has the op topology, enable-mask histograms, vendor-vs-template diff patterns, the blob-to-op mappings observed during diagnosis, and a "do not re-litigate" list of dead ends already explored.
Important correction (also captured in the docs file): openrknn has never actually executed any FP16 transformer model in OWN mode. The mobilesam_encoder and lprnet entries in the CI suite are labeled "init + query only" because extract_npu_data() can't find regcmd/taskbo in their FB layout and silently falls through to the vendor proxy for run. SmolVLM shards are the first FP16 models that reach openrknn's OWN run path at all, which is why they're a better starting target than mobilesam.
Two narrow patch rules for transformer op lowerings (em=0x18 input-consuming REFORMAT, em=0x0d non-softmax CNA_FEATURE_DATA_ADDR) — both gated so CNN byte-exactness is preserved.
Dev scaffolding env vars (ORKNN_DUMP_BO1_PRE, ORKNN_DEBUG_BLOBS, ORKNN_DEBUG_PATCH, ORKNN_ORACLE_PATCH).
Still missing:
Core blocker (§0) — ORKNN_ORACLE_PATCH mode still hangs the NPU submit even when loading vendor's post-init BO byte-for-byte + rebasing every known DMA register. Something beyond regcmd content is also missing.
Patch rule coverage (§2/§3) — patch_regcmd_addresses can't resolve op→auxiliary-blob mapping for em=0x0d non-softmax ops (info lives in FB field 9 which openrknn doesn't parse).
FB extractor generalization (§5) — mobilesam_encoder / lprnet regcmd/taskbo live in an FB field layout extract_npu_data() doesn't recognize.
Only 12 ops, 1183 tasks, 3 FB segments → fits on a screen.
Op mix: Reshape / Mul / Transpose / exNorm / ConvExSwish / Conv / Add / Reshape → every transformer op class we need except attention itself.
MLP shards are layout-identical fused vs unfused (attention is the only part that changes with fusion rules), so this target stays stable across rknn-toolkit2 configurations.
Phase 0 — Resolve the ORKNN_ORACLE_PATCH hang (PREREQUISITE)
The open question at the top of everything.ORKNN_ORACLE_PATCH loads vendor's post-init weight BO verbatim into openrknn's weight BO and rebases every known DMA-bearing register (0x0010, 0x1070, 0x1110, 0x4020, 0x4048, 0x5018, 0x5020, 0x502c, 0x5038, 0x504c, 0x6070, 0x701c) from vendor's bases to openrknn's bases. The rebased BO is byte-exact with what vendor's runtime produces before submit.
It still hangs segment 0. That tells us at least one of the following is wrong independent of regcmd content:
Activation BO sizing. Vendor allocates ~3× openrknn's activation BO (28 MB vs 9.3 MB for l0_mlp — probably one slice per core even in single-core mode). Max observed activation offset in the regcmd is 7.9 MB which fits in both, so it shouldn't matter… but it might. Test: bump openrknn's activation BO to 3× and retry oracle-patch.
Single-task isolation. Add an ORKNN_MAX_TASKS=N env var that truncates the first segment to N tasks. Find the smallest N that still hangs — that narrows the failure to a single task whose content we can inspect directly.
struct rknpu_submit byte comparison. Hex-dump openrknn's ioctl args against vendor's (via extended intercept_swap.c already logging task_base_addr, iommu_domain_id, subcore_task[0..2]). Anything differing beyond task_obj_addr is a bug.
Per-segment MEM_SYNC(TO_DEVICE). Currently openrknn syncs the weight BO once after patching. Try adding a pre-submit sync of task + weight + activation BOs for every segment and retry.
IOMMU attach state. Our out-of-tree kernel patch (patches/kernel/0001-rocket-iommu-attach-caching-and-submit-error-propagation.patch) caches IOMMU attaches per rocket core. Confirm whether the vendor rknpu kernel module has similar IOMMU handling for transformer shards specifically — a re-attach stall could look like a hang.
Exit criterion.ORKNN_ORACLE_PATCH=... bench_throughput --model l0_mlp.rknn --strategy pinned completes segment 0 submit and returns an output tensor (content correctness is the next phase's problem; we just need the submit to not time out). Until this passes, phases 1+ are not meaningful — they'd all be blind changes.
Estimated cost: 1–3 days of iterative bisection. This is the highest-risk phase because we don't yet know what's wrong.
Phase 1 — Byte-exact l0_mlp.rknn (FB field 9 schema decode)
With phase 0 done, we know the patch rules are the only remaining gap. The concrete blocker is that em=0x0d non-softmax ops (Transpose, exNorm, general compute) have a per-op auxiliary LUT blob reference that's not in input_tensors[] — it's encoded in the FB operator record's field 9 (a 70-u32 vector of undocumented metadata per op).
Example: for l0_mlp.rknn, op 3 (Transpose) needs CNA_DCOMP_ADDR0 = wt_base + 0x910180 which is blob[17] (size 8192, type 6). The op's only declared input is tensor 15 (the activation), so the existing tensor_weight_blob[input_tensors[1]] lookup returns nothing.
Plan:
Dump field 9 for every op in l0_mlp, l0_attn (fused), and a few known-working CNN models side-by-side. Look for u32 values that map to known blob offsets. The most likely shape is a {blob_tidx, stride, count, ...} struct, possibly packed.
Write a parse_op_aux_blob() helper in openrknn_model.c that extracts the auxiliary blob reference per op and stores it in struct orknn_op_info (add aux_blob_tidx / aux_blob_off fields).
Extend the case 0x1110: branch in patch_regcmd_addresses to resolve the aux blob for em=0x0d non-softmax tasks, with op-type dispatch (Transpose / exNorm / Mul).
Extend the case 0x1070: branch similarly for the subset of ops where the FEATURE_DATA_ADDR walks a scratch tensor rather than the primary input.
Validate via diff_regcmd.py: 0 non-DMA diffs, 0 DMA-class diffs on l0_mlp over all 643 unique regcmd sections. This is the objective gate — not "my changes look reasonable".
Exit criterion.ORKNN_OWN=... bench_throughput --model l0_mlp.rknn --strategy pinned runs without ORKNN_ORACLE_PATCH set, completes 3 segments per inference, and produces an output tensor whose cosine similarity against a vendor-run reference is > 0.999.
Estimated cost: 3–5 days (dominated by FB schema RE — unknown until somebody tries). exNorm is the hardest op because it references 4+ blobs per op in different slots.
With l0_mlp running, attention shards are the next piece. Two flavors depending on how upstream is compiled:
Fused (smolvlm_subshards_fused/): uses exSDPAttention × 32 (one per NanoTiled Q-chunk), plus Conv × 4, Mul × 2, exNorm × 1. Much simpler lowering — exSDPAttention is a single fused op that mobilesam_encoder also uses, so its patch rules likely generalize across both models.
Unfused (smolvlm_subshards/): uses exMatMul × 64 + Softmax × 32 + Mul × 34. Way more tasks (19k for unfused l0_attn vs ~6k for fused) but each op is simpler. Useful as a stress test for the patching hot path.
Start with fused. exSDPAttention is the only genuinely-new op type here and also appears in mobilesam_encoder — solving it unlocks both models.
Run the oracle/template diff on smolvlm_subshards_fused/l0_attn.rknn.
Add whatever patch rules are needed for exSDPAttention, Mul, and any tile-routing quirks.
Verify byte-exact.
Verify end-to-end cosine > 0.999.
Exit criterion. Both l0_attn.rknn (fused and unfused) run end-to-end under OWN mode with cosine > 0.999 vs vendor.
Estimated cost: 2–3 days on top of Phase 1.
Phase 3 — Full 24-shard SmolVLM pipeline
All 24 shards (l0..l11 × {attn, mlp}) running under openrknn, orchestrated by the upstream project's HybridSplitEncoder Python glue.
Run scripts/validate.py from the upstream repo against openrknn: layer-by-layer cosine similarity for all 24 shards.
Run scripts/run_inference.py: full vision encoder + image embed export, compared against CPU reference.
Measure end-to-end vision encoder latency under OWN mode. Reference: vendor stack is 1970 ms on the same shards (from librocketnpu: SmolVLM ViT sharding vendor-stack reproduction #79's README). We expect similar — openrknn is just a runtime; it doesn't change NPU compute time.
Exit criterion: end-to-end vision encoder cosine > 0.99 vs CPU reference, all 24 shards byte-exact against vendor oracle.
Estimated cost: 1–2 days (mostly just running tests and fixing per-layer edge cases).
Phase 4 — Add SmolVLM shards to CI
Add smolvlm_l0_mlp.rknn and smolvlm_l0_attn.rknn (fused variants) to openrknn/tests/ground_truth.json as fp16_run targets, with expected cosine thresholds against a captured vendor reference.
Check in the captured Phase 2.5 segmentation ground-truth artifact (segmentation_ground_truth/smolvlm_l0_*.json).
Ensure ci_validate.sh Phase 2 + Phase 2.5 both cover the new models without doubling CI runtime.
Exit criterion. Phase 2.5 shows 7 clean, 0 gating, 0 allowlisted (5 existing CNN + 2 new SmolVLM shards). Any future openrknn patch that regresses transformer support is caught immediately.
Estimated cost: 1 day.
Phase 5 — Lift FB extractor for mobilesam_encoder / lprnet
Separate from the SmolVLM work but cheap to do once the transformer patch rules exist. Both models currently bail out of extract_npu_data() with regcmd(0) or taskbo(0) not found and silently fall through to the vendor proxy. Their regcmd/taskbo blobs live in a different FB field layout that openrknn's scanner doesn't recognize — probably just a missing blob-type heuristic.
Identify the correct FB field for regcmd/taskbo in these models (they're newer rknn-toolkit2 outputs, likely use weight_data_v6 at vt offset 0x2C with different blob type bytes).
Extend extract_npu_data() to find them.
Run l0_mlp-style oracle diffs against both models and fill in whatever patch rules are still missing.
Promote both from fp16_parse to fp16_run in the CI suite.
Exit criterion.mobilesam_encoder and lprnet run end-to-end under OWN mode in Phase 2; Phase 2.5 shows byte-exact diffs.
Estimated cost: 2–3 days — smaller scope than SmolVLM once the op patch rules are done.
Phase 6 — Upstream the intercept_swap.c enhancements
Nice-to-have. During this session I extended librocketnpu/tests/intercept_swap.c to log task_base_addr, iommu_domain_id, and subcore_task[3..4] in the SWAP: SUBMIT line. That's useful for future submit-parameter diagnosis and should be upstreamed as its own small commit (not tied to this roadmap's correctness work).
Clean up the on-board hand-edited version vs the repo's canonical librocketnpu/tests/intercept_swap.c.
Small PR with the logging diff + a note in the file header about what new fields are captured and why.
Non-goals
Things explicitly out of scope for this roadmap:
Mesa rocket / openrknn cross-pollination. This issue is strictly about openrknn's OWN path consuming vendor-compiled .rknn files. Mesa rocket produces its own regcmd from ONNX — that's a completely different engineering surface tracked separately.
Writing our own FP16 lowering for any op. openrknn is a runtime, not a compiler. If the vendor's compiler can't produce an op, openrknn doesn't either.
Optimizing vs vendor perf. Matching vendor latency for transformer shards is the goal, not beating it. openrknn doesn't compile the regcmd, so any perf delta vs vendor is noise from submit/ioctl plumbing.
Relationship to other roadmaps
Parallel to but independent of #68 (main openrknn roadmap). Most feature issues there (zero-copy I/O, dynamic shape, multi-core batch, LSTM) are orthogonal to transformer support and can land in either order. The exceptions:
Summary
Tracking the multi-phase work needed to run FP16 transformer models (SmolVLM / ViT-family, and secondarily mobilesam_encoder / lprnet) end-to-end under openrknn's OWN path. This issue is the umbrella for ViT-specific work; it hangs off the main openrknn roadmap at #68.
All phases below are gated by a single prerequisite that is currently unresolved (see §0 below). Without unblocking that, none of the patch-rule work lower in the list will produce a runnable model — the oracle-patch scouting mode already proves that even byte-exact vendor regcmd bytes aren't enough to get SmolVLM shards to complete a submit.
Background
Context lives in
openrknn/docs/fp16-transformer-support.md(landed in #78). Read that end-to-end first — it has the op topology, enable-mask histograms, vendor-vs-template diff patterns, the blob-to-op mappings observed during diagnosis, and a "do not re-litigate" list of dead ends already explored.Important correction (also captured in the docs file): openrknn has never actually executed any FP16 transformer model in OWN mode. The mobilesam_encoder and lprnet entries in the CI suite are labeled "init + query only" because
extract_npu_data()can't find regcmd/taskbo in their FB layout and silently falls through to the vendor proxy for run. SmolVLM shards are the first FP16 models that reach openrknn's OWN run path at all, which is why they're a better starting target than mobilesam.Current state (master @ f468ea2, 2026-04-12)
Landed in #78:
MAX_PATCHED_OFFSETShardcap 4096 → dynamicm->task_count).CNA_FEATURE_DATA_ADDR) — both gated so CNN byte-exactness is preserved.ORKNN_DUMP_BO1_PRE,ORKNN_DEBUG_BLOBS,ORKNN_DEBUG_PATCH,ORKNN_ORACLE_PATCH).Still missing:
ORKNN_ORACLE_PATCHmode still hangs the NPU submit even when loading vendor's post-init BO byte-for-byte + rebasing every known DMA register. Something beyond regcmd content is also missing.patch_regcmd_addressescan't resolve op→auxiliary-blob mapping forem=0x0dnon-softmax ops (info lives in FB field 9 which openrknn doesn't parse).extract_npu_data()doesn't recognize.Bisection target
external/smolvlm_rk3588_full_npu_native/smolvlm_subshards/l0_mlp.rknn(build via #79's Dockerfile + poad42/smolvlm_rk3588_full_npu_native).Why l0_mlp and not l0_attn:
openrknn/docs/fp16-transformer-support.md §8.Oracle capture (one-time per model):
Phase 0 — Resolve the
ORKNN_ORACLE_PATCHhang (PREREQUISITE)The open question at the top of everything.
ORKNN_ORACLE_PATCHloads vendor's post-init weight BO verbatim into openrknn's weight BO and rebases every known DMA-bearing register (0x0010, 0x1070, 0x1110, 0x4020, 0x4048, 0x5018, 0x5020, 0x502c, 0x5038, 0x504c, 0x6070, 0x701c) from vendor's bases to openrknn's bases. The rebased BO is byte-exact with what vendor's runtime produces before submit.It still hangs segment 0. That tells us at least one of the following is wrong independent of regcmd content:
ORKNN_MAX_TASKS=Nenv var that truncates the first segment to N tasks. Find the smallest N that still hangs — that narrows the failure to a single task whose content we can inspect directly.struct rknpu_submitbyte comparison. Hex-dump openrknn's ioctl args against vendor's (via extendedintercept_swap.calready loggingtask_base_addr,iommu_domain_id,subcore_task[0..2]). Anything differing beyondtask_obj_addris a bug.MEM_SYNC(TO_DEVICE). Currently openrknn syncs the weight BO once after patching. Try adding a pre-submit sync of task + weight + activation BOs for every segment and retry.patches/kernel/0001-rocket-iommu-attach-caching-and-submit-error-propagation.patch) caches IOMMU attaches per rocket core. Confirm whether the vendorrknpukernel module has similar IOMMU handling for transformer shards specifically — a re-attach stall could look like a hang.Exit criterion.
ORKNN_ORACLE_PATCH=... bench_throughput --model l0_mlp.rknn --strategy pinnedcompletes segment 0 submit and returns an output tensor (content correctness is the next phase's problem; we just need the submit to not time out). Until this passes, phases 1+ are not meaningful — they'd all be blind changes.Estimated cost: 1–3 days of iterative bisection. This is the highest-risk phase because we don't yet know what's wrong.
Phase 1 — Byte-exact
l0_mlp.rknn(FB field 9 schema decode)With phase 0 done, we know the patch rules are the only remaining gap. The concrete blocker is that
em=0x0dnon-softmax ops (Transpose, exNorm, general compute) have a per-op auxiliary LUT blob reference that's not ininput_tensors[]— it's encoded in the FB operator record's field 9 (a 70-u32 vector of undocumented metadata per op).Example: for
l0_mlp.rknn, op 3 (Transpose) needsCNA_DCOMP_ADDR0 = wt_base + 0x910180which is blob[17] (size 8192, type 6). The op's only declared input is tensor 15 (the activation), so the existingtensor_weight_blob[input_tensors[1]]lookup returns nothing.Plan:
l0_mlp,l0_attn(fused), and a few known-working CNN models side-by-side. Look for u32 values that map to known blob offsets. The most likely shape is a{blob_tidx, stride, count, ...}struct, possibly packed.parse_op_aux_blob()helper inopenrknn_model.cthat extracts the auxiliary blob reference per op and stores it instruct orknn_op_info(addaux_blob_tidx/aux_blob_offfields).case 0x1110:branch inpatch_regcmd_addressesto resolve the aux blob forem=0x0dnon-softmax tasks, with op-type dispatch (Transpose / exNorm / Mul).case 0x1070:branch similarly for the subset of ops where the FEATURE_DATA_ADDR walks a scratch tensor rather than the primary input.diff_regcmd.py:0 non-DMA diffs, 0 DMA-class diffsonl0_mlpover all 643 unique regcmd sections. This is the objective gate — not "my changes look reasonable".Exit criterion.
ORKNN_OWN=... bench_throughput --model l0_mlp.rknn --strategy pinnedruns withoutORKNN_ORACLE_PATCHset, completes 3 segments per inference, and produces an output tensor whose cosine similarity against a vendor-run reference is > 0.999.Estimated cost: 3–5 days (dominated by FB schema RE — unknown until somebody tries). exNorm is the hardest op because it references 4+ blobs per op in different slots.
Phase 2 — l0_attn.rknn (exMatMul / exSDPAttention)
With
l0_mlprunning, attention shards are the next piece. Two flavors depending on how upstream is compiled:smolvlm_subshards_fused/): usesexSDPAttention × 32(one per NanoTiled Q-chunk), plusConv × 4,Mul × 2,exNorm × 1. Much simpler lowering —exSDPAttentionis a single fused op that mobilesam_encoder also uses, so its patch rules likely generalize across both models.smolvlm_subshards/): usesexMatMul × 64+Softmax × 32+Mul × 34. Way more tasks (19k for unfused l0_attn vs ~6k for fused) but each op is simpler. Useful as a stress test for the patching hot path.Start with fused.
exSDPAttentionis the only genuinely-new op type here and also appears in mobilesam_encoder — solving it unlocks both models.smolvlm_subshards_fused/l0_attn.rknn.exSDPAttention,Mul, and any tile-routing quirks.Exit criterion. Both
l0_attn.rknn(fused and unfused) run end-to-end under OWN mode with cosine > 0.999 vs vendor.Estimated cost: 2–3 days on top of Phase 1.
Phase 3 — Full 24-shard SmolVLM pipeline
All 24 shards (l0..l11 × {attn, mlp}) running under openrknn, orchestrated by the upstream project's
HybridSplitEncoderPython glue.scripts/validate.pyfrom the upstream repo against openrknn: layer-by-layer cosine similarity for all 24 shards.scripts/run_inference.py: full vision encoder + image embed export, compared against CPU reference.Estimated cost: 1–2 days (mostly just running tests and fixing per-layer edge cases).
Phase 4 — Add SmolVLM shards to CI
smolvlm_l0_mlp.rknnandsmolvlm_l0_attn.rknn(fused variants) toopenrknn/tests/ground_truth.jsonasfp16_runtargets, with expected cosine thresholds against a captured vendor reference.segmentation_ground_truth/smolvlm_l0_*.json)..rknnfiles themselves (git-lfs? separate release? rebuildable via the Dockerfile from librocketnpu: SmolVLM ViT sharding vendor-stack reproduction #79?).ci_validate.shPhase 2 + Phase 2.5 both cover the new models without doubling CI runtime.Exit criterion. Phase 2.5 shows
7 clean, 0 gating, 0 allowlisted(5 existing CNN + 2 new SmolVLM shards). Any future openrknn patch that regresses transformer support is caught immediately.Estimated cost: 1 day.
Phase 5 — Lift FB extractor for mobilesam_encoder / lprnet
Separate from the SmolVLM work but cheap to do once the transformer patch rules exist. Both models currently bail out of
extract_npu_data()withregcmd(0) or taskbo(0) not foundand silently fall through to the vendor proxy. Their regcmd/taskbo blobs live in a different FB field layout that openrknn's scanner doesn't recognize — probably just a missing blob-type heuristic.weight_data_v6at vt offset 0x2C with different blob type bytes).extract_npu_data()to find them.l0_mlp-style oracle diffs against both models and fill in whatever patch rules are still missing.fp16_parsetofp16_runin the CI suite.Exit criterion.
mobilesam_encoderandlprnetrun end-to-end under OWN mode in Phase 2; Phase 2.5 shows byte-exact diffs.Estimated cost: 2–3 days — smaller scope than SmolVLM once the op patch rules are done.
Phase 6 — Upstream the
intercept_swap.cenhancementsNice-to-have. During this session I extended
librocketnpu/tests/intercept_swap.cto logtask_base_addr,iommu_domain_id, andsubcore_task[3..4]in the SWAP: SUBMIT line. That's useful for future submit-parameter diagnosis and should be upstreamed as its own small commit (not tied to this roadmap's correctness work).librocketnpu/tests/intercept_swap.c.Non-goals
Things explicitly out of scope for this roadmap:
.rknnfiles. Mesa rocket produces its own regcmd from ONNX — that's a completely different engineering surface tracked separately..rkllmformat, which is a completely different container outside openrknn's scope.Relationship to other roadmaps
Parallel to but independent of #68 (main openrknn roadmap). Most feature issues there (zero-copy I/O, dynamic shape, multi-core batch, LSTM) are orthogonal to transformer support and can land in either order. The exceptions:
Tooling references
All of these already exist and are ready to use:
LD_PRELOAD=librocketnpu/tests/intercept_swap.so DUMP_ALL_BOS=1 bench_throughput ...ORKNN_DUMP_BO1=/path,ORKNN_DUMP_BO1_PRE=/path,ORKNN_DUMP_TASKBO=/pathopenrknn/tests/diff_regcmd.py --oracle --template --task-bo --submit-txt --show-dma --em-filter 0x0dORKNN_ORACLE_PATCH=/path ORKNN_ORACLE_{WT,ACT,IN,OUT}_BASE=0x...ORKNN_DEBUG_BLOBS=1,ORKNN_DEBUG_PATCH=1struct.unpack_from('<8I Q', task_bo, i*40)for 40-byterknpu_taskentries (fields:flags, op_idx, enable_mask, int_mask, int_clear, int_status, regcfg_amount, regcfg_offset, regcmd_addr)RE references (existing):
~/projects/ida/librknnrt/docs/model-parsing.md— FB schema (field 0..21)~/projects/ida/librknnrt/docs/graph-executor.md— op routing, per-nodetargetbyte~/projects/ida/librknnrt/docs/rknn-model-format.md— weight data section semantics, version-gated field layouts~/projects/ida/librknnrt/librknnrt.so.c— full Hex-Rays decompile, look atsub_2DC018(v>5 operator extractor) for per-op field decoding hints🤖 Generated with Claude Code