Skip to content

Roadmap: openrknn FP16 transformer — DONE: bit-exact, 12% faster than vendor #80

Description

@widgetii

Summary

Tracking the multi-phase work needed to run FP16 transformer models (SmolVLM / ViT-family, and secondarily mobilesam_encoder / lprnet) end-to-end under openrknn's OWN path. This issue is the umbrella for ViT-specific work; it hangs off the main openrknn roadmap at #68.

All phases below are gated by a single prerequisite that is currently unresolved (see §0 below). Without unblocking that, none of the patch-rule work lower in the list will produce a runnable model — the oracle-patch scouting mode already proves that even byte-exact vendor regcmd bytes aren't enough to get SmolVLM shards to complete a submit.

Background

Context lives in openrknn/docs/fp16-transformer-support.md (landed in #78). Read that end-to-end first — it has the op topology, enable-mask histograms, vendor-vs-template diff patterns, the blob-to-op mappings observed during diagnosis, and a "do not re-litigate" list of dead ends already explored.

Important correction (also captured in the docs file): openrknn has never actually executed any FP16 transformer model in OWN mode. The mobilesam_encoder and lprnet entries in the CI suite are labeled "init + query only" because extract_npu_data() can't find regcmd/taskbo in their FB layout and silently falls through to the vendor proxy for run. SmolVLM shards are the first FP16 models that reach openrknn's OWN run path at all, which is why they're a better starting target than mobilesam.

Current state (master @ f468ea2, 2026-04-12)

Landed in #78:

  • Dedup overflow bugfix (MAX_PATCHED_OFFSETS hardcap 4096 → dynamic m->task_count).
  • Two narrow patch rules for transformer op lowerings (em=0x18 input-consuming REFORMAT, em=0x0d non-softmax CNA_FEATURE_DATA_ADDR) — both gated so CNN byte-exactness is preserved.
  • Dev scaffolding env vars (ORKNN_DUMP_BO1_PRE, ORKNN_DEBUG_BLOBS, ORKNN_DEBUG_PATCH, ORKNN_ORACLE_PATCH).

Still missing:

  • Core blocker (§0) — ORKNN_ORACLE_PATCH mode still hangs the NPU submit even when loading vendor's post-init BO byte-for-byte + rebasing every known DMA register. Something beyond regcmd content is also missing.
  • Patch rule coverage (§2/§3) — patch_regcmd_addresses can't resolve op→auxiliary-blob mapping for em=0x0d non-softmax ops (info lives in FB field 9 which openrknn doesn't parse).
  • FB extractor generalization (§5) — mobilesam_encoder / lprnet regcmd/taskbo live in an FB field layout extract_npu_data() doesn't recognize.

Bisection target

external/smolvlm_rk3588_full_npu_native/smolvlm_subshards/l0_mlp.rknn (build via #79's Dockerfile + poad42/smolvlm_rk3588_full_npu_native).

Why l0_mlp and not l0_attn:

  • Only 12 ops, 1183 tasks, 3 FB segments → fits on a screen.
  • Op mix: Reshape / Mul / Transpose / exNorm / ConvExSwish / Conv / Add / Reshape → every transformer op class we need except attention itself.
  • MLP shards are layout-identical fused vs unfused (attention is the only part that changes with fusion rules), so this target stays stable across rknn-toolkit2 configurations.
  • Reference data (blob layout, op→LUT mapping, enable_mask histogram, first-cycle task breakdown) is all in openrknn/docs/fp16-transformer-support.md §8.

Oracle capture (one-time per model):

# On the board, with vendor kernel 6.1.115-vendor-rk35xx booted
rm -rf /tmp/rknn_dump
LD_PRELOAD=/root/npu-research/librocketnpu/tests/intercept_swap.so \
DUMP_ALL_BOS=1 \
  /root/npu-research/openrknn/tests/bench_throughput \
    --model /root/npu-research/smolvlm_shard/smolvlm_subshards/l0_mlp.rknn \
    --workers 1 --duration 1 --warmup 0 --strategy pinned --label vendor
# -> /tmp/rknn_dump/sub1_bo_{000..004}.bin + submit_1.txt

Phase 0 — Resolve the ORKNN_ORACLE_PATCH hang (PREREQUISITE)

The open question at the top of everything. ORKNN_ORACLE_PATCH loads vendor's post-init weight BO verbatim into openrknn's weight BO and rebases every known DMA-bearing register (0x0010, 0x1070, 0x1110, 0x4020, 0x4048, 0x5018, 0x5020, 0x502c, 0x5038, 0x504c, 0x6070, 0x701c) from vendor's bases to openrknn's bases. The rebased BO is byte-exact with what vendor's runtime produces before submit.

It still hangs segment 0. That tells us at least one of the following is wrong independent of regcmd content:

  • Activation BO sizing. Vendor allocates ~3× openrknn's activation BO (28 MB vs 9.3 MB for l0_mlp — probably one slice per core even in single-core mode). Max observed activation offset in the regcmd is 7.9 MB which fits in both, so it shouldn't matter… but it might. Test: bump openrknn's activation BO to 3× and retry oracle-patch.
  • Single-task isolation. Add an ORKNN_MAX_TASKS=N env var that truncates the first segment to N tasks. Find the smallest N that still hangs — that narrows the failure to a single task whose content we can inspect directly.
  • struct rknpu_submit byte comparison. Hex-dump openrknn's ioctl args against vendor's (via extended intercept_swap.c already logging task_base_addr, iommu_domain_id, subcore_task[0..2]). Anything differing beyond task_obj_addr is a bug.
  • Per-segment MEM_SYNC(TO_DEVICE). Currently openrknn syncs the weight BO once after patching. Try adding a pre-submit sync of task + weight + activation BOs for every segment and retry.
  • IOMMU attach state. Our out-of-tree kernel patch (patches/kernel/0001-rocket-iommu-attach-caching-and-submit-error-propagation.patch) caches IOMMU attaches per rocket core. Confirm whether the vendor rknpu kernel module has similar IOMMU handling for transformer shards specifically — a re-attach stall could look like a hang.

Exit criterion. ORKNN_ORACLE_PATCH=... bench_throughput --model l0_mlp.rknn --strategy pinned completes segment 0 submit and returns an output tensor (content correctness is the next phase's problem; we just need the submit to not time out). Until this passes, phases 1+ are not meaningful — they'd all be blind changes.

Estimated cost: 1–3 days of iterative bisection. This is the highest-risk phase because we don't yet know what's wrong.


Phase 1 — Byte-exact l0_mlp.rknn (FB field 9 schema decode)

With phase 0 done, we know the patch rules are the only remaining gap. The concrete blocker is that em=0x0d non-softmax ops (Transpose, exNorm, general compute) have a per-op auxiliary LUT blob reference that's not in input_tensors[] — it's encoded in the FB operator record's field 9 (a 70-u32 vector of undocumented metadata per op).

Example: for l0_mlp.rknn, op 3 (Transpose) needs CNA_DCOMP_ADDR0 = wt_base + 0x910180 which is blob[17] (size 8192, type 6). The op's only declared input is tensor 15 (the activation), so the existing tensor_weight_blob[input_tensors[1]] lookup returns nothing.

Plan:

  • Dump field 9 for every op in l0_mlp, l0_attn (fused), and a few known-working CNN models side-by-side. Look for u32 values that map to known blob offsets. The most likely shape is a {blob_tidx, stride, count, ...} struct, possibly packed.
  • Write a parse_op_aux_blob() helper in openrknn_model.c that extracts the auxiliary blob reference per op and stores it in struct orknn_op_info (add aux_blob_tidx / aux_blob_off fields).
  • Extend the case 0x1110: branch in patch_regcmd_addresses to resolve the aux blob for em=0x0d non-softmax tasks, with op-type dispatch (Transpose / exNorm / Mul).
  • Extend the case 0x1070: branch similarly for the subset of ops where the FEATURE_DATA_ADDR walks a scratch tensor rather than the primary input.
  • Validate via diff_regcmd.py: 0 non-DMA diffs, 0 DMA-class diffs on l0_mlp over all 643 unique regcmd sections. This is the objective gate — not "my changes look reasonable".

Exit criterion. ORKNN_OWN=... bench_throughput --model l0_mlp.rknn --strategy pinned runs without ORKNN_ORACLE_PATCH set, completes 3 segments per inference, and produces an output tensor whose cosine similarity against a vendor-run reference is > 0.999.

Estimated cost: 3–5 days (dominated by FB schema RE — unknown until somebody tries). exNorm is the hardest op because it references 4+ blobs per op in different slots.


Phase 2 — l0_attn.rknn (exMatMul / exSDPAttention)

With l0_mlp running, attention shards are the next piece. Two flavors depending on how upstream is compiled:

  • Fused (smolvlm_subshards_fused/): uses exSDPAttention × 32 (one per NanoTiled Q-chunk), plus Conv × 4, Mul × 2, exNorm × 1. Much simpler lowering — exSDPAttention is a single fused op that mobilesam_encoder also uses, so its patch rules likely generalize across both models.
  • Unfused (smolvlm_subshards/): uses exMatMul × 64 + Softmax × 32 + Mul × 34. Way more tasks (19k for unfused l0_attn vs ~6k for fused) but each op is simpler. Useful as a stress test for the patching hot path.

Start with fused. exSDPAttention is the only genuinely-new op type here and also appears in mobilesam_encoder — solving it unlocks both models.

  • Run the oracle/template diff on smolvlm_subshards_fused/l0_attn.rknn.
  • Add whatever patch rules are needed for exSDPAttention, Mul, and any tile-routing quirks.
  • Verify byte-exact.
  • Verify end-to-end cosine > 0.999.

Exit criterion. Both l0_attn.rknn (fused and unfused) run end-to-end under OWN mode with cosine > 0.999 vs vendor.

Estimated cost: 2–3 days on top of Phase 1.


Phase 3 — Full 24-shard SmolVLM pipeline

All 24 shards (l0..l11 × {attn, mlp}) running under openrknn, orchestrated by the upstream project's HybridSplitEncoder Python glue.

  • Run scripts/validate.py from the upstream repo against openrknn: layer-by-layer cosine similarity for all 24 shards.
  • Run scripts/run_inference.py: full vision encoder + image embed export, compared against CPU reference.
  • Measure end-to-end vision encoder latency under OWN mode. Reference: vendor stack is 1970 ms on the same shards (from librocketnpu: SmolVLM ViT sharding vendor-stack reproduction #79's README). We expect similar — openrknn is just a runtime; it doesn't change NPU compute time.
  • Exit criterion: end-to-end vision encoder cosine > 0.99 vs CPU reference, all 24 shards byte-exact against vendor oracle.

Estimated cost: 1–2 days (mostly just running tests and fixing per-layer edge cases).


Phase 4 — Add SmolVLM shards to CI

  • Add smolvlm_l0_mlp.rknn and smolvlm_l0_attn.rknn (fused variants) to openrknn/tests/ground_truth.json as fp16_run targets, with expected cosine thresholds against a captured vendor reference.
  • Check in the captured Phase 2.5 segmentation ground-truth artifact (segmentation_ground_truth/smolvlm_l0_*.json).
  • Decide on storage for the .rknn files themselves (git-lfs? separate release? rebuildable via the Dockerfile from librocketnpu: SmolVLM ViT sharding vendor-stack reproduction #79?).
  • Ensure ci_validate.sh Phase 2 + Phase 2.5 both cover the new models without doubling CI runtime.

Exit criterion. Phase 2.5 shows 7 clean, 0 gating, 0 allowlisted (5 existing CNN + 2 new SmolVLM shards). Any future openrknn patch that regresses transformer support is caught immediately.

Estimated cost: 1 day.


Phase 5 — Lift FB extractor for mobilesam_encoder / lprnet

Separate from the SmolVLM work but cheap to do once the transformer patch rules exist. Both models currently bail out of extract_npu_data() with regcmd(0) or taskbo(0) not found and silently fall through to the vendor proxy. Their regcmd/taskbo blobs live in a different FB field layout that openrknn's scanner doesn't recognize — probably just a missing blob-type heuristic.

  • Identify the correct FB field for regcmd/taskbo in these models (they're newer rknn-toolkit2 outputs, likely use weight_data_v6 at vt offset 0x2C with different blob type bytes).
  • Extend extract_npu_data() to find them.
  • Run l0_mlp-style oracle diffs against both models and fill in whatever patch rules are still missing.
  • Promote both from fp16_parse to fp16_run in the CI suite.

Exit criterion. mobilesam_encoder and lprnet run end-to-end under OWN mode in Phase 2; Phase 2.5 shows byte-exact diffs.

Estimated cost: 2–3 days — smaller scope than SmolVLM once the op patch rules are done.


Phase 6 — Upstream the intercept_swap.c enhancements

Nice-to-have. During this session I extended librocketnpu/tests/intercept_swap.c to log task_base_addr, iommu_domain_id, and subcore_task[3..4] in the SWAP: SUBMIT line. That's useful for future submit-parameter diagnosis and should be upstreamed as its own small commit (not tied to this roadmap's correctness work).

  • Clean up the on-board hand-edited version vs the repo's canonical librocketnpu/tests/intercept_swap.c.
  • Small PR with the logging diff + a note in the file header about what new fields are captured and why.

Non-goals

Things explicitly out of scope for this roadmap:

  • Mesa rocket / openrknn cross-pollination. This issue is strictly about openrknn's OWN path consuming vendor-compiled .rknn files. Mesa rocket produces its own regcmd from ONNX — that's a completely different engineering surface tracked separately.
  • Writing our own FP16 lowering for any op. openrknn is a runtime, not a compiler. If the vendor's compiler can't produce an op, openrknn doesn't either.
  • LLM decode / KV-cache / dynamic shapes. These would be separate features tracked under openrknn: dynamic-shape model support (rknn_set_input_shape(s)) #61 (dynamic-shape) and/or openrknn: zero-copy I/O via rknn_create_mem / rknn_set_io_mem #58–59 (zero-copy I/O). SmolVLM's vision encoder is static shape — that's why it's tractable. The language model half of SmolVLM uses RKNN-LLM's .rkllm format, which is a completely different container outside openrknn's scope.
  • Optimizing vs vendor perf. Matching vendor latency for transformer shards is the goal, not beating it. openrknn doesn't compile the regcmd, so any perf delta vs vendor is noise from submit/ioctl plumbing.

Relationship to other roadmaps

Parallel to but independent of #68 (main openrknn roadmap). Most feature issues there (zero-copy I/O, dynamic shape, multi-core batch, LSTM) are orthogonal to transformer support and can land in either order. The exceptions:

Tooling references

All of these already exist and are ready to use:

  • Oracle capture: LD_PRELOAD=librocketnpu/tests/intercept_swap.so DUMP_ALL_BOS=1 bench_throughput ...
  • Template dump: ORKNN_DUMP_BO1=/path, ORKNN_DUMP_BO1_PRE=/path, ORKNN_DUMP_TASKBO=/path
  • Diff: openrknn/tests/diff_regcmd.py --oracle --template --task-bo --submit-txt --show-dma --em-filter 0x0d
  • Oracle-patch fallback: ORKNN_ORACLE_PATCH=/path ORKNN_ORACLE_{WT,ACT,IN,OUT}_BASE=0x...
  • Debug log: ORKNN_DEBUG_BLOBS=1, ORKNN_DEBUG_PATCH=1
  • Reference FB/regcmd walker: Python struct.unpack_from('<8I Q', task_bo, i*40) for 40-byte rknpu_task entries (fields: flags, op_idx, enable_mask, int_mask, int_clear, int_status, regcfg_amount, regcfg_offset, regcmd_addr)

RE references (existing):

  • ~/projects/ida/librknnrt/docs/model-parsing.md — FB schema (field 0..21)
  • ~/projects/ida/librknnrt/docs/graph-executor.md — op routing, per-node target byte
  • ~/projects/ida/librknnrt/docs/rknn-model-format.md — weight data section semantics, version-gated field layouts
  • ~/projects/ida/librknnrt/librknnrt.so.c — full Hex-Rays decompile, look at sub_2DC018 (v>5 operator extractor) for per-op field decoding hints

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions