openrknn: SmolVLM vision-language demo for Orange Pi 5+ - #95
Merged
Merged
Conversation
Self-contained demo script that runs the full SmolVLM pipeline: 1. Image input (or video frame extraction via ffmpeg) 2. SigLIP vision encoder on NPU (24 FP16 shards via openrknn OWN mode) 3. SmolVLM-256M language model on CPU (HuggingFace transformers) 4. Natural-language image description output with timing breakdown Features: - --compare flag: side-by-side openrknn vs vendor rknnlite2 timing - --video/--frame: extract frames from video files - --prompt: custom questions about the image - Auto-discovers openrknn library and shard directory Vision encoder: ~1.8s on NPU (24 shards, bit-exact with vendor, 12% faster) Language model: ~350ms/token on CPU (SmolVLM-256M-Instruct) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
First pass at attention shard DMA address fixes: - exSDPAttention DST: route to input[3] (output buffer) instead of output[0]. Eliminates 3072 DST diffs. - exSDPAttention SRC: route REFORMAT tasks to input[2] (V tensor). Partially correct — some sub-tasks need input[1] (K) or input[3] (output) instead. Eliminates 384 SRC diffs. - pc2/pc3 heuristic: extend exNorm detection to include exSDPAttention ops for the gap-case swap. Progress: 11,691 → 8,235 DMA-class diffs (30% reduction). Remaining: 4962 SRC, 2122 CNA_DCOMP, 759 DST, 352 CNA_FEAT, 40 PPU. The exSDPAttention REFORMAT structure is more complex than exNorm: different sub-tasks within the same op read from different inputs (Q/K/V/output), requiring per-sub-task routing similar to exNorm's em0d_block_idx tracking. Tracked by #96. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Self-contained demo that runs full SmolVLM vision-language pipeline on the Orange Pi 5+ NPU:
--compareflag for openrknn vs vendor rknnlite2 timing comparisonUsage
source /root/npu-research/venv/bin/activate python3 openrknn/examples/smolvlm_demo.py --image photo.jpg python3 openrknn/examples/smolvlm_demo.py --video clip.mp4 --frame 100 python3 openrknn/examples/smolvlm_demo.py --image photo.jpg --compareTest plan
🤖 Generated with Claude Code