Skip to content

openrknn: SmolVLM vision-language demo for Orange Pi 5+ - #95

Merged
widgetii merged 2 commits into
masterfrom
openrknn/smolvlm-demo
May 8, 2026
Merged

widgetii merged 2 commits into
masterfrom
openrknn/smolvlm-demo

Conversation

@widgetii

Copy link
Copy Markdown
Owner

Summary

Self-contained demo that runs full SmolVLM vision-language pipeline on the Orange Pi 5+ NPU:

  • Vision encoder: 24 FP16 transformer shards on RK3588 NPU via openrknn (~1.8s)
  • Language model: SmolVLM-256M-Instruct on CPU (~350ms/token)
  • Supports image input, video frame extraction, custom prompts
  • Optional --compare flag for openrknn vs vendor rknnlite2 timing comparison

Usage

source /root/npu-research/venv/bin/activate
python3 openrknn/examples/smolvlm_demo.py --image photo.jpg
python3 openrknn/examples/smolvlm_demo.py --video clip.mp4 --frame 100
python3 openrknn/examples/smolvlm_demo.py --image photo.jpg --compare

Test plan

  • Tested on board with dog and bus images
  • Vision encoder timing ~1.8s (24 shards)
  • Language model generates text output
  • Clean output when stderr suppressed

🤖 Generated with Claude Code

Self-contained demo script that runs the full SmolVLM pipeline:
1. Image input (or video frame extraction via ffmpeg)
2. SigLIP vision encoder on NPU (24 FP16 shards via openrknn OWN mode)
3. SmolVLM-256M language model on CPU (HuggingFace transformers)
4. Natural-language image description output with timing breakdown

Features:
- --compare flag: side-by-side openrknn vs vendor rknnlite2 timing
- --video/--frame: extract frames from video files
- --prompt: custom questions about the image
- Auto-discovers openrknn library and shard directory

Vision encoder: ~1.8s on NPU (24 shards, bit-exact with vendor, 12% faster)
Language model: ~350ms/token on CPU (SmolVLM-256M-Instruct)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
First pass at attention shard DMA address fixes:
- exSDPAttention DST: route to input[3] (output buffer) instead of
  output[0]. Eliminates 3072 DST diffs.
- exSDPAttention SRC: route REFORMAT tasks to input[2] (V tensor).
  Partially correct — some sub-tasks need input[1] (K) or input[3]
  (output) instead. Eliminates 384 SRC diffs.
- pc2/pc3 heuristic: extend exNorm detection to include exSDPAttention
  ops for the gap-case swap.

Progress: 11,691 → 8,235 DMA-class diffs (30% reduction).
Remaining: 4962 SRC, 2122 CNA_DCOMP, 759 DST, 352 CNA_FEAT, 40 PPU.

The exSDPAttention REFORMAT structure is more complex than exNorm:
different sub-tasks within the same op read from different inputs
(Q/K/V/output), requiring per-sub-task routing similar to exNorm's
em0d_block_idx tracking.

Tracked by #96.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@widgetii
widgetii merged commit 63a1e99 into master May 8, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant