Skip to main content

Streaming Vision Agent on Jetson

Introduction

Most Jetson vision demos stop at single-frame detection (each frame is independent) or short offline clip understanding (run a VLM once over a few seconds of recorded video). Neither keeps state across a continuous live stream, so after an object leaves the view — or after the clip ends — you usually cannot ask “what just happened a moment ago?” with evidence. A Streaming Vision Agent keeps a short online rolling multimodal memory on the edge — visual embeddings, episodic events, and semantic facts — and answers questions with evidence frames and clips while the camera is still running.

This wiki deploys a real-time demo on Seeed Jetson devices (verified on reComputer Mini J5012 · JetPack 7.2). A USB camera feeds a browser UI; two independent Qwen3-VL-2B instances handle recognition and Ask so answering does not block background memory writes.

tip

The design is inspired by WorldMM (CVPR 2026) multimodal memory ideas. This demo targets an online rolling window on Jetson — it is not a reproduction of the paper’s offline EgoLife benchmarks. See Inspiration & acknowledgments.

Verified on reComputer Mini (Jetson AGX Orin 64GB) with JetPack 7.2 (L4T R39.2.0).

Overview

LayerRole
Visual memoryVLM2Vec frame embeddings + JPEG evidence (~every 5 s)
Episodic memoryQwen3-VL-2B #1 — appear / move / disappear events (~every 45 s)
Semantic factsEntity state (is_at / absent_from / usually_at) + timeline
AskRetrieve memory → Qwen3-VL-2B #2 answers with trajectory + evidence

Open http://<jetson-ip>:8790 for live video, rolling memory, and Ask.

Camera ──► visual @ ~5s (VLM2Vec)
└──► episodic @ ~45s (Qwen3-VL-2B recognition)
Ask ──► retrieve memory ──► Qwen3-VL-2B answer

Supported Hardware

ItemConfiguration
DevicesreComputer J501 Mini
VerifiedreComputer J501 Mini · JetPack 7.2 (L4T 39.2.0)
RAM / Disk64 GB RAM recommended · ≥50 GB free disk for models + venv
CameraUSB UVC / V4L2 (/dev/video0)

Installation

1. Clone the repository

git clone https://github.com/xbs0325/Streaming-Vision-Agent-Orin.git
cd Streaming-Vision-Agent-Orin

2. Create the Jetson Python environment

bash script/jetson_setup.sh

Default venv path: ~/leucus/.venv-worldmm (override with WORLDMM_VENV).

Activate and set environment variables:

source "${WORLDMM_VENV:-$HOME/leucus/.venv-worldmm}/bin/activate"
export PYTHONPATH="$PWD/src:$PWD:$PYTHONPATH"
export WORLDMM_ATTN_IMPL=sdpa
export WORLDMM_QWEN_DEVICE_MAP=cuda:0
export WORLDMM_DTYPE=bfloat16
export WORLDMM_MODELS="${WORLDMM_MODELS:-$HOME/leucus/models/worldmm}"
export HF_HOME="$WORLDMM_MODELS/hf_home"
unset HF_ENDPOINT
caution

Do not set HF_ENDPOINT=https://hf-mirror.com on this stack — it can break huggingface_hub downloads.

3. Download models

bash script/jetson_download_models.sh
ModelRequired for default dual-2B live
Qwen3-VL-2B-InstructYes (loaded twice: recognition + Ask)
Qwen3-Embedding-4BYes
Qwen2-VL-2B-Instruct + VLM2Vec-V2.0Yes (visual memory)
Qwen3-VL-8B-InstructOptional (WORLDMM_DOWNLOAD_8B=1 or --episodic-model 8b)

Qwen weights download via ModelScope; VLM2Vec via Hugging Face (huggingface.co). First download may take a while depending on network.

Run the Live Demo

Connect a USB camera, then:

bash run.sh
# or:
python script/orin_live.py --ui-port 8790 --window-min 8 \
--visual-interval 5 --episodic-interval 45

Open in a browser:

http://<jetson-ip>:8790/

Default runtime is dual 2B (separate model instances, locks, and CUDA streams). Optional flags:

FlagMeaning
--episodic-model 8bStronger recognition with Qwen3-VL-8B
--shared-2bOne 2B for both roles (lower VRAM; Ask waits on recognition)
--window-min 10Longer rolling memory window

Smoke test (optional)

Short capture + pipeline check:

python script/orin_smoke.py --vlm qwen3vl-2b --seconds 20 \
--vlm2vec-base "$WORLDMM_MODELS/Qwen2-VL-2B-Instruct"

Demo Results

Short clips on Seeed files CDN showing rolling memory and Ask answers in the live UI.

note

In some Ask turns (see the clips and evidence thumbnails above), the text answer may name one object while the retrieved evidence JPEG / clip shows a different object from an earlier moment in the rolling window. That is expected with a short dual-2B memory demo: retrieval can attach the nearest visual evidence rather than a perfect identity match. Prefer center-framed, one-object-at-a-time interactions for cleaner results.

What You Should See

SceneExpected behavior
Steady desk viewLive frame + status pills; episodic / visual counters increase over time
Place / remove an object (e.g. earphone case) in center viewEpisodic event after confirmation; Ask can answer “what was just put down?”
Ask “What was just put down?”Answer cites timeline / facts; optional evidence JPEG or clip

Recognition uses center-biased triggers and FULL + CENTER crop panels to reduce distraction from hands / mouse at the image edge.

Models Used in This Demo

RoleDefault model
Episodic recognitionQwen3-VL-2B-Instruct (instance #1)
Ask / answerQwen3-VL-2B-Instruct (instance #2)
Text embeddingQwen3-Embedding-4B
Visual embeddingQwen2-VL-2B-Instruct + VLM2Vec-V2.0 LoRA

Troubleshooting

IssueFix
Cannot open /dev/video0Check ls /dev/video*; try --device /dev/video1
huggingface_hub FileMetadataErrorunset HF_ENDPOINT; avoid hf-mirror
Hub / transformers conflictKeep huggingface_hub>=0.34,<1 (pinned in jetson_setup.sh)
OOM / very slowDo not run other heavy GPU demos in parallel; try --shared-2b or longer --episodic-interval
Ask feels blockedConfirm you are not using --shared-2b; default dual-2B should answer on a separate stream
Port in usefuser -k 8790/tcp then relaunch

Resources

Inspiration & Acknowledgments

This edge demo is inspired by WorldMM — a dynamic multimodal memory agent for long video reasoning (CVPR 2026 Highlight). We adapt the three-memory idea (visual / episodic / semantic) to a real-time rolling window on Jetson.

@inproceedings{yeo2026worldmm,
title = {WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning},
author = {Yeo, Woongyeong and Kim, Kangsan and Yoon, Jaehong and Hwang, Sung Ju},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {25599-25609}
}

Also thanks to HippoRAG, VLM2Vec, and the upstream WorldMM implementation (Apache-2.0).

Tech Support & Product Discussion

Thank you for choosing our products! We are here to provide you with different support to ensure that your experience with our products is as smooth as possible. We offer several communication channels to cater to different preferences and needs.

Loading Comments...