Deploy AI Models on Jetson: What Can Run and How
This page covers the questions that come up most often when running AI models on NVIDIA Jetson — LLMs, vision-language models, generative models, embodied/VLA models, speech, and computer vision:
- What model size (in B, billion parameters) can my device run?
- Which models can I deploy?
- How do I deploy them?
Three ways to deploy:
- Online platform (recommended): browse and deploy from the reComputer AI Lab — 100+ optimized CV / LLM / VLM models for reComputer Jetson with one-command Docker deployment and benchmarks. Pick a model on the page, copy the command, run it.
- One-command CLI: jetson-examples —
reComputer run <model>. - Manual: you pick the engine (Ollama, llama.cpp, MLC, TensorRT Edge-LLM, TensorRT-Model-Connect) and follow the tutorials listed below.
1. What model size can my device run?
"Model size" is expressed in B (billion parameters) — a 7B model has 7 billion parameters, and larger B numbers need more memory. What matters is your Jetson module's unified memory (the module, not the carrier board). For a quantized (Q4_K_M) model:
8 GB · Orin Nano
1B – 4B models
Llama 3.2 1B/3B, Qwen3.5-4B, Gemma4 E4B, Live VLM WebUI
16 GB · Orin NX
7B – 8B models
Llama3 8B, DeepSeek-R1 7B, Llama2-7B (MLC), LLaVA 7B, LocateAnything
32 GB · AGX Orin
up to 27B (quantized)
Qwen3.5-27B Q4_K_M (see JetPack 6.2 vs 7.2 benchmark below)
64 GB · AGX Orin
30B – 35B+
Nemotron-3-Nano-30B (Q4_K_M), Qwen3.6-35B (UD-Q6_K), GPT-OSS
128 GB · Jetson Thor
35B+ and larger
Nemotron3 33B, JoyAI VLM / ASR / TTS stack
Memory use includes the KV cache and system overhead, so a model whose weights nearly fill the RAM will not run stably. Measured on AGX Orin 32 GB, a Qwen3.5-27B Q4_K_M model used about 24.6 GB of the 30 GB available after load — that is a realistic ceiling for 32 GB devices.
2. Which models can I deploy?
2.1 One-command deployment (jetson-examples)
Install once, then run any supported model with one command:
sudo apt install python3-pip
pip3 install jetson-examples
Qwen Series
reComputer run qwen3.5-4breComputer run qwen3.6-35bLlama Family
reComputer run llama3reComputer run llama3.2reComputer run llavaGemma Family
reComputer run gemma4NVIDIA Models (64 GB)
reComputer run nemotron-3-nanoreComputer run gpt-ossVLM (Vision-Language)
reComputer run live-vlm-webuireComputer run locateanythingYOLO & Detection
reComputer run ultralytics-yoloreComputer run yolo11reComputer run yolo26reComputer run yolo26-tensorrtreComputer run yolov10reComputer run depth-anything-v2reComputer run depth-anything-v3reComputer run nanoowlTooling
reComputer run ollamareComputer run whisperreComputer run llama-factoryNote: sizes above are for the shown quantization. Higher precision (Q8 / FP16) roughly doubles or more the weight size. Each example also needs enough free disk (image + model) — LLaVA FP16 needs about 27.4 GB total. Check the example list in the jetson-examples README for exact image sizes.
2.2 Manual deployment: full tutorial index
Each category below links to a complete tutorial on this wiki. Pick the one that matches your task and follow it on the device.
Engine at a glance
Easy setup
Q4 quantized
GGUF, full control
Production · JP 6.2
On-device build · JP 7.2
1. General LLM
VRAM estimate: reliable — params × bits/8 + KV cache + system. Q4: 7B ≈ 4-5 GB, 27B ≈ 15-18 GB.
- DeepSeek-R1 7B with Ollama
- DeepSeek 1.5B with MLC ~60 tok/s
- Llama2-7B Q4 with MLC
- GPT-OSS 20B with llama.cpp
- Structured LLM output with Langchain
- RAG with LlamaIndex + ChromaDB
- Local AI Assistant (AnythingLLM)
- Local OpenClaw (Clawdbot)
- Fine-tune with Llama-Factory
- Develop reComputer with Clawdbot
- LLM hardware interface control
- Voice control motor by LLM
2. VLM (Vision-Language)
VRAM estimate: roughly — LLM part + vision encoder (0.3-4B) + input resolution/tiles. Higher resolution scales VRAM non-linearly.
- Run VLM on reComputer
- Live VLM WebUI 7 VLMs · real-time
- JoyAI-VL-Interaction Thor · voice/video
- VLM with speech interaction
- LLaVA warehouse guard Ollama llava-llama3 8B
- Zero-Shot Detection
3. Generative & World models
VRAM estimate: not by parameters alone — DiT patch/frame count, temporal attention, U-Net vs DiT vs MoE change actual VRAM 3-5x at the same size. Measure on device.
- Stable Diffusion (text-to-image)
- ComfyUI
reComputer run comfyui - AudioCraft (music gen)
reComputer run audiocraft - Wan / video generation, world models — no verified Jetson deployment yet; VRAM must be measured per model
4. Embodied / VLA
VRAM estimate: not by parameters — vision encoder + camera count + control rate + action chunking. Measure on device.
- GR00T N1.5 + LeRobot SO-101 (Thor)
- GR00T N1.6 + LeRobot SO-101 (AGX Orin)
- GR00T N1.7 + reBot Arm (J601)
- GraspNet visual grasping (reBot-DM)
- Control SO-Arm by OpenClaw (Thor)
- Control reBot by NemoClaw (Thor)
- Jetson-Claw starter (8 GB)
- GR00T N1.7 full-weight TensorRT (JP 7.2)
- Voice-control reBot Arm B601
- Microduck RL on Jetson
- Microduck RL environment
- Microduck RL official policies
- Microduck RL custom motion
5. Speech (ASR / TTS)
VRAM estimate: small models — usually fine on 8 GB. Whisper/Riva run alongside an LLM in a pipeline.
6. Computer Vision (YOLO & detection)
VRAM estimate: small — YOLO models typically run on 8 GB with TensorRT, and scale with input resolution. See the one-command list in section 2.1.
- YOLOv8 with TensorRT
- YOLOv8 with DeepStream + TensorRT
- Train and deploy YOLOv8
- YOLOv8 custom classification
- YOLOv5 object detection
- YOLOv11 + Depth Camera
- YOLOv26 dual camera system
- Depth Anything V3
- MaskCam
- Traffic Management (DeepStream)
- DashCamNet multicamera
- Multi-GMSL cameras
- Four-camera fisheye surround view
- Multi-task vision inference engine
- Streaming Vision Agent (Qwen3-VL)
- Frigate NVR
- Security X-ray scan
- Industrial vision monitoring
- NVBlox 3D mapping
- AI NVR
3. How do I deploy?
AI Lab (online platform)
- Open reComputer AI Lab Models (Jetson pre-filtered)
- Filter by device (e.g. Jetson Orin Nano) and browse CV / LLM / VLM models with benchmarks
- Copy the one-command Docker command and run it on your device
100+ optimized models, no setup, includes benchmark numbers.
One-command CLI
Install jetson-examples, then deploy any model from section 2.1:
sudo apt install python3-pip
pip3 install jetson-examples
# Deploy a model (example)
reComputer run qwen3.5-4bExposes an OpenAI-compatible API endpoint, usable directly with curl or any OpenAI SDK.
Manual deployment
Pick an engine from section 2.2 and follow its tutorial on your device:
- Ollama — easiest start, large model library
- MLC LLM — quantized (Q4) compiled models
- llama.cpp — lightweight, GGUF, full control
- TensorRT Edge-LLM / Model-Connect — production speed
All tutorials are verified on reComputer Jetson devices.
4. Performance benchmark: JetPack 6.2 vs 7.2
Same Qwen3.5-27B Q4_K_M model (llama.cpp) on AGX Orin-class hardware, JetPack 6.2 vs 7.2:
| Metric | JetPack 6.2 | JetPack 7.2 | Change |
|---|---|---|---|
| Memory after model load | 24.6 GB / 30 GB | 14.7 GB / 30 GB | ~40% lower |
| GPU frequency during inference | 930 MHz | 1.36 GHz | higher boost |
| Prompt processing | 18.2 tok/s | 25.8 tok/s | ~41.8% faster |
| Token generation | 4.3 tok/s | 5.5 tok/s | ~27.9% faster |
Details: JetPack 7.2 Deep Dive.
5. FAQ
My model is too big for one device — what can I do?
- Use a smaller quantization (Q4_K_M instead of Q8/FP16).
- Limit the context length (
--max-sequence-length) to shrink the KV cache. - Distribute inference across multiple reComputer devices with llama.cpp RPC (tutorial).
Do I need a cloud GPU?
No. Everything above runs fully on-device. Data stays on the edge and there are no recurring API fees.
More resources
GenAI topic overview for reComputer-Jetson
Dataset → train → export → deploy
Troubleshooting and usage questions
One-command deployment repo
Tech Support & Product Discussion
Thank you for choosing our products! We are here to provide you with different support to ensure that your experience with our products is as smooth as possible. We offer several communication channels to cater to different preferences and needs.