Skip to main content

Deploy AI Models on Jetson: What Can Run and How

This page covers the questions that come up most often when running AI models on NVIDIA Jetson — LLMs, vision-language models, generative models, embodied/VLA models, speech, and computer vision:

  1. What model size (in B, billion parameters) can my device run?
  2. Which models can I deploy?
  3. How do I deploy them?

Three ways to deploy:

  • Online platform (recommended): browse and deploy from the reComputer AI Lab — 100+ optimized CV / LLM / VLM models for reComputer Jetson with one-command Docker deployment and benchmarks. Pick a model on the page, copy the command, run it.
  • One-command CLI: jetson-examples — reComputer run <model>.
  • Manual: you pick the engine (Ollama, llama.cpp, MLC, TensorRT Edge-LLM, TensorRT-Model-Connect) and follow the tutorials listed below.

1. What model size can my device run?​

"Model size" is expressed in B (billion parameters) — a 7B model has 7 billion parameters, and larger B numbers need more memory. What matters is your Jetson module's unified memory (the module, not the carrier board). For a quantized (Q4_K_M) model:

8 GB · Orin Nano

1B – 4B models

Llama 3.2 1B/3B, Qwen3.5-4B, Gemma4 E4B, Live VLM WebUI

16 GB · Orin NX

7B – 8B models

Llama3 8B, DeepSeek-R1 7B, Llama2-7B (MLC), LLaVA 7B, LocateAnything

32 GB · AGX Orin

up to 27B (quantized)

Qwen3.5-27B Q4_K_M (see JetPack 6.2 vs 7.2 benchmark below)

64 GB · AGX Orin

30B – 35B+

Nemotron-3-Nano-30B (Q4_K_M), Qwen3.6-35B (UD-Q6_K), GPT-OSS

128 GB · Jetson Thor

35B+ and larger

Nemotron3 33B, JoyAI VLM / ASR / TTS stack

Memory use includes the KV cache and system overhead, so a model whose weights nearly fill the RAM will not run stably. Measured on AGX Orin 32 GB, a Qwen3.5-27B Q4_K_M model used about 24.6 GB of the 30 GB available after load — that is a realistic ceiling for 32 GB devices.

2. Which models can I deploy?​

2.1 One-command deployment (jetson-examples)​

Install once, then run any supported model with one command:

sudo apt install python3-pip
pip3 install jetson-examples

Qwen Series

Qwen3.5-4B~2.5 GB (Q4_K_M) · Orin Nano 8 GB · JP 6.1/6.2/6.2.1
reComputer run qwen3.5-4b
Qwen3.6-35B~24 GB+ (UD-Q6_K) · AGX Orin 64 GB · JP 6.1/6.2/6.2.1
reComputer run qwen3.6-35b

Llama Family

Llama3 8B~4.9 GB (Ollama Q4) · Orin NX 16 GB · JP 5.1.1-6.0
reComputer run llama3
Llama 3.2 1B/3B~1.3 / 2.0 GB (Ollama Q4) · Orin Nano 8 GB · JP 5.1.1-6.x
reComputer run llama3.2
LLaVA 7B(VLM) ~15 GB total (FP16) · Orin NX 16 GB · JP 5.1.1-6.0
reComputer run llava

Gemma Family

Gemma4 E4B~2.5 GB (Q4_K_M) · Orin Nano 8 GB · JP 6.1/6.2/6.2.1
reComputer run gemma4

NVIDIA Models (64 GB)

Nemotron-3-Nano-30B~24.5 GB (Q4_K_M) · AGX Orin 64 GB · JP 6.1/6.2/6.2.1
reComputer run nemotron-3-nano
GPT-OSS~39 GB (Q4_K) · AGX Orin 64 GB · JP 6.1/6.2/6.2.1
reComputer run gpt-oss

VLM (Vision-Language)

Live VLM WebUI7 VLMs (Ollama Q4) · 8 GB min · JP 6.0-7.1
reComputer run live-vlm-webui
LocateAnything~7.3 GB weights (BF16) · Orin NX 16 GB · JP 6.x
reComputer run locateanything

YOLO & Detection

Ultralytics YOLOdetection / segmentation / pose / classification · 8 GB + · JP 4.6-6.2
reComputer run ultralytics-yolo
YOLO11· 8 GB + · JP 5.1.1-6.x
reComputer run yolo11
YOLO26· 8 GB + · JP 5.1.1-7.1
reComputer run yolo26
YOLO26 TensorRT C++native, no Docker · JP 6.0-7.2
reComputer run yolo26-tensorrt
YOLOv10· 8 GB + · JP 5.1.1-6.0
reComputer run yolov10
Depth Anything V2monocular depth · 16 GB · JP 5.1.1-5.1.3
reComputer run depth-anything-v2
Depth Anything V3· 16 GB · JP 6.1-6.2.1
reComputer run depth-anything-v3
NanoOWLopen-vocabulary detection · 8 GB + · JP 5.1.1-6.x
reComputer run nanoowl

Tooling

Ollama(any model) size depends on model + quant (usually Q4) · 8 GB + · JP 5.1.1-6.x
reComputer run ollama
Whisper(STT) size depends on Whisper variant (small/base/large) · 8 GB + · JP 5.1.1-6.x
reComputer run whisper
Llama-Factory(fine-tune) training footprint depends on model + LoRA/QLoRA · 16 GB + · JP 5.1.1-5.1.3
reComputer run llama-factory

Note: sizes above are for the shown quantization. Higher precision (Q8 / FP16) roughly doubles or more the weight size. Each example also needs enough free disk (image + model) — LLaVA FP16 needs about 27.4 GB total. Check the example list in the jetson-examples README for exact image sizes.

2.2 Manual deployment: full tutorial index​

Each category below links to a complete tutorial on this wiki. Pick the one that matches your task and follow it on the device.

Engine at a glance

Ollama
Easy setup
MLC LLM
Q4 quantized
llama.cpp
GGUF, full control
TensorRT Edge-LLM
Production · JP 6.2
TensorRT-Model-Connect
On-device build · JP 7.2

2. VLM (Vision-Language)

VRAM estimate: roughly — LLM part + vision encoder (0.3-4B) + input resolution/tiles. Higher resolution scales VRAM non-linearly.

3. Generative & World models

VRAM estimate: not by parameters alone — DiT patch/frame count, temporal attention, U-Net vs DiT vs MoE change actual VRAM 3-5x at the same size. Measure on device.

5. Speech (ASR / TTS)

VRAM estimate: small models — usually fine on 8 GB. Whisper/Riva run alongside an LLM in a pipeline.

3. How do I deploy?​

AI Lab (online platform)

  1. Open reComputer AI Lab Models (Jetson pre-filtered)
  2. Filter by device (e.g. Jetson Orin Nano) and browse CV / LLM / VLM models with benchmarks
  3. Copy the one-command Docker command and run it on your device

100+ optimized models, no setup, includes benchmark numbers.

One-command CLI

Install jetson-examples, then deploy any model from section 2.1:

sudo apt install python3-pip
pip3 install jetson-examples

# Deploy a model (example)
reComputer run qwen3.5-4b

Exposes an OpenAI-compatible API endpoint, usable directly with curl or any OpenAI SDK.

Manual deployment

Pick an engine from section 2.2 and follow its tutorial on your device:

  • Ollama — easiest start, large model library
  • MLC LLM — quantized (Q4) compiled models
  • llama.cpp — lightweight, GGUF, full control
  • TensorRT Edge-LLM / Model-Connect — production speed

All tutorials are verified on reComputer Jetson devices.

4. Performance benchmark: JetPack 6.2 vs 7.2​

Same Qwen3.5-27B Q4_K_M model (llama.cpp) on AGX Orin-class hardware, JetPack 6.2 vs 7.2:

MetricJetPack 6.2JetPack 7.2Change
Memory after model load24.6 GB / 30 GB14.7 GB / 30 GB~40% lower
GPU frequency during inference930 MHz1.36 GHzhigher boost
Prompt processing18.2 tok/s25.8 tok/s~41.8% faster
Token generation4.3 tok/s5.5 tok/s~27.9% faster

Details: JetPack 7.2 Deep Dive.

5. FAQ​

Can I run DeepSeek on Jetson?

Yes. DeepSeek-R1 7B runs well on Orin NX 16 GB devices via Ollama (tutorial); the MLC-quantized 1.5B variant reaches about 60 tok/s on Orin NX (tutorial).

My model is too big for one device — what can I do?

  • Use a smaller quantization (Q4_K_M instead of Q8/FP16).
  • Limit the context length (--max-sequence-length) to shrink the KV cache.
  • Distribute inference across multiple reComputer devices with llama.cpp RPC (tutorial).

Do I need a cloud GPU?

No. Everything above runs fully on-device. Data stays on the edge and there are no recurring API fees.

More resources​

Generative AI Intro

GenAI topic overview for reComputer-Jetson

Fine-tune with Llama-Factory

Dataset → train → export → deploy

Jetson FAQ

Troubleshooting and usage questions

jetson-examples

One-command deployment repo

Tech Support & Product Discussion​

Thank you for choosing our products! We are here to provide you with different support to ensure that your experience with our products is as smooth as possible. We offer several communication channels to cater to different preferences and needs.

Loading Comments...