HomeModel Memory FitQwen2.5-VL 3B Instruct › Jetson Orin Nano 8GB
Model Memory Fit · Qwen2.5-VL 3B Instruct on Jetson Orin Nano 8GB

Qwen2.5-VL 3B Instruct on Jetson Orin Nano 8GB: FITS

Method v1.0 · dataset 2026-09-07 · verified 2026-09-07 · last updated September 2026

Qwen2.5-VL 3B Instruct at GGUF Q4_K_M / INT4 (AWQ, MLC q4f16), 4096 tokens × 1 sequence on Jetson Orin Nano 8GB: 4.6 GB of 8.0 GB — FITS (58%). Weights 1.8 GB (class A) + KV cache 0.2 GB + runtime 0.5 GB + OS 0.9 GB + reserve. Largest context that fits at this concurrency: 32768 tokens; 11 concurrent sequences at 4096.

Q4 quantisation · 4096-token context · 1 sequence · llama.cpp · headless. Change any of those on the live engine.

Memory breakdown

PartAmountEvidenceBasis
Weights1.8 GBclass AQwen2.5-VL-3B-Instruct-Q4_K_M.gguf: 1.93 GB published file size
KV cache0.2 GBclass D2 × layers × KV heads × head dim × 5376 tokens × 2 bytes × 1 sequence
Vision tower0.8 GBclass A/DGGUF mmproj Q8_0 (vision tower) 0.845 GB published file size
Runtime overhead0.5 GBclass Ellama.cpp / Ollama: 400 MB base + 6% of weights (CUDA context + compute buffer for the prompt batch (n_batch 512))
OS headroom0.9 GBclass Eheadless
Reserve0.4 GBclass E10% of subtotal
Total4.6 GBof 8.0 GB — 58% — FITS

Weights source

  • ggml-org (llama.cpp maintainers) · hf:ggml-org/Qwen2.5-VL-3B-Instruct-GGUF · verified 2026-09-07 · source · class A
    Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf 1.93 GB (published file size)

Headroom

Largest context that fits at 1 sequence: 32768 tokens. Max concurrent sequences at 4096 tokens: 11. Sequence ladder: 1 comfortable (≤60%) · 11 likely viable (≤85%) · fails at 18.

Other quantisations on Jetson Orin Nano 8GB

QuantisationWeight sourceTotalVerdict
FP16 / BF16estimated9.3 GBDOES NOT FIT
FP8 (E4M3)estimated5.9 GBFITS
INT8 / GGUF Q8_0estimated6.1 GBFITS
GGUF Q4_K_M / INT4 (AWQ, MLC q4f16)measured artefact4.6 GBFITS
NVFP4estimated4.4 GBFITS

Qwen2.5-VL 3B Instruct on other modules

Constraints

TypeRequiredAvailableUtilizationStatusEvidence · confidenceSource / method
memory4,723 MB8,192 MB58%PASSD · MEDIUMNVIDIA · nvidia.com:jetson-modules-spec-table · verified 2026-09-07

Sources

  • NVIDIA · nvidia.com:jetson-modules-spec-table · verified 2026-09-07 · source · class A
    Jetson Orin Nano 8GB: 8 GB
  • Qwen (Alibaba) · hf:Qwen/Qwen2.5-VL-3B-Instruct · config.json · verified 2026-09-07 · source · class A
    layers 36, kv_heads 2, head_dim 128, max context 128000
  • ggml-org (llama.cpp maintainers) · hf:ggml-org/Qwen2.5-VL-3B-Instruct-GGUF · verified 2026-09-07 · source · class A
    Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf 1.93 GB (published file size)
  • ggml-org (llama.cpp maintainers) · hf:ggml-org/Qwen2.5-VL-3B-Instruct-GGUF · verified 2026-09-07 · source · class A
    GGUF mmproj Q8_0 (vision tower) 0.845 GB
  • NVIDIA · nvidia-blog:mastering-llm-techniques-inference-optimization · Key-value caching · verified 2026-09-07 · source · class A
    KV cache size per token = 2 × (num_layers) × (num_heads × dim_head) × precision_in_bytes
  • NVIDIA · nvidia-docs:cuda-for-tegra-appnote · Memory Management · verified 2026-09-07 · source · class A
    Jetson iGPU and CPU share the same DRAM: OS, runtime and model all draw on one pool.

Assumptions

  • Weights: Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf: 1.93 GB published file size (class A).
  • KV cache: 2 × 36 layers × 2 KV heads × 128 head dim × 5376 tokens × 2 bytes × 1 sequence = 189 MB (class D formula).
  • 1 image × 1280 tokens added to the context (dynamic: (H/28)×(W/28) tokens after 2×2 spatial merge; 1280 assumed for a ~1 MP image; class D).
  • Vision tower: GGUF mmproj Q8_0 (vision tower) 0.845 GB published file size (class A).
  • Runtime overhead: llama.cpp / Ollama: 400 MB base + 6% of weights (CUDA context + compute buffer for the prompt batch (n_batch 512)) (class E).
  • OS 900 MB (headless) + 10% reserve from the shared memory budget (class E).

Method and limitations

total = weights (published artefact size, else parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes × sequences) + vision tower (VLM) + runtime overhead + OS headroom + reserve, on the memory budget shared with Camera Stream Capacity. Verdict: FITS < 85% of module memory ≤ TIGHT < 100% ≤ DOES_NOT_FIT. Vision pipelines take their per-part figures from the memory estimator and are rebuilt on the same budget.

  • Weights are class A when a published quantised artefact matches; otherwise parameters × bytes-per-parameter is class D for block quants (Q4, INT8, NVFP4) and class A for fixed-width types (FP16, FP8).
  • KV cache, runtime overhead, OS headroom and reserve are class D/E formulas — engineering models, not vendor measurements. Validate on device with sudo tegrastats; llama.cpp prints its own KV and compute-buffer sizes at load.
  • Full method, evidence classes and the confidence rule: Model Memory Fit methodology.

Change one input and re-run.

The live engine keeps every source line and gives you a fresh permalink and summary.

OPEN THIS WORKLOAD LIVE →