Qwen2.5-VL 7B Instruct on Jetson AGX Orin 64GB: FITS
Method v1.0 · dataset 2026-09-07 · verified 2026-09-07 · last updated September 2026
Qwen2.5-VL 7B Instruct at GGUF Q4_K_M / INT4 (AWQ, MLC q4f16), 4096 tokens × 1 sequence on Jetson AGX Orin 64GB: 7.7 GB of 64 GB — FITS (12%). Weights 4.4 GB (class A) + KV cache 0.3 GB + runtime 0.7 GB + OS 0.9 GB + reserve. Largest context that fits at this concurrency: 65536 tokens; 148 concurrent sequences at 4096.
Q4 quantisation · 4096-token context · 1 sequence · llama.cpp · headless. Change any of those on the live engine.
Memory breakdown
| Part | Amount | Evidence | Basis |
|---|---|---|---|
| Weights | 4.4 GB | class A | Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf: 4.683 GB published file size |
| KV cache | 0.3 GB | class D | 2 × layers × KV heads × head dim × 5376 tokens × 2 bytes × 1 sequence |
| Vision tower | 0.8 GB | class A/D | GGUF mmproj Q8_0 (vision tower) 0.853 GB published file size |
| Runtime overhead | 0.7 GB | class E | llama.cpp / Ollama: 400 MB base + 6% of weights (CUDA context + compute buffer for the prompt batch (n_batch 512)) |
| OS headroom | 0.9 GB | class E | headless |
| Reserve | 0.7 GB | class E | 10% of subtotal |
| Total | 7.7 GB | of 64 GB — 12% — FITS | |
Weights source
- ggml-org (llama.cpp maintainers) · hf:ggml-org/Qwen2.5-VL-7B-Instruct-GGUF · verified 2026-09-07 · source · class A
Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf 4.683 GB (published file size)
Headroom
Largest context that fits at 1 sequence: 65536 tokens. Max concurrent sequences at 4096 tokens: 148. Sequence ladder: 64 comfortable (≤60%) · 64 likely viable (≤85%) · fails at —.
Other quantisations on Jetson AGX Orin 64GB
| Quantisation | Weight source | Total | Verdict |
|---|---|---|---|
| FP16 / BF16 | estimated | 19 GB | FITS |
| FP8 (E4M3) | estimated | 11 GB | FITS |
| INT8 / GGUF Q8_0 | estimated | 11 GB | FITS |
| GGUF Q4_K_M / INT4 (AWQ, MLC q4f16) | measured artefact | 7.7 GB | FITS |
| NVFP4 | estimated | 7.3 GB | FITS |
Qwen2.5-VL 7B Instruct on other modules
Constraints
| Type | Required | Available | Utilization | Status | Evidence · confidence | Source / method |
|---|---|---|---|---|---|---|
| memory | 7,909 MB | 65,536 MB | 12% | PASS | D · MEDIUM | NVIDIA · nvidia.com:jetson-modules-spec-table · verified 2026-09-07 |
Sources
- NVIDIA · nvidia.com:jetson-modules-spec-table · verified 2026-09-07 · source · class A
Jetson AGX Orin 64GB: 64 GB
- Qwen (Alibaba) · hf:Qwen/Qwen2.5-VL-7B-Instruct · config.json · verified 2026-09-07 · source · class A
layers 28, kv_heads 4, head_dim 128, max context 128000
- ggml-org (llama.cpp maintainers) · hf:ggml-org/Qwen2.5-VL-7B-Instruct-GGUF · verified 2026-09-07 · source · class A
Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf 4.683 GB (published file size)
- ggml-org (llama.cpp maintainers) · hf:ggml-org/Qwen2.5-VL-7B-Instruct-GGUF · verified 2026-09-07 · source · class A
GGUF mmproj Q8_0 (vision tower) 0.853 GB
- NVIDIA · nvidia-blog:mastering-llm-techniques-inference-optimization · Key-value caching · verified 2026-09-07 · source · class A
KV cache size per token = 2 × (num_layers) × (num_heads × dim_head) × precision_in_bytes
- NVIDIA · nvidia-docs:cuda-for-tegra-appnote · Memory Management · verified 2026-09-07 · source · class A
Jetson iGPU and CPU share the same DRAM: OS, runtime and model all draw on one pool.
Assumptions
- Weights: Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf: 4.683 GB published file size (class A).
- KV cache: 2 × 28 layers × 4 KV heads × 128 head dim × 5376 tokens × 2 bytes × 1 sequence = 294 MB (class D formula).
- 1 image × 1280 tokens added to the context (dynamic: (H/28)×(W/28) tokens after 2×2 spatial merge; 1280 assumed for a ~1 MP image; class D).
- Vision tower: GGUF mmproj Q8_0 (vision tower) 0.853 GB published file size (class A).
- Runtime overhead: llama.cpp / Ollama: 400 MB base + 6% of weights (CUDA context + compute buffer for the prompt batch (n_batch 512)) (class E).
- OS 900 MB (headless) + 10% reserve from the shared memory budget (class E).
Method and limitations
total = weights (published artefact size, else parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes × sequences) + vision tower (VLM) + runtime overhead + OS headroom + reserve, on the memory budget shared with Camera Stream Capacity. Verdict: FITS < 85% of module memory ≤ TIGHT < 100% ≤ DOES_NOT_FIT. Vision pipelines take their per-part figures from the memory estimator and are rebuilt on the same budget.
- Weights are class A when a published quantised artefact matches; otherwise parameters × bytes-per-parameter is class D for block quants (Q4, INT8, NVFP4) and class A for fixed-width types (FP16, FP8).
- KV cache, runtime overhead, OS headroom and reserve are class D/E formulas — engineering models, not vendor measurements. Validate on device with
sudo tegrastats; llama.cpp prints its own KV and compute-buffer sizes at load. - Full method, evidence classes and the confidence rule: Model Memory Fit methodology.
Links
- Live engine (this exact workload): https://edgeaistack.ai/engines/model-memory-fit/?mode=vlm&hw=jetson_agx_orin&model=qwen2.5-vl-7b-instruct&quant=q4&ctx=4096&seq=1&rt=llama_cpp&os=headless&mv=1.0&dv=2026-09-07
- Methodology: /methodology/model-memory-fit/
- Dataset: /datasets/model-memory.json
- Hugging Face repository: Qwen/Qwen2.5-VL-7B-Instruct
Change one input and re-run.
The live engine keeps every source line and gives you a fresh permalink and summary.