HomeModel Memory Fit › Gemma 2 2B IT
Model Memory Fit · LLM

Gemma 2 2B IT: memory fit on every Jetson module.

Method v1.0 · dataset 2026-09-07 · verified 2026-09-07 · last updated September 2026

2.61 B parameters · 26 layers · 8,192-token max context. Hugging Face: google/gemma-2-2b-it.

Model facts

FactValueNote
Parameters2.614 B
Layers26
Hidden size2304
Attention heads8
KV heads4
Head dimension256
Max context8,192 tokens
dtypebfloat16
Vocabulary256,000

Hugging Face (mirror: unsloth/gemma-2-2b-it) · hf:unsloth/gemma-2-2b-it · config.json · verified 2026-09-07

“"num_attention_heads": 8, "num_key_value_heads": 4, "head_dim": 256”

Parameters source

  • Hugging Face (mirror: unsloth/gemma-2-2b-it, identical weights to google/gemma-2-2b-it) · hf:unsloth/gemma-2-2b-it · verified 2026-09-07 · model.safetensors (blob metadata via HF API) · class A
    “model.safetensors size 5228717512 bytes”

Published quantised artefacts

A file the publisher or a community mirror actually ships, used as the class-A weight figure when the requested quant matches. Any quant without a row here falls back to parameters × bytes-per-parameter (class D for block quants).

QuantFormatFileSizeSource
GGUF Q4_K_M / INT4 (AWQ, MLC q4f16)GGUF Q4_K_Mgemma-2-2b-it-Q4_K_M.gguf1.709 GBbartowski (community GGUF) · hf:bartowski/gemma-2-2b-it-GGUF · verified 2026-09-07

Jetson tokens/s measurements

Published or archived throughput numbers, not modelled. Class C (external measured benchmark).

ModuleRuntimeQuantMeasuredSource
Jetson Orin Nano Superunspecified (archived benchmark)34.97 tok/sNVIDIA Jetson AI Lab (archive) · jetson-ai-lab:benchmarks.html · verified 2026-09-07

Every module × every quantisation

Verdict and total memory at 4096-token context, 1 sequence, llama.cpp, headless. Q4 cells link to the static breakdown page; every other cell links to the live engine at that quantisation.

Notes

google/gemma-2-2b-it config.json/README return 401 (gated); unsloth mirror (byte-identical weights) used.

Method and data: Model Memory Fit methodology. Full registry: model-memory.json.

Change the context, concurrency or runtime.

The live engine covers any context length, sequence count, KV precision and runtime, with a permanent link.

OPEN MODEL MEMORY FIT →