GPU Sizing — methodology

Method version 2.0.2 · engine: /engines/gpu-sizing/ · reads the hardware catalog (vendor peak TOPS, memory, price) and its own model compute table.

A legacy planning envelope. It has its own method, described here; it does not share Compute Fit's or Hardware Match's method, and nothing on this page describes theirs.

Contents

  1. Purpose, successors and transition status
  2. Inputs
  3. Required compute and platform capacity
  4. The platform score
  5. What the result says
  6. Evidence limitations
  7. Method changelog

1. Purpose, successors and transition status

GPU Sizing estimates how much compute an inference model needs, in TOPS, and recommends one catalog platform by a heuristic score. It answers POST /api/v1/gpu-sizing. The unit of the whole method is TOPS: it computes no frame rate, reads no benchmark database row to decide anything, and filters no platform out — every catalog platform is scored and the highest score is the recommendation.

Successors. Compute Fit answers the throughput question for one named platform from measured benchmark rows, with JetPack, runtime and memory checks, and recommends no platform. Hardware Match chooses a platform and node count from a complete-configuration check of the whole workload. Every GPU Sizing result carries scope_disclosure: Compute Fit as the owner of the question, a hand-off link that opens it with the same model family and precision, and the three differences — catalog TOPS rather than the benchmark database, a planning envelope rather than measured frames per second, and no JetPack, thermal or model-memory checks.

Transition status. Live and legacy. The route and the page stay; the plan is to answer this page from Compute Fit and retire this second implementation. Until that happens this page describes the method the route actually runs, and its URL will be kept after retirement.

2. Inputs

All five inputs are required; none is defaulted. A missing one is a 400 naming it, and so is a stated value that is not accepted or a key that is not an input.

InputAcceptedWhat it does
model_archa model on the engine's compute table (YOLOv8 / YOLO11 / YOLO12 n–x, ResNet, EfficientNet, EfficientDet, MobileNet, SSD-MobileNet, SAM, DeepLabv3) or a catalog modelSets the model's FP16 compute figure.
precisionfp32, fp16, bf16, int8, int4Scales the compute figure; selects the platform's INT8 or FP16 TOPS.
batch_sizeinteger 1–32Multiplies the compute figure.
latency_classrealtime, low, moderate, batch_okScore bonus when the class is on the platform's latency list (section 4).
throughput_classlow, medium, high, very_highScore bonus when the class is on the platform's throughput list.

3. Required compute and platform capacity

required_tops = model_tops_fp16 × precision_multiplier × batch_size precision_multiplier: fp32 2.0 · fp16 1.0 · bf16 1.0 · int8 0.5 · int4 0.25 headroom = platform_tops ÷ required_tops (platform INT8 TOPS for int8/int4, FP16 TOPS otherwise)

A model's FP16 compute figure comes from the engine's own table, hand-calibrated from published GFLOPs. A model not on the table falls back to its catalog GFLOPs divided by a family factor (4.5 for YOLOv8, 6.25 for YOLO11 and YOLO12, 2.0 for ResNet, MobileNet and EfficientNet, 6.25 for other families, 4.0 when no family matches). Platform capacity is the catalog's vendor peak TOPS; where the catalog figure is not the FP16/INT8 figure (Jetson Thor's catalog peak is FP4) the engine overrides it.

4. The platform score

Each of the eleven scored platforms gets an additive score; the platforms are sorted by it, the top one is recommended_platform and the next two are alternatives. Nothing is eliminated: a platform that fails the memory rule is scored down, and when the top pick fails it the result warns.

  • Memory: −50 when the model's memory at the precision exceeds 80 % of the platform's memory, or when the platform has no memory figure and the model is over 100 MB.
  • Headroom: +30 for 2–20×, +15 for 20× and above, +20 for 1–2×, +5 for 0.7–1×, −20 below 0.7×.
  • Latency class: +20 when the requested class is on the platform's list, −10 otherwise. This is a membership test on each platform's list, not an ordered preference — an open finding recorded for the successor, which treats latency as an ordered requirement.
  • Throughput class: +20 when it is on the platform's list; model fit: +15 when the model is on the platform's good-for list; precision fit: +10 for integer precision on an NPU/TPU accelerator or floating point on a Jetson; runtime: the runtime-efficiency figure × 10; price: when headroom is at least 1×, max(1, round(12 − log2(price ÷ 50))) points.

5. What the result says

One kernel constraint, inference in TOPS (required against the platform envelope), class D; compute_requirement_estimate, platform_compute_envelope and planning_headroom (1 ÷ that constraint's utilization); the recommended platform with its cost, tagline and strengths; two alternatives; the kernel verdict fields; and scope_disclosure. A benchmark row matching the top pick, model and precision may raise the displayed confidence; it never changes the pick.

6. Evidence limitations

  • Platform TOPS are vendor peaks; a model rarely sustains them. Headroom in TOPS is not headroom in frames per second.
  • Model compute figures are hand-calibrated from GFLOPs, and the fallback divisors are empirical factors — class D/E, not measurements.
  • The latency, throughput and good-for lists, the score weights, the 80 % memory rule and the 100 MB rule are heuristics.
  • No JetPack, thermal, decode or model-memory-fit check is made. For a platform you already have, use Compute Fit; to choose a platform, use Hardware Match.

7. Method changelog

MethodDateChange
2.0.22026-09-28One accepted-key list (WP-02): a request key that is not one of the five inputs (or a documented alias) is a 400 naming the key, where it used to be ignored.
2.0.12026-09-28Present-invalid inputs are rejected (WP-01): a batch_size that is not an integer from 1 to 32 is a 400 naming the field. Before, a value like "abc" passed a numeric comparison and was sized at 0.001 TOPS (an AGX Orin at 275,000× headroom). No verdict change for valid inputs.
2.02026-03-14The planning-envelope compute model: required TOPS from the model's compute figure, precision and batch; platforms scored against catalog TOPS. Later releases without a method change added the kernel constraint (headroom read from the constraint), scope_disclosure with the Compute Fit hand-off, and a rounding fix so the headroom and the estimate agree.

Method changes bump the method version; results carry it in provenance.method_version with the dataset version of the catalog the engine reads.