GPU sizing is not “parameters times two, divided by HBM.” That answers one narrow question: can a checkpoint’s nominal payload fit? A production service must also hold attention state, allocator blocks, kernels, collectives, graph captures and transient activations—then do useful work quickly enough under a distribution of prompts, outputs and arrivals.
The practical answer is a pair, not a number: a feasible topology and a measured operating envelope. The topology must fit every persistent and transient byte. The envelope must meet TTFT, inter-token latency and throughput objectives at the percentiles that matter. Mathematics eliminates impossible designs and supplies lower bounds. Only a controlled benchmark turns those bounds into capacity.
Fit is necessary, never sufficient. A configuration that starts can still miss every service-level objective; excellent offline tokens/s can coexist with unacceptable queueing latency.
Phase 0 · Fix the language
Count the same thing on both sides
A gigabyte is 109 bytes; a gibibyte is 230 bytes. An 80 GB nameplate converts to 74.51 GiB arithmetically, but packaging, firmware reservations, ECC mode and driver reporting mean runtime-visible capacity must be queried rather than inferred from the label. Network links are commonly quoted in gigabits/s while memory links use gigabytes/s. Thus 400 Gb/s is 50 GB/s raw line rate, not 400 GB/s, before overhead.
E2E ≈ TTFT + (O − 1) × mean_ITLThe minus one matters because the first token ends TTFT. More important: an average ITL cannot establish a p99 E2E objective. Say whether “tokens/s” means per request, replica, GPU or fleet. The vLLM metrics reference separates queue, prefill, generation and inter-token components.
Arrival rate also does not equal concurrency. Under stable steady state, Little’s Law gives average in-flight requests:
N̄ = λ × W̄It does not size a burst buffer or prove a tail SLA. Preserve the joint trace of arrival gaps, input length, output length, prefix identity and priority. Correlations—long prompts that also generate long answers—are where averages fail.
Phase 1 · Reconstruct the model
Parameters are a consequence of architecture
Start from the exact checkpoint revision, not a family nickname. Record hidden width H, layers L, query heads Hq, KV heads Hkv, head dimension D, vocabulary, tied embeddings, intermediate width, expert count and experts selected per token. Usually D = H / Hq, but modern architectures can split positional and non-positional dimensions.
Llama 3.1 70B is a clean dense example: 80 layers, width 8,192, 64 query heads, 8 KV heads and head dimension 128. That 8:1 ratio is grouped-query attention (GQA). Multi-head attention has one KV head per query head; multi-query attention has one KV head shared by all query heads. See the original MQA and GQA papers.
Weight memory
M_weights,payload = P_total × bits_per_weight / 8This is a payload floor, not allocated memory. Quantized checkpoints also carry scales, zero points, group metadata, packing alignment and sometimes higher-precision embeddings or output heads. A 4-bit format is not automatically 0.5 bytes per parameter end to end. Sum actual tensor byte sizes, inspect the quantization config, then measure memory after load.
Llama 3.1 70B, BF16
Using 70.6 billion as a rounded nominal count, BF16 payload arithmetic is 141.2 GB = 131.50 GiB; this is not the checkpoint’s exact tensor byte count. Two “80 GB” nameplates total 160 GB, or 149.01 GiB by conversion, but driver-visible capacity may differ—and the service still needs KV and workspace. TP=2 is not a credible production fit even before fragmentation.
MoE: two parameter counts, two constraints
A mixture-of-experts model needs total parameters for residency but only a routed subset for each token’s expert MLP work. Mixtral 8×7B’s config declares eight local experts and selects two per token; the official model card reports 46.7B total and 12.9B active parameters. BF16 weight payload is 93.4 GB = 86.99 GiB, even though per-token linear FLOPs resemble a smaller dense model.
Do not turn “active parameters” directly into decode bytes. Across a batch, the union of selected experts can approach all experts; hot experts may be reused while cold experts create uneven traffic. Capacity padding, routing imbalance, shared experts and all-to-all exchange add work not represented by Pactive.
Phase 2 · Account for state
The KV cache is architecture × retained tokens
For conventional MHA, GQA or MQA with uniform full attention, logical cache bytes are:
M_KV = B × S × L × 2 × H_kv × D × b_kvB is live sequences, S retained tokens per sequence, two represents key and value, and bkv is bytes per cached element. This excludes page rounding, allocator metadata, alignment and temporary attention workspace.
Llama 3.1 70B at 32-way concurrency
BF16 cache per token is 2 × 80 × 8 × 128 × 2 = 327,680 bytes = 320 KiB. If each sequence retains 4,096 input plus 512 output tokens, one consumes 1.40625 GiB and 32 consume exactly 45 GiB cluster-wide.
With ideal KV-head sharding, TP=2 carries 22.5 GiB of KV and 65.75 GiB of weights per rank: 88.25 GiB before runtime memory, so it cannot fit an 80 GB-class device. TP=4 reduces that logical sum to 44.13 GiB per rank. That is feasibility—not a latency prediction.
Sliding, hybrid and paged caches
M_KV = B × Σ_layer [S_layer × 2 × H_kv,layer × D_layer × b_kv]For a sliding-window layer, Slayer = min(S, W); full-attention layers retain the full sequence. Hybrid models can have recurrent or state-space layers with fixed per-sequence state. Runtime managers may group layers into common page shapes and pad to the largest group. “Context window × bytes/token” can overstate an ideal sliding cache yet understate real padded allocation.
Paged allocation maps logical blocks to non-contiguous physical blocks; it does not make KV free. The PagedAttention paper reports low waste for its evaluated allocator, while current runtimes add prefix caching, offload and hybrid policies with their own metadata and eviction behavior.
MLA is not “GQA with fewer heads”
Multi-head latent attention stores a compressed joint KV latent plus positional key state, with exact contents dependent on the implementation. The public DeepSeek-V3 config describes width 7,168; 61 layers; 128 attention heads; non-positional Q/K head dimension 128; RoPE head dimension 64; value head dimension 128; and KV latent rank 512. Its first three layers are dense; later MoE layers expose 256 routed experts, select eight, and add one shared expert. A common absorbed MLA representation has this BF16 floor:
bytes/token = 61 × (512 + 64) × 2 = 70,272 bytesDeepSeek-V3, 64 live 16K-token sequences
The latent-plus-RoPE cache is 68.625 KiB/token. Exactly 1,048,576 retained tokens consume 68.625 GiB logically. Plan conservatively for 68.625 GiB on every rank when the engine replicates latent state. An evenly sequence-sharded eight-way cache would be a purely hypothetical 8.578 GiB/rank lower bound; do not use it unless the exact engine version and layout prove that mapping. The technical report gives 671B total and 37B activated parameters: FP8 payload is roughly 671 GB before metadata, while linear compute follows the active path plus attention and routing.
Prefix caching changes compute demand, not the existence of cached bytes. Report token hit rate, eviction rate and prefix-length distribution; request hit rate alone is misleading.
Phase 3 · Bound the work
Prefill and decode live on different rooflines
For a dense decoder-only transformer, a useful first-order linear-layer cost is 2 × active parameters × tokens, counting a multiply-add as two FLOPs. A naive full-attention accounting adds an approximately quadratic term:
F_prefill ≈ 2 × P_active × S + 4 × L × S² × HThe second term covers QK scores and applying attention probabilities to V; it is not a universal model-FLOPs estimator. Causal kernels avoid work on masked positions, FlashAttention changes memory traffic, and sliding/hybrid attention, MoE routing, embeddings, norms and logits add or remove work. For Llama 3.1 70B at 8,192 tokens, this deliberately conservative arithmetic gives a 1,156.71 TFLOP linear term and 175.92 TFLOP attention term—13.2% of the simplified total.
Decode processes one new token per live sequence, reads retained KV and traverses the model. The right bound is a roofline, not a fixed efficiency percentage:
t_step ≥ max(F_step / C_effective, D_HBM / BW_effective, t_collectives)At batch one, a dense decode step may be close to streaming weight bytes once, so BW / weight bytes is a useful upper bound on per-request tokens/s. At batch B, a fused matrix multiply can reuse a weight tile across sequences. In the ideal weight-only limit, aggregate tokens/s grows toward B × BW / weight bytes, until KV reads, compute, scheduling or communication bind. “Every output token loads all weights” without a batch qualifier is wrong.
The original Roofline model frames this as FLOPs per byte. Batching moves decode rightward by amortizing weight traffic; long-context KV reads can pull it back toward bandwidth pressure. Effective compute and bandwidth must be measured for the kernel, dtype, batch and topology—brochure peaks are ceilings.
From service demand to replicas
input_tokens/s = λ × E[S_in] output_tokens/s = λ × E[S_out]Divide demand by measured sustainable throughput only after filtering benchmark points that meet the latency SLO. The fastest saturated run is irrelevant if p99 TTFT is 12 seconds and the target is 800 ms. Tie headroom to failure domains, uncertainty and rolling updates—not a universal constant.
Demand, concurrency and replicas
Assume 6 requests/s, mean 1,200 input tokens and 300 output tokens. Demand is 7,200 input tok/s and 1,800 output tok/s. Suppose—not predict—that the same mixed trace on one candidate replica sustains 12,000 input tok/s and 900 output tok/s while meeting every latency percentile. The planning ratios are ⌈7,200/12,000⌉ = 1 and ⌈1,800/900⌉ = 2; their maximum makes two replicas only a lower-bound candidate. Independent peak prefill and decode runs cannot be combined this way.
If measured mean TTFT is 0.600 s and mean ITL is 25 ms on that trace, mean E2E is 0.600 + (300−1) × 0.025 = 8.075 s. Little’s Law, using means in steady state, gives 6 × 8.075 = 48.45 average in-flight requests. The two-replica service must still pass the mixed trace. If it does, an illustrative N+1 policy would make three replicas the conditional deployment floor; none of these means protects p99 by itself.
Phase 4 · Map onto hardware
Parallelism changes both bytes and wires
Tensor parallelism is attractive inside a high-bandwidth NVLink/NVSwitch island. It also inserts collectives into the critical path. KV sharding is not guaranteed to scale with TP: if KV heads divide across ranks it often shards cleanly; with MQA or TP larger than the KV-head count, implementations may replicate heads or form replication groups. Decode-context parallelism can shard cache along sequence, at extra communication cost.
Pipeline parallelism places layers on stages, splitting layer-owned weights and KV but adding traversal and bubbles. Data parallelism provides independent replicas and is usually the cleanest throughput scale once one replica fits. Expert parallelism distributes expert weights but leaves attention, shared experts and dense components replicated or sharded by another axis; all-to-all traffic makes topology part of the model. The Megatron-LM paper derives intra-layer TP.
Form factor is topology
An H100 SXM module in an HGX baseboard is not interchangeable with an H100 PCIe card. NVIDIA lists up to 3.35 TB/s HBM and 900 GB/s total NVLink bandwidth for H100 SXM, with different power and interconnect characteristics for PCIe, on its H100 specification page. These are vendor peak conventions, not application throughput.
Never compare aggregate bidirectional NVLink to one-way NIC payload bandwidth. For inter-node TP or EP, measure collective bus bandwidth and tail latency across exact rail count, oversubscription, GPU-NIC affinity and message sizes. Nominal link rate is not NCCL algorithmic bandwidth.
Phase 5 · Give the runtime room
The engine owns memory the model sheet cannot see
M_device ≥ M_weights + M_KV + M_graphs + M_activations + M_collectives + M_allocator + M_transient + marginChunked prefill caps tokens admitted to a prefill iteration, limiting activation spikes and reducing how badly a long prompt blocks decode. It can improve ITL fairness while increasing a long prompt’s TTFT. Continuous batching packs work dynamically; the batch is a changing token budget, not a static request count. CUDA graphs reduce launch overhead but reserve capture buffers. Paged caches reduce fragmentation but block rounding and copy-on-write remain.
vLLM and SGLang are policies, not interchangeable labels
This article checked the release pages for vLLM v0.30.0 and SGLang v0.5.21, plus each project’s latest documentation, on 04 October 2026. “Latest” docs can move; pin the package or image digest and archive resolved arguments with every result.
Current engine docs describe gpu_memory_utilization as a per-instance model-executor limit with default 0.92, and allow an explicit kv_cache_memory_bytes. max_num_batched_tokens bounds tokens per iteration; max_num_seqs sizes runner and graph capacity; chunked prefill uses remaining token budget. The docs say several scheduler defaults are testing conveniences.
mem_fraction_static covers weights plus KV pool, while activations and graph buffers sit outside it. Current docs compute it from detected reserved memory, falling back to 0.88—not a timeless constant. chunked_prefill_size, max_running_requests and per-phase graph capture change the OOM/latency frontier. RadixAttention prefix caching is enabled unless disabled.
Prefix caching helps only when prefixes repeat and remain resident. Paged allocation helps when block policy matches workload. Graphs help captured shapes and increase reserved memory. For both engines, record attention backend, quantization kernels, KV dtype, page size, maximum length, scheduler limits, speculative decoding, prefix policy, parallel axes and graph shapes.
A single “framework overhead” percentage hides resources with different scaling laws. Replace it with measured fixed reservations, graph memory, peak prefill activations, collective workspace, allocator slack and an explicit operational margin.
Phase 6 · Prove the envelope
Benchmark the service you intend to operate
An analytic sheet should output constraints and uncertainty, not “predicted latency.” Use it to choose candidate topologies. Then run the same model revision, tokenizer, quantization, engine build and GPU firmware that production will use.
Workload: use example D’s mixed trace—6 RPS, mean 1,200 input and 300 output tokens—with explicit p99 TTFT, ITL and E2E thresholds.
Memory floor: example B rejects BF16 Llama 3.1 70B at ideal TP=2. Ideal TP=4 has a 44.13 GiB/rank logical weight-plus-KV floor for 32 live 4,608-token sequences, leaving capacity to be measured for runtime state.
Candidate topology: begin with one four-GPU TP group inside a single high-bandwidth node. Do not cross a slower inter-node fabric until measurement shows a reason.
Acceptance: replay the mixed trace against one, then two groups. Two four-GPU replicas become the steady-state candidate only if the eight-GPU service sustains offered load while every latency percentile, memory peak and queue bound passes. A conditional N+1 policy would test three groups—12 GPUs—under one-group failure before approval.
- Freeze identity.Checkpoint commit and tensors; engine/container digest; driver, CUDA, kernels, NCCL; GPU SKU, form factor, clocks and topology.
- Replay distributions.Joint input/output lengths, arrivals, prefix reuse, priorities, cancellations and multi-turn think time. Include cold and warm cache.
- Sweep admission.Offered RPS, running requests, token budget, cache fraction, chunk size and parallel layout. Hold seed and dataset constant.
- Observe every stage.Queue, TTFT, ITL/TPOT and E2E at p50/p90/p95/p99; input/output tok/s; KV occupancy, evictions, preemptions; HBM, SM, network and power.
- Find sustainable capacity.Highest load meeting every SLO without an unbounded queue, OOM, thermal drift or error-rate regression during a long soak.
- Exercise failure.Rolling update, one replica unavailable, cache cold start, burst, long-prompt outliers and network degradation.
Report confidence intervals and run-to-run variance. Separate measurements from interpolation; label multi-node extrapolation. The vLLM benchmark suite provides harnesses, but not your traffic distribution or acceptance criteria.
A final sizing checklist
The useful GPU count is the smallest tested topology whose worst relevant constraint still clears the service objective. Everything before that sentence is disciplined elimination. Everything after it is operations.
Source and version notes
What was checked
Primary sources accessed 04 October 2026. Model configs on “main” are mutable unless a commit is shown; pin them before reproduction.
- Meta Llama 3.1 70B config — dense GQA dimensions.
- Mixtral 8×7B config, commit 125c431 — expert and attention dimensions.
- DeepSeek-V3 config, commit e815299 and report — 61-layer MLA and MoE values.
- NVIDIA H100 specifications — HBM, form factor and interconnect ceilings.
- vLLM engine arguments, prefix caching, and benchmarking.
- SGLang server arguments and tuning guide.
- PagedAttention, GQA, MQA, Megatron-LM, and Roofline.
Method note: examples are arithmetic demonstrations, not GPU benchmarks. No customer data or proprietary calculator code was used in external research. Values labeled nominal, logical, floor or ceiling require measurement before a performance or purchasing decision.