Pick a model, a bit-width, and a machine. This calculator applies the tiered decode law — fitted and validated on measurements from 7B to 753B, including two pre-registered hardware tests — to predict decode speed, memory fit, and quality cost before you download a single weight.
--no-mmap + -ot exps=CPU). Estimates ±25% off-VRAM; all-in-VRAM is one-sided — measured ≥0.90× the prediction every time, typically 1.1–1.8× faster. v1.1 context term: each generated token re-reads the whole KV cache from the tier KV lives on (ηkv≈0.7, single-point calibration: measured 20.02→16.12 tok/s at depth 16384; term prompted by u/RogerAI--fyi). MLA models cache ~10× less; SWA long-context slopes use global layers only [est].
pip install quantprobe && quantprobe calibrate measures your machine (RAM stream, disk, GPU sustained clocks) and anchors predictions by default (pre-registered leave-one-out gate #64: median error 19% → 5.8% across 5 arms; --no-anchors disables).