What changed
Artificial Analysis ran its 100K context benchmark on the Gemma 4 31B model on Groq 3 LPX and measured 3,431 output tokens/second, a figure NVIDIA’s post reported on 24 August 2026 1. The same day The Register worked out what that model occupies on the hardware: a little over 31 GB, or just under 64 LPUs of SRAM capacity, with NVIDIA running it at FP8 2. NVIDIA’s post pairs Groq 3 LPX with Vera Rubin NVL72 to serve multiagent systems powered by 2T+ parameter models 1.
What one rack actually holds
At 31 billion parameters, The Register’s own arithmetic puts the model inside a single LPX rack regardless of the data type used to store the weights 2. What decides the fit at this size, I think, is the rack’s SRAM budget, and precision starts to matter for sizing only at the boundary where a model stops fitting a rack. The concrete quantity behind that is just under 64 LPUs of SRAM capacity at FP8 2. The table below sets the figures side by side.
| Figure | Value | Where it originates |
|---|---|---|
| Interactivity, 100K context benchmark | 3,431 output tokens/second | NVIDIA’s post, reporting Artificial Analysis’ run 1 |
| Benchmarked model | Gemma 4 31B | NVIDIA’s post 1 |
| Weights at FP8 | a little over 31 GB | The Register’s own derivation 2 |
| SRAM capacity needed | just under 64 LPUs | The Register’s own derivation 2 |
| Footprint | a single LPX rack, regardless of data type | The Register’s own derivation 2 |
Impact on your team
This lands on anyone picking an inference provider for long-context work, and on anyone sizing a rack around a model they already run. In my experience the useful move is to compute your own model’s weight footprint at your own precision and check it against the accelerator’s SRAM budget before you quote anybody’s tokens-per-second, because on this hardware capacity is bought in accelerator count. What that changes is the shortlist: a provider that cannot hold your model in one rack sits in a different bracket from one that is merely slower. Sizing comes before speed.