What changed
OpenAI repriced the GPT-5.6 family on July 30 2026: Luna costs 80% less, Terra 20% less, and Sol did not move 15. The change lands on v1/responses and v1/chat/completions 1. Priority Processing was renamed Fast mode the same day, and requests still sending service_tier: "priority" keep working 13.
The new floor
| Model | Before | After |
|---|---|---|
| gpt-5.6-sol | $5.00 / $30.00 | unchanged 25 |
| gpt-5.6-terra | $2.50 / $15.00 4 | $2.00 / $12.00 2 |
| gpt-5.6-luna | $1.00 / $6.00 4 | $0.20 / $1.20 2 |
(input / output per 1M tokens, short context.)
A floor that falls fivefold barely discounts the invoice you already pay. It changes what you can afford to run at all, which I think is the real story here: per-chunk reranking on every retrieval stops being an experiment you have to justify. Luna sits 3.75x under the last generation’s mid tier, gpt-5.4-mini at $0.75 / $4.50 2, on both sides of the meter.
The fan-out arithmetic flips sign. On input, Luna cost $1.00 per Mtok 4 against $2.50 for Terra 4, so five Luna calls came to five times $1.00, twice the price of one Terra call. At $0.20 2 against $2.00 2, five times $0.20 is half that single call. Fanning out was the expensive option; it is now the cheap one.
The objection is fair: an 80% cut on a tier your task cannot use saves nothing, and a cascade still costs engineering time plus an eval harness. Neither point survives the sign flip above, which changes which experiments fit the budget rather than trimming the calls you already make.
Fast mode is a flat 2x
Fast mode bills exactly twice standard on input and output across all three 5.6 tiers, luna at $0.40 / $2.40 and sol at $10.00 / $60.00 2. The 2.5x speedup is tied to one named model, gpt-5.6-sol 13, though the guide’s opening line promises up to 2.5x faster without naming a tier 3. Read that as where the number is anchored rather than as a complaint about silence: the doubling is certain everywhere, the measurement is not, so on Terra and Luna I measure before flipping the flag. Two traps sit underneath. The response object reports priority for GPT-5.6 and earlier even when the request said fast 3, so telemetry keyed on "fast" never fires; long context, fine-tuned models, and embeddings are out of scope entirely 3.
Impact on your team
Re-run tier selection this week instead of banking the cut on paper: anything parked on Terra purely for headroom now pays ten times Luna’s input price, and the saving lands only if the router moves. Audit service_tier at the project level before the next invoice, because a project default of Fast doubles every unlabelled request while the response object still says priority, so the dashboard will not tell you 3. What I would ignore is Fast mode as a general latency fix outside Sol, where the 2x is certain and the 2.5x is anchored to another model 13.