What changed
OpenAI released GPT Transcribe and GPT Live Transcribe on July 28 1. The file model bills $0.0045 per minute and the streaming one $0.017 per minute 2; Microsoft listed both in Foundry the next day 7. The rewritten transcription guide now routes each workflow to one recommended starting model 6.
Three things move at once
Reading this as a version bump is fair on the surface: same vendor, same task, a newer id on the same call. It still breaks: the hint field renames, the endpoints split 34, and the outputs you consume move elsewhere.
| Capability | gpt-transcribe | gpt-live-transcribe |
|---|---|---|
| Price per minute | $0.0045 2 | $0.017 2 |
| File endpoint | v1/audio/transcriptions 3 | not supported 4 |
| Realtime endpoint | v1/realtime/transcription_sessions 3 | v1/realtime/transcription_sessions 4 |
| Word timestamps, speaker labels | routed to whisper-1 or gpt-4o-transcribe-diarize 6 | none returned 5 |
| Language hint field | languages 6 | languages 5 |
The rename is cheap and fails at runtime rather than at review time.
The expensive break sits downstream. Anything consuming word timestamps, subtitles, or speaker labels stays on whisper-1 or gpt-4o-transcribe-diarize 6, because gpt-live-transcribe returns none of them 5. There the upgrade means rewriting the consumer. I would adopt on the file path where the product only needs text, and keep the old model wherever the transcript feeds a timeline.
A streaming session carries the context hints in one object 5:
"transcription": {
"model": "gpt-live-transcribe",
"prompt": "A customer support call about a premium plan and account AC-42.",
"keywords": ["premium plan", "AC-42", "billing"],
"languages": ["en", "fr"],
"delay": "low"
}
Where the price cut actually lands
The quarter-below-its-predecessor headline 8 describes one of the two models: $0.0045 against $0.006 for gpt-4o-transcribe 2. gpt-4o-mini-transcribe still bills $0.003 2, so the recommended path costs half again as much per minute as the tier OpenAI does not recommend for new integrations 6. And gpt-realtime-whisper bills the same $0.017 2: the streaming meter is not new. Artificial Analysis measured a 3.31% word error rate, ninth among the roughly 50 systems it tracks 8. Size your accuracy expectations to that rank.
Impact on your team
There is no deadline, and that is the useful part: the 4o models keep working for existing integrations 6, so read “recommended starting model” as routing advice for new work. Before any model id changes, grep the transcription call sites for language: and for every consumer of word timestamps or speaker labels; those two greps decide whether this is a one-line change or a sprint. Keep gpt-4o-mini-transcribe for bulk batch where $0.003 beats $0.0045 2. If you open streaming sessions, benchmark delay on your own telephony audio: the docs name starting points from minimal to xhigh and publish no milliseconds 5.