# GPT-Realtime-2
GPT-Realtime-2 is [[OpenAI]]'s speech-to-speech model for voice agents, released on 7 May 2026 in the [[OpenAI Realtime API]]. Audio goes in, audio comes out, with no separate speech-to-text and text-to-speech steps in between. OpenAI calls it its first voice model with "[[GPT-5]]-class reasoning".
The reason it exists: voice agents used to be fast talkers with weak brains. GPT-Realtime-2 can reason before answering, call several tools at once, and keep talking while it works ("let me check that"). That's what turns a voice bot that answers questions into one that can finish a task, like rebooking a trip or scheduling a house visit.
It shipped together with [[GPT-Realtime-Translate]] and [[GPT-Realtime-Whisper]]. Release write-up: [[2026-05-07 GPT-Realtime-2 - OpenAI's voice model gets GPT-5-class reasoning]].
## Specifications and pricing
| | GPT-Realtime-2 |
|---|---|
| Model ID | `gpt-realtime-2` |
| Endpoint | `v1/realtime` only (WebRTC, WebSocket or SIP) |
| Input / output | Text, audio, image in; text, audio out |
| [[Context Window\|Context window]] | 128,000 tokens (GPT-Realtime-1.5: 32,000) |
| Max output | 32,000 tokens (GPT-Realtime-1.5: 4,096) |
| Knowledge cutoff | 30 September 2024 |
| Reasoning effort | `minimal`, `low` (default), `medium`, `high`, `xhigh` |
| Audio tokens | $32 in, $0.40 cached, $64 out per 1M |
| Text tokens | $4 in, $0.40 cached, $24 out per 1M (1.5: $16 out) |
| Image tokens | $5 in, $0.50 cached per 1M |
| Features | Function calling, prompt caching |
## What's new compared with GPT-Realtime-1.5
- **Preambles**: optional short phrases ("one moment while I look into it") before the real answer
- **Parallel tool calls**, which the model can announce out loud
- **Recovery**: it says it's having trouble instead of failing silently
- **Specialized vocabulary**: better at proper nouns, domain terms and healthcare words
- **Tone control**: calm, empathetic or upbeat depending on the situation
Higher reasoning effort means more latency and more output tokens. Low is the default for a reason: in a live conversation, a long silence feels broken.
## Benchmarks
OpenAI's numbers against GPT-Realtime-1.5 (values from OpenAI's chart as reproduced by The Decoder):
- **Big Bench Audio**, at `high`: 96.6% vs 81.4%
- **Audio MultiChallenge**, at `xhigh`: 48.5% vs 34.7%
OpenAI's post calls these "15.2%" and "13.8%" higher; they are percentage-point gains. Zillow, an early tester, reported 95% call success on its hardest adversarial benchmark, against 69% before (after prompt optimization).
## GPT-Realtime-2.1 and 2.1 mini
On 6 July 2026, OpenAI released two follow-ups:
- **`gpt-realtime-2.1`**: better alphanumeric recognition (order numbers, codes), silence and noise handling, and interruption behavior. Same specs and same price as GPT-Realtime-2
- **`gpt-realtime-2.1-mini`**: a distilled, faster and cheaper reasoning model. Audio: $10 in, $0.30 cached, $20 out per 1M tokens; text: $0.60 / $2.40
OpenAI also said p95 latency dropped by at least 25% across its Realtime voice models thanks to better caching. Early forum feedback was mixed: one developer saw no real difference between 2 and 2.1, another reported that 2.1 mini stopped calling function tools in a SIP flow that worked with the older mini.
OpenAI's deprecations page names the 2.1 models as the replacements for older ones: `gpt-realtime` and `gpt-4o-realtime` go to `gpt-realtime-2.1` on 20 January 2027; `gpt-realtime-mini`, `tts-1`, `tts-1-hd` and the `gpt-4o-mini-tts` snapshots go to `gpt-realtime-2.1-mini`.
## Where it sits now
Since 10 September 2026, OpenAI's audio guide tells you to start new conversational voice apps with GPT-Live instead, a full-duplex model that delegates reasoning and tools to a backend model ($0.05 per minute plus backend usage). GPT-Realtime-2.x stays the choice when you want the Realtime API's own session and tool model.
My take: if you're already on the Realtime API, move to `gpt-realtime-2.1` (it costs the same) and re-test your tools before switching. For a new voice product, evaluate GPT-Live first.
## References
- OpenAI, "Advancing voice intelligence with new models in the API": https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/
- Model docs: https://developers.openai.com/api/docs/models/gpt-realtime-2, https://developers.openai.com/api/docs/models/gpt-realtime-2.1, https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini, https://developers.openai.com/api/docs/models/gpt-realtime-1.5
- Pricing: https://developers.openai.com/api/docs/pricing
- Changelog: https://developers.openai.com/api/docs/changelog
- Deprecations: https://developers.openai.com/api/docs/deprecations
- Audio guide (GPT-Live recommendation): https://developers.openai.com/api/docs/guides/audio
- OpenAI Developer Community, gpt-realtime-2.1 announcement thread: https://community.openai.com/t/new-realtime-models-on-the-api-gpt-realtime-2-1-and-gpt-realtime-2-1-mini/1385896
- The Decoder (benchmark values): https://the-decoder.com/openais-new-voice-model-brings-gpt-5-level-reasoning-to-real-time-conversations/
- Hacker News discussion: https://news.ycombinator.com/item?id=48051991
## Related
- [[OpenAI Realtime API]]
- [[GPT-Realtime-Translate]]
- [[GPT-Realtime-Whisper]]
- [[2026-05-07 GPT-Realtime-2 - OpenAI's voice model gets GPT-5-class reasoning]]
- [[OpenAI]]
- [[Text-to-Speech (TTS)]]
- [[AI Agents]]
- [[NVIDIA VoiceChat-11B]]