# GPT-Realtime-2 GPT-Realtime-2 is [[OpenAI]]'s speech-to-speech model for voice agents, released on 7 May 2026 in the [[OpenAI Realtime API]]. Audio goes in, audio comes out, with no separate speech-to-text and text-to-speech steps in between. OpenAI calls it its first voice model with "[[GPT-5]]-class reasoning". The reason it exists: voice agents used to be fast talkers with weak brains. GPT-Realtime-2 can reason before answering, call several tools at once, and keep talking while it works ("let me check that"). That's what turns a voice bot that answers questions into one that can finish a task, like rebooking a trip or scheduling a house visit. It shipped together with [[GPT-Realtime-Translate]] and [[GPT-Realtime-Whisper]]. Release write-up: [[2026-05-07 GPT-Realtime-2 - OpenAI's voice model gets GPT-5-class reasoning]]. ## Specifications and pricing | | GPT-Realtime-2 | |---|---| | Model ID | `gpt-realtime-2` | | Endpoint | `v1/realtime` only (WebRTC, WebSocket or SIP) | | Input / output | Text, audio, image in; text, audio out | | [[Context Window\|Context window]] | 128,000 tokens (GPT-Realtime-1.5: 32,000) | | Max output | 32,000 tokens (GPT-Realtime-1.5: 4,096) | | Knowledge cutoff | 30 September 2024 | | Reasoning effort | `minimal`, `low` (default), `medium`, `high`, `xhigh` | | Audio tokens | $32 in, $0.40 cached, $64 out per 1M | | Text tokens | $4 in, $0.40 cached, $24 out per 1M (1.5: $16 out) | | Image tokens | $5 in, $0.50 cached per 1M | | Features | Function calling, prompt caching | ## What's new compared with GPT-Realtime-1.5 - **Preambles**: optional short phrases ("one moment while I look into it") before the real answer - **Parallel tool calls**, which the model can announce out loud - **Recovery**: it says it's having trouble instead of failing silently - **Specialized vocabulary**: better at proper nouns, domain terms and healthcare words - **Tone control**: calm, empathetic or upbeat depending on the situation Higher reasoning effort means more latency and more output tokens. Low is the default for a reason: in a live conversation, a long silence feels broken. ## Benchmarks OpenAI's numbers against GPT-Realtime-1.5 (values from OpenAI's chart as reproduced by The Decoder): - **Big Bench Audio**, at `high`: 96.6% vs 81.4% - **Audio MultiChallenge**, at `xhigh`: 48.5% vs 34.7% OpenAI's post calls these "15.2%" and "13.8%" higher; they are percentage-point gains. Zillow, an early tester, reported 95% call success on its hardest adversarial benchmark, against 69% before (after prompt optimization). ## GPT-Realtime-2.1 and 2.1 mini On 6 July 2026, OpenAI released two follow-ups: - **`gpt-realtime-2.1`**: better alphanumeric recognition (order numbers, codes), silence and noise handling, and interruption behavior. Same specs and same price as GPT-Realtime-2 - **`gpt-realtime-2.1-mini`**: a distilled, faster and cheaper reasoning model. Audio: $10 in, $0.30 cached, $20 out per 1M tokens; text: $0.60 / $2.40 OpenAI also said p95 latency dropped by at least 25% across its Realtime voice models thanks to better caching. Early forum feedback was mixed: one developer saw no real difference between 2 and 2.1, another reported that 2.1 mini stopped calling function tools in a SIP flow that worked with the older mini. OpenAI's deprecations page names the 2.1 models as the replacements for older ones: `gpt-realtime` and `gpt-4o-realtime` go to `gpt-realtime-2.1` on 20 January 2027; `gpt-realtime-mini`, `tts-1`, `tts-1-hd` and the `gpt-4o-mini-tts` snapshots go to `gpt-realtime-2.1-mini`. ## Where it sits now Since 10 September 2026, OpenAI's audio guide tells you to start new conversational voice apps with GPT-Live instead, a full-duplex model that delegates reasoning and tools to a backend model ($0.05 per minute plus backend usage). GPT-Realtime-2.x stays the choice when you want the Realtime API's own session and tool model. My take: if you're already on the Realtime API, move to `gpt-realtime-2.1` (it costs the same) and re-test your tools before switching. For a new voice product, evaluate GPT-Live first. ## References - OpenAI, "Advancing voice intelligence with new models in the API": https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/ - Model docs: https://developers.openai.com/api/docs/models/gpt-realtime-2, https://developers.openai.com/api/docs/models/gpt-realtime-2.1, https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini, https://developers.openai.com/api/docs/models/gpt-realtime-1.5 - Pricing: https://developers.openai.com/api/docs/pricing - Changelog: https://developers.openai.com/api/docs/changelog - Deprecations: https://developers.openai.com/api/docs/deprecations - Audio guide (GPT-Live recommendation): https://developers.openai.com/api/docs/guides/audio - OpenAI Developer Community, gpt-realtime-2.1 announcement thread: https://community.openai.com/t/new-realtime-models-on-the-api-gpt-realtime-2-1-and-gpt-realtime-2-1-mini/1385896 - The Decoder (benchmark values): https://the-decoder.com/openais-new-voice-model-brings-gpt-5-level-reasoning-to-real-time-conversations/ - Hacker News discussion: https://news.ycombinator.com/item?id=48051991 ## Related - [[OpenAI Realtime API]] - [[GPT-Realtime-Translate]] - [[GPT-Realtime-Whisper]] - [[2026-05-07 GPT-Realtime-2 - OpenAI's voice model gets GPT-5-class reasoning]] - [[OpenAI]] - [[Text-to-Speech (TTS)]] - [[AI Agents]] - [[NVIDIA VoiceChat-11B]]