# OpenAI Realtime API The Realtime API is [[OpenAI]]'s API for live audio. Instead of uploading a file and waiting for a result, you keep a connection open and stream audio in and out. Browsers connect over WebRTC, servers over [[WebSockets]], and phone systems over SIP. Why a separate API? A voice conversation can't wait for a request/response round trip per sentence. The session has to handle turn-taking, interruptions, tool calls and audio playback while the user keeps talking. ## Session types and models Since May 2026, the Realtime API has three kinds of sessions, each with its own model: | Session | Endpoint | Model | Price | |---|---|---|---| | Speech-to-speech voice agent | `v1/realtime` | [[GPT-Realtime-2]], `gpt-realtime-2.1`, `gpt-realtime-2.1-mini` | Per token (audio $32 / $64 per 1M for 2.x; $10 / $20 for 2.1 mini) | | Live translation | `v1/realtime/translations` | [[GPT-Realtime-Translate]] | $0.034 per minute | | Live transcription | `v1/realtime/transcription_sessions` | [[GPT-Realtime-Whisper]], `gpt-live-transcribe` | $0.017 per minute | Older realtime models (`gpt-realtime`, `gpt-4o-realtime`, `gpt-realtime-mini`) are scheduled for removal on 20 January 2027. Voice agents in the Realtime API support function tools and [[Model Context Protocol (MCP)|MCP]] tools, voice activity detection, and server-side controls. The [[OpenAI Agents SDK]] adds guardrails on top. The API supports EU data residency, and OpenAI runs classifiers that can halt sessions violating its policies. ## Realtime API or GPT-Live? In September 2026, OpenAI made GPT-Live 1 generally available through a separate `v1/live/sessions` endpoint. GPT-Live is full-duplex: it can listen while speaking, and it hands reasoning and tools off to a backend model while it keeps the conversation going. OpenAI's audio guide now recommends GPT-Live for new conversational voice apps, and the Realtime API when you need its own session and tool model, or dedicated translation and transcription sessions. ## References - Realtime API guide: https://developers.openai.com/api/docs/guides/realtime - Audio and voice guide: https://developers.openai.com/api/docs/guides/audio - Realtime translation guide: https://developers.openai.com/api/docs/guides/realtime-translation - Realtime transcription guide: https://developers.openai.com/api/docs/guides/realtime-transcription - Pricing: https://developers.openai.com/api/docs/pricing - Deprecations: https://developers.openai.com/api/docs/deprecations - Changelog: https://developers.openai.com/api/docs/changelog - OpenAI, "Advancing voice intelligence with new models in the API": https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/ ## Related - [[GPT-Realtime-2]] - [[GPT-Realtime-Translate]] - [[GPT-Realtime-Whisper]] - [[2026-05-07 GPT-Realtime-2 - OpenAI's voice model gets GPT-5-class reasoning]] - [[OpenAI SDK]] - [[OpenAI Agents SDK]] - [[OpenAI]] - [[NVIDIA VoiceChat-11B]]