# OpenRouter adds unified audio transcription [[OpenRouter]] added speech-to-text to its API. A new endpoint (`POST /api/v1/audio/transcriptions`) accepts base64-encoded audio and returns the transcript as JSON, using the same API key as chat completions, so you don't need a separate STT provider or server. Two model families are available: [[Whisper]]-class models priced per audio second, and newer STT models priced per token (discoverable with the `?output_modalities=transcription` filter). Pricing is the model's catalog rate with no OpenRouter markup, and every response includes a usage object with duration, token counts, and the actual cost of that request. When a model is hosted by several providers, requests load-balance automatically. Current limits: a 60-second processing timeout (a throughput constraint; audio length itself isn't capped), no URL-based audio input (you upload the bytes), and no built-in SRT/VTT subtitle output. For anyone already routing [[Large Language Models (LLMs)]] traffic through OpenRouter, this makes meeting transcripts, voice notes, and podcast-to-text pipelines a one-endpoint addition instead of a second vendor relationship. ## References - Announcement: https://openrouter.ai/blog/tutorials/transcription-on-openrouter/ ## Related - [[OpenRouter]] - [[Whisper]]