# OpenRouter adds unified audio transcription
[[OpenRouter]] added speech-to-text to its API. A new endpoint (`POST /api/v1/audio/transcriptions`) accepts base64-encoded audio and returns the transcript as JSON, using the same API key as chat completions, so you don't need a separate STT provider or server.
Two model families are available: [[Whisper]]-class models priced per audio second, and newer STT models priced per token (discoverable with the `?output_modalities=transcription` filter). Pricing is the model's catalog rate with no OpenRouter markup, and every response includes a usage object with duration, token counts, and the actual cost of that request. When a model is hosted by several providers, requests load-balance automatically.
Current limits: a 60-second processing timeout (a throughput constraint; audio length itself isn't capped), no URL-based audio input (you upload the bytes), and no built-in SRT/VTT subtitle output.
For anyone already routing [[Large Language Models (LLMs)]] traffic through OpenRouter, this makes meeting transcripts, voice notes, and podcast-to-text pipelines a one-endpoint addition instead of a second vendor relationship.
## References
- Announcement: https://openrouter.ai/blog/tutorials/transcription-on-openrouter/
## Related
- [[OpenRouter]]
- [[Whisper]]