# VoiceStudio VoiceStudio is an open-source desktop app that calls itself "the fully-local [[ElevenLabs]] alternative". It does voice cloning, voice design (describe a voice in words, get a voice), video dubbing, dictation, transcription, stories and audiobooks, in 646 languages. Everything runs on your own hardware. Remote services are optional, and usage analytics require your consent. It's maintained by debpalash on GitHub. The repo was created on 2026-04-09 and reached roughly 53k stars by early October 2026, with commits landing daily. That's a LOT of attention for a six-month-old project. ## What it does - Voice cloning: pick a voice or add a clean reference recording, type your text, generate - Voice design: create a new voice from a text description instead of a sample - Video dubbing with timed speech over existing video - A floating dictation widget for speech-to-text anywhere on the desktop - Transcription (see [[Speech-to-Text (STT)]]) - Long-form generation for stories, audiobooks and batch jobs (the same need [[ebook2audiobook]] covers) - A local API and a [[Model Context Protocol (MCP)]] server, so AI agents can drive it. It also ships agent skills (`npx skills add debpalash/VoiceStudio`) and an install prompt you can paste into [[Claude Code]] or another coding agent - Optional remote workers to offload generation to another machine (e.g., an Intel Mac using a GPU box as backend) The default engine is k2-fsa's OmniVoice. You can switch engines and manage local models from the app. OmniVoice is also one of the open models [[Supertonic]] benchmarks itself against. ## Platforms and hardware - macOS (`.dmg`), Windows (`.exe`), Linux (`.AppImage`, `.deb`) and [[Docker]] - One-line installer: `curl -fsSL https://voicestudio.sh/install | sh` (PowerShell variant on Windows). Settings, projects and models are kept across upgrades - CUDA acceleration on NVIDIA GPUs, Metal (MPS) on Apple Silicon. CPU-only works too, just slower, and the CPU build of PyTorch takes about 5 GB of disk - Windows on ARM is experimental. Intel Macs only run the UI and need a remote backend The desktop shell is now [[Electron]]. Up to 0.5.3 it was built with [[Tauri]]; that shell was retired, and existing Tauri users have to install the Electron app separately. The GitHub topics still list `tauri` and `mlx` ([[MLX]]), so trust the README over the tags. ## VoiceStudio vs VoiceBox vs ElevenLabs [[VoiceBox]] is the closest sibling. It's open source and local-first too, for people who'd rather not rent their voice from a cloud provider, but its scope is narrower. | | VoiceStudio | VoiceBox | ElevenLabs | |---|---|---|---| | Hosting | Local, optional remote workers | Local | Cloud (SaaS) | | License | AGPL-3.0 | MIT | Proprietary, subscription | | Default model | OmniVoice (others selectable) | [[Qwen3-TTS]] | Their own | | Scope | Cloning, voice design, dubbing, dictation, transcription, audiobooks | Cloning, TTS, multi-track timeline editor, transcription | TTS, cloning, voice changing, dubbing, SFX, agents | | Desktop shell | Electron (Tauri until 0.5.3) | Tauri | Web and mobile apps | | Agent access | Local API, MCP, agent skills | REST API | API, hosted MCP | | Platforms | macOS, Windows, Linux, Docker | macOS, Windows (Linux planned) | Browser, iOS, Android | VoiceBox sticks to cloning and composing. VoiceStudio goes after the whole ElevenLabs feature set, locally, and adds dictation on top (which puts it near [[Wispr Flow]]). That breadth is also the risk: lots of features to keep working while the architecture is still changing. ## Things to know before using it - The license is [[Affero General Public License (AGPL)|AGPL-3.0]]. If you modify it and offer it as a network service, you have to publish your changes. Irrelevant for personal use; it matters if you build a product on top - Models have their own licenses. The app's license says nothing about the weights, so check each model before commercial use - Only clone voices with permission. The README says so explicitly (see [[Voice Cloning]]) - Expect breaking upgrades. The Tauri-to-Electron switch shows the architecture is still moving ## References - GitHub repository: https://github.com/debpalash/VoiceStudio - Website: https://voicestudio.sh - README (fetched 2026-10-04): install options, hardware table, Electron migration notice, license section - GitHub repository metadata (fetched 2026-10-04): creation date, stars, license, topics ## Related - [[VoiceBox]] - [[ElevenLabs]] - [[Supertonic]] - [[Text-to-Speech (TTS)]] - [[Voice Cloning]] - [[Speech-to-Text (STT)]] - [[ebook2audiobook]] - [[Local-First Software]] - [[Open Source]] - [[Model Context Protocol (MCP)]] - [[Electron]] - [[Tauri]] - [[Artificial Intelligence (AI)]] - [[Knowii Voice AI]] - [[Voice Clone Studio]] - [[Voice Clone Lab]]