# FrankenTTS FrankenTTS runs [[Qwen3-TTS]] (the 0.6B model from [[Qwen]]) on a plain CPU. It's a clean-room reimplementation of the model in pure [[Rust]], which gives you [[Text-to-Speech (TTS)]] and zero-shot [[Voice Cloning]] with no Python, no PyTorch and no GPU. It works on macOS, Linux and Windows, and inside a browser tab. Running an open TTS model usually means a Python environment, a few GB of dependencies and ideally an NVIDIA card. FrankenTTS replaces all of that with one binary and one ~1.77 GB model file. That's what makes speech possible in a CLI tool, a static web page or a laptop without a GPU. Jeffrey Emanuel built it (GitHub `Dicklesworthstone`, also the author of [[Coding Agent Session Search (cass)]] and [[Beads Viewer]]), as a sibling of his FrankenMarkdown project. The repo was created on 2026-08-06 and he announced it on LinkedIn around September 2026, saying it grew from "a quick-and-dirty, fun project" into what he considers the state of the art for TTS on consumer hardware without a GPU. ## How it works The pipeline follows Qwen3-TTS step by step: 1. A BPE tokenizer turns the text into tokens, and a 1,024-float voice vector says *who* is speaking 2. A 28-layer [[Transformers|transformer]] predicts one codec token per 80 ms frame (12.5 frames per second) 3. A 5-layer "microdecoder" adds 15 residual codes per frame, for 16 code groups in total 4. A causal convolutional codec turns the codes into 24 kHz audio Weights are [[AI Quantization|quantized]] to 8 bits (W8A8: 8-bit weights and 8-bit activations), which is how the model fits in ~1.77 GB and runs fast on a CPU. He also checked the port against the PyTorch reference. The README claims all 28 layers match exactly on argmax and the codec codes are 100% identical. Quantized ports tend to drift from the original, so if that claim holds, the output tokens are the same as Qwen's own. ## Voices and cloning There are 18 built-in voices. For your own, a speaker encoder squeezes a short recording into a single 1,024-float vector (an "x-vector"). No fine-tuning and no training run. A small denoiser (FastEnhancer, 207K parameters) cleans the enrollment audio first. According to the author, an iPhone voice memo or a MacBook microphone is good enough. ## Browser and CLI frankentts.com is a fully static site on [[Cloudflare Pages]]. The model runs in the tab through WebAssembly with SharedArrayBuffer threads (iOS gets a slower single-threaded build), and you can record yourself in the page to create a voice profile. A voice is only 1,024 numbers, so the profile plus your text fit in one compressed URL. Anyone can open that "say this in my voice" link, and no server is involved. For the command line, install with [[Homebrew]] or the install script, then: ```bash ftts pull # download the model ftts say "Hello there" out.m4a # synthesize to a file ``` The CLI keeps the model loaded in RAM for 10 minutes after a generation, so follow-up calls skip the load time. ## Performance Numbers from the author (real-time factor above 1 means faster than playback): | Setup | Speed | |---|---| | Native CLI on a Mac | 1.4 to 1.6x real time | | Browser (WebAssembly, threads) | 0.31 to 0.43x real time | | Time to first audio | ~450 ms | The native build keeps up with playback. In the browser, one second of audio takes roughly 2.5 to 3 seconds to generate: OK for a sentence, slow for an article. ## Compared to other voice tools [[VoiceBox]] and [[VoiceStudio]] wrap open models in a desktop app. FrankenTTS goes one level down and rewrites the inference itself, which is why it can run in a browser at all. Compared to [[ElevenLabs]], there's no per-character bill and the audio stays on your machine. The voice-in-a-URL idea worries me a bit, though. Cloning someone's voice now takes a phone recording and a link. Before building on it: - It's a young, single-author project: ~45 GitHub stars in early October 2026, last push on 2026-10-03 - The speed and accuracy numbers are self-reported, with no independent benchmarks yet - The code is MIT plus an extra rider that names OpenAI and Anthropic, so read it before commercial use. The Qwen weights stay Apache-2.0 - Only the 0.6B model is supported, not the larger Qwen3-TTS 1.7B or the VoiceDesign variants ## References - Jeffrey Emanuel's LinkedIn announcement (around September 2026): https://www.linkedin.com/posts/jeffreyemanuel_im-pleased-to-introduce-frankenttscom-ugcPost-7492557602313097216-uUfq - Website and in-browser demo: https://frankentts.com - Source code (README, license, repo stats): https://github.com/Dicklesworthstone/franken_tts ## Related - [[Qwen3-TTS]] - [[Qwen]] - [[Text-to-Speech (TTS)]] - [[Voice Cloning]] - [[Rust]] - [[Cloudflare Pages]] - [[Homebrew]] - [[VoiceBox]] - [[VoiceStudio]] - [[ElevenLabs]] - [[Supertonic]] (another on-device TTS model, much smaller, via ONNX) - [[Coding Agent Session Search (cass)]] (same author) - [[Beads Viewer]] (same author) - [[Voice Clone Studio]] - [[Voice Clone Lab]] - [[LuxTTS]] - [[Running AI Models Locally]] - [[Transformers.js]]