# Supertonic Supertonic is an open-weight [[Text-to-Speech (TTS)]] model built to run entirely on your own device. No cloud, no API key, no GPU required. It ships as a handful of [[ONNX]] files plus ready-made examples for Python, Node.js, the browser, iOS, Rust and more, so you can drop offline speech into almost any app. Why care? Most good TTS lives behind a paid API, and the open models that compete on quality sit in the 0.7B to 2B parameter class. Supertonic went the other way: it's tiny (66M parameters for v1 and v2, about 99M for v3) and very fast. Supertone claims it can turn a whole web page into audio in under a second. The tradeoff is expressiveness and, for some listeners, accuracy (more on that below). One thing to know before you build on it: **the project is archived**. Supertone announced the end of development on July 23, 2026, and the code and weights now live in a `supertone-oss-archive` organization. They still work and keep their licenses; nobody fixes bugs anymore. ## Who's behind it Supertone, Inc. was a South Korean voice AI company. Fast Company describes it as the brainchild of language processing and machine learning expert Lee Kyogu, and reports that HYBE (the entertainment group behind BTS) bought it for $36 million. HYBE, then still called Big Hit, had first invested ₩4 billion in February 2021. Its commercial products were Shift (real-time voice changer), Clear (voice separation, de-noise, de-reverb), Air (dialogue matching), Play (TTS app), the Supertone API and Voice Builder. On July 23, 2026, HYBE disclosed that shareholders had approved Supertone's dissolution. Shift, Clear and Air moved to a new company, Antinode Audio (supertone.ai now redirects there). Play, the API and Voice Builder were shut down; Voice Builder closed on August 31, 2026. The code and weights were moved to the archive organization in September 2026, with roughly 13.8K stars and 1.5K forks at the time of writing. ## How it works The design comes from the paper *SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System* (Kim et al., Supertone, arXiv 2503.23108, first posted March 2025, revised September 2025). Three parts: - **Speech autoencoder**: compresses audio into a low-dimensional continuous latent, with temporal compression, using ConvNeXt blocks. It decodes straight to 44.1 kHz audio. - **Text-to-latent module**: maps text to those latents with flow matching. The number of steps is a speed/quality knob (the paper uses 32; the released model's benchmarks use 2 or 5). - **Utterance-level duration predictor**: estimates how long the whole sentence should last. The interesting choice: it reads **raw characters**. No grapheme-to-phoneme (G2P) step and no external aligner; cross-attention learns the text-speech alignment during training. That matters in practice. As a r/rust commenter put it, many earlier small on-device models (he cites Piper and Kokoro) lean on espeak for phonemes, a GPL-licensed C library; Supertonic doesn't. Character input is also why it reads "$5.2M", "(212) 555-0142 ext. 402" or "30kph" without a text-normalization pass. The paper model had 44M parameters and was trained in English only: the autoencoder on 11,167 hours from about 14,000 speakers (public datasets plus Supertone's internal data), the text-to-latent and duration modules on 945 hours from 2,576 speakers (LJSpeech, VCTK, Hi-Fi TTS, LibriTTS). Later work from the same team (Length-Aware RoPE, self-purifying flow matching, RobustSpeechFlow) targets alignment and the classic flow-matching failure: skipped or repeated words. The released package has four ONNX graphs (text encoder, duration predictor, vector estimator, vocoder) and ten preset voice style files (F1 to F5, M1 to M5). ## Speed claims Supertone's own numbers for v1 (2 inference steps): | Setup | Characters per second | Real-time factor | |---|---|---| | M4 Pro, CPU (ONNX) | 912 to 1,263 | 0.015 to 0.012 | | M4 Pro, WebGPU | 996 to 2,509 | 0.014 to 0.006 | | RTX 4090 (PyTorch) | 2,615 to 12,164 | 0.005 to 0.001 | | Kokoro, M4 Pro CPU (ONNX) | 104 to 117 | 0.144 to 0.124 | | ElevenLabs Flash v2.5 (API) | 144 to 287 | 0.133 to 0.057 | | OpenAI TTS-1 (API) | 37 to 82 | 0.471 to 0.201 | The "167× faster than real time" headline is the best WebGPU figure (RTF 0.006). Read these with care: the cloud APIs were timed from Seoul, so network latency is baked in, and the RTX 4090 row uses the PyTorch model, which was never released. The local comparison with Kokoro on the same CPU is the fair one, and it's still a ~10× gap. Independent reports point the same way. A community WebGPU port (Transformers.js) claims ~1,750 characters per second after optimization and a ~5-hour audiobook generated in under 3 minutes. Supertone's own demo shows an Onyx Boox Go 6 e-reader in airplane mode at RTF 0.3, and a Raspberry Pi running it in real time. ## Accuracy and quality For v3, Supertone reports word/character error rates on the Minimax-MLS-test benchmark against much larger open models. English WER is 2.06 (VoxCPM2 2.11, OmniVoice 2.02, Qwen3-TTS 2.25, Supertonic 2 2.52). It is weaker in some languages: Vietnamese 4.49 vs 0.79 for OmniVoice, Finnish 5.40, Japanese CER 4.61. Supertone also says v3 cuts repeat/skip failures and improves speaker similarity over v2. Community reports on v1 (late 2025) are mixed: - One r/MachineLearning user who converts newsletters to audio found it "messes up words... once every 1/30 words or so" and recommended Kokoro instead. - Another, who runs it as an Android TTS backend, found the prosody "really monotonous", closer to old concatenative TTS and tiring for long texts. - The PageEcho developer, who ships it in an iOS e-book reader, described it as "very fast on mobile" and "natural and stable for long-form reading". - On Hacker News, the standout was text normalization: closed competitors stumbling on things like "Wed. June 23rd" while Supertonic read them correctly. My read: Supertonic wins on speed, size and messy input (numbers, units, currencies). It is not the model to pick for expressive narration. Test it on your own text before committing. ## Languages and voices - **v1**: English only. - **v2**: 5 languages (English, Korean, Spanish, Portuguese, French). - **v3**: 31 languages, including French, Dutch, German, Japanese, Arabic and Hindi; `lang="na"` handles text in an unknown language. v3 also adds 10 inline expression tags such as `<laugh>`, `<breath>` and `<sigh>`. Voices are 10 fixed presets. The open weights don't include voice cloning: the style encoder that turns a recording into a voice file was never released. Voice Builder (paid, hosted) filled that gap until it closed. Since then, a community project (`supertonic.embed`) recovers a voice style from a WAV by gradient descent; it needs an NVIDIA GPU with about 2.5 to 10 GB of memory. ## Where it runs Everything goes through ONNX Runtime, which is the whole point of the design. Official examples cover Python, Node.js, the browser ([[WebGPU]] or WASM via [[ONNX Runtime Web]]), Java, C++, C#, Go, Swift, iOS, Rust and Flutter. Output is 44.1 kHz 16-bit WAV, and batch inference is supported. The v1 model card says GPU mode was not tested; the CPU path is the main target. Community builds extend the reach: [[Transformers.js]] support, an MNN port (fp32/fp16/int8), a Kotlin Multiplatform SDK, an Android system-wide TTS engine, a Pinokio one-click install, browser extensions (Read Aloud, TLDRL), the PageEcho iOS reader, and Aftertone, which speaks replies from Cursor and [[Claude Code]] locally. ## Install and use The archive README pins a snapshot and avoids the old namespace: ```bash git clone https://github.com/supertone-oss-archive/supertonic.git cd supertonic && python3.11 -m venv .venv && source .venv/bin/activate python -m pip install huggingface_hub hf download supertone-oss-archive/supertonic-3 \ --revision aafc6e32416a594460b32413efc49d7fe4ce6d46 --local-dir assets python -m pip install -r py/requirements.txt cd py && python example_onnx.py --n-test 1 --text "Hello from Supertonic." --lang en ``` There is also a Python SDK (`pip install supertonic`): ```python from supertonic import TTS tts = TTS(auto_download=True) style = tts.get_voice_style(voice_name="M1") wav, duration = tts.synthesize("Your text here.", voice_style=style, lang="en") tts.save_audio(wav, "output.wav") ``` Careful: older SDK releases auto-download from the original `Supertone` Hugging Face namespace. The archive recommends downloading `assets/` yourself and passing `model_dir` with `auto_download=False`. SDK v1.3.1 (May 2026) added `supertonic serve`, a local HTTP server with an OpenAI-compatible `/v1/audio/speech` endpoint. ## License - **Code**: MIT. - **Weights**: BigScience OpenRAIL-M. Commercial use is allowed, with use-based restrictions (Attachment A). One of them is easy to miss: you may not publish generated content "without expressly and intelligibly disclaiming that the information and/or content is machine generated". The restrictions also flow down to your users. Not everyone likes this. One r/LocalLLaMA commenter called OpenRAIL "a trainwreck" that only burdens people who follow the rules. A r/rust commenter noted that only the ONNX model and the code to load it were released, not the model implementation; the repository holds inference examples only. ## Versions | Version | Date | Parameters | Languages | Notes | |---|---|---|---|---| | Supertonic 1 | November 2025 | ~66M | 1 (English) | 6 extra voices added December 2025 (10 total); PyPI package December 2025 | | Supertonic 2 | January 6, 2026 | ~66M | 5 | Voice Builder launched January 22, 2026 | | Supertonic 3 | April 29, 2026 | ~99M | 31 | Expression tags, fewer repeat/skip failures, v2-compatible ONNX interface | | Archive | September 2026 | | | Development ended; code and weights moved to `supertone-oss-archive` | ## Alternatives - **Kokoro** (82M, Apache): the usual comparison. About 10× slower in Supertone's benchmark, but two commenters above preferred it for clean, reliable output ("a dry reader, but always gives clean, sane output"). [[Audiblez]] uses it. - **Piper**: "the voice is a little robotic but it's reliable and lightweight", per a r/TextToSpeech commenter. - [[Qwen3-TTS]]: 0.6B and 1.7B open models with voice cloning; much heavier. - [[VibeVoice]]: Microsoft's open TTS; its Realtime 0.5B variant is in Supertone's comparisons. - [[ElevenLabs]] and [[Gemini 3.1 Flash TTS]]: cloud APIs with more expressive voices and cloning, at the cost of sending your text to a server and paying network latency. - Community follow-ups: TTSLibre, an effort to build a fully open (MIT code, CC0 data and weights) successor to Supertonic and Kokoro. ## References - GitHub repository (archived; the original `supertone-inc/supertonic` URL redirects): https://github.com/supertone-oss-archive/supertonic - Hugging Face model card, Supertonic 1: https://huggingface.co/Supertone/supertonic - Hugging Face model card, Supertonic 3 (archive): https://huggingface.co/supertone-oss-archive/supertonic-3 - Supertonic 3 model license (OpenRAIL-M): https://huggingface.co/supertone-oss-archive/supertonic-3/blob/main/LICENSE - Research demo page for the paper: https://supertonictts.github.io/ - Paper, SupertonicTTS (arXiv 2503.23108): https://arxiv.org/abs/2503.23108 - Paper, RobustSpeechFlow (arXiv 2605.22083): https://arxiv.org/abs/2605.22083 - supertonictts.com (third-party explainer site, states it is not affiliated with Supertone): https://supertonictts.com/ - Antinode Audio (where supertone.ai now redirects): https://supertone.ai - Fast Company, HYBE, Midnatt and Supertone (May 2023): https://www.fastcompany.com/90896187/hybe-bts-kpop-midnatt-supertone-ai-lee-hyun - Wikipedia, HYBE (investment and dissolution): https://en.wikipedia.org/wiki/Hybe_Corporation - HN, Supertonic launch: https://news.ycombinator.com/item?id=46028650 - HN, Show HN Supertonic 2: https://news.ycombinator.com/item?id=46510535 - HN, voice cloning without the style encoder: https://news.ycombinator.com/item?id=49845641 - supertonic.embed (voice style extractor): https://github.com/kdrkdrkdr/supertonic.embed - Reddit r/LocalLLaMA, Supertonic WebGPU: https://www.reddit.com/r/LocalLLaMA/comments/1p5r6vp/ - Reddit r/rust, open-source on-device TTS model: https://www.reddit.com/r/rust/comments/1p4ohus/ - Reddit r/MachineLearning, Supertonic 66M: https://www.reddit.com/r/MachineLearning/comments/1pj11sm/ - Reddit r/TextToSpeech, offline TTS on mobile: https://www.reddit.com/r/TextToSpeech/comments/1prd3vi/ - Reddit r/LocalTextToSpeech, TTSLibre: https://www.reddit.com/r/LocalTextToSpeech/comments/1w89n7l/ ## Related - [[Text-to-Speech (TTS)]] - [[ONNX]] - [[ONNX Runtime Web]] - [[WebGPU]] - [[Transformers.js]] - [[Edge AI]] - [[Running AI Models Locally]] - [[AI Open Weight Models]] - [[Voice Cloning]] - [[Qwen3-TTS]] - [[VibeVoice]] - [[ElevenLabs]]