# Open-LLM-VTuber
Open-LLM-VTuber (OLV) is an open-source AI companion you talk to with your voice. It listens through your microphone, sends what you said to an LLM, speaks the answer back, and shows an animated Live2D character (the 2D anime-style rigged avatars VTubers use) whose expressions follow the conversation. Every piece can run locally, so it works offline.
The name explains the origin. The project started as an attempt to recreate Neuro-sama, the closed-source AI VTuber that streams on Twitch, using open-source parts that also run outside Windows. Today the docs pitch it mostly as a personal companion ("virtual girlfriend, boyfriend, cute pet") and a desktop pet.
## The problem it solves
Talking to a local LLM by voice usually means wiring speech recognition, a model and a TTS engine together yourself, then fighting audio feedback because the mic hears the assistant. OLV packages that loop, adds a face, and lets you swap every component through one config file (`conf.yaml`).
## How it works
A Python (FastAPI) backend runs the voice loop and serves a web frontend over WebSocket:
1. **Listen**: voice activity detection plus speech recognition (sherpa-onnx, FunASR, Faster-Whisper, Whisper.cpp, Groq Whisper, Azure...)
2. **Think**: any LLM ([[Ollama]], [[LM Studio]], vLLM, GGUF, or OpenAI-compatible, Claude, Gemini, DeepSeek, Mistral APIs)
3. **Speak**: [[Text-to-Speech (TTS)|TTS]] (sherpa-onnx, MeloTTS, Coqui, GPT-SoVITS, CosyVoice, Bark, Edge TTS, Fish Audio, Azure...)
4. **Animate**: the LLM emits emotion tags like `[joy]`, which get mapped to the Live2D model's expressions
The quick-start default is Ollama + sherpa-onnx (SenseVoiceSmall) + Edge TTS. Swap Ollama for an API and the speech recognition for Groq Whisper and you don't need a GPU at all.
## Key features
- **Voice interruption without headphones**: you can cut the AI off mid-sentence and it won't hear its own voice
- **Visual perception**: it can look at your webcam, your screen or a screenshot (opt-in)
- **Desktop pet mode** (Electron client): transparent, always on top, click-through, so the character sits on your desktop while you work
- **Tool calling through [[Model Context Protocol (MCP)|MCP]]** (time, DuckDuckGo search, BrowserBase browser control with a live view), added in v1.2.0
- **Long-term memory** through Letta (v1.2.0). The README still says built-in memory is "temporarily removed", so Letta is the route
- **Live streaming hook**: a Bilibili live-comment client and an interface for other platforms
- Group chat between several AI characters, proactive speaking, "inner thoughts" display, chat history, TTS translation (chat in one language, voice in another)
- Custom Live2D models (Cubism 5; Cubism 2 support was dropped in v1.2.0)
## Install and usage
Free. You need Git, [[FFmpeg]] (mandatory), `uv` and Python 3.10 to 3.12.
1. Download the `Open-LLM-VTuber-v1.x.x.zip` from the Releases page (the docs insist: not the "Download ZIP" button) or clone the repo
2. `uv sync`, then `uv run run_server.py` once to generate the config, then edit `conf.yaml`
3. Open `http://localhost:12393` in **Chrome** (the docs say Edge and Safari have known issues)
4. Optional: the Electron desktop client (`.exe` / `.dmg`) from the same Releases page for pet mode
Running it on another machine than the browser needs HTTPS (browsers only open the mic in a secure context). There's a Dockerfile and a separate Docker config repo since February 2026. Minimum hardware is "a computer; a Raspberry Pi works too" if you use APIs; for local models they recommend an Apple Silicon Mac or a decent GPU.
## License
It's mixed, and it's moving:
- **Backend**: MIT ("Copyright (c) 2025 Yi-Ting Chiu"). GitHub shows the license as "Other" because the LICENSE file carves out the Live2D models
- **Frontend** (web and Electron clients, `Open-LLM-VTuber-Web`): "Open-LLM-VTuber License 1.0" since v1.2.0, i.e. Apache 2.0 plus extra conditions. Free for personal, educational and non-profit use, and explicitly free for "VTuber streaming and content creation" including monetized YouTube or Twitch channels. A commercial license is required to sell hosted access, rebrand and resell it, or embed it in a paid product. The maintainers said in the v1.2.0 notes they'd move the backend to the same license "around v1.3 or v1.4"; that hasn't happened yet
- **Live2D sample models**: Live2D Inc.'s own terms, NOT covered by MIT. Under the Free Material License, individuals and small businesses with less than 10 million yen in yearly sales get broader rights than bigger companies, and each sample character has its own terms
That last point isn't theoretical. In July 2026 Live2D Inc. sent a DMCA notice that got the sibling project linked from the README (LLM-Live2D-Desktop-Assistant, by OLV's second-biggest contributor) blocked on GitHub, for redistributing the Cubism Core runtime and the "Shizuku" sample model. The main OLV repo is still up.
## Maturity
- **Stars**: about 14,000, 1,670 forks, 157 open issues (GitHub API, 2026-10-03)
- **Created**: 2023-11-24
- **Last release**: v1.2.1 (2025-08-26). Nothing since
- **Activity**: a burst of commits in January and February 2026 (Docker, FireRedASR), then one README fix on 2026-05-15. Nothing since. The README announces a full v2.0 rewrite "in its early discussion and planning phase" and asks people not to open feature requests for v1
- Translation: v1 is stable-ish and feature-frozen, v2 doesn't exist yet. If you build on it, expect to maintain your own fork
## Who's behind it
Yi-Ting Chiu (GitHub `t41372`, Tempe, Arizona), who created it and wrote 643 commits, with `ylxmf2005` (139 commits, lead on the frontend) and 80+ other contributors. It lives in the Open-LLM-VTuber GitHub organization, takes donations through Buy Me a Coffee, and runs its dev community on Zulip (plus Discord and a QQ group). A big share of the community is Chinese-speaking: parts of the docs default to Chinese.
## Limits and caveats
- **Latency**: a pipeline of separate speech recognition, LLM and TTS steps adds up. An independent review measured about 1.2 seconds before the answer starts with a 7B model on an M2 Pro Mac mini. Fine for chatting, sluggish next to end-to-end voice models
- **Stalled project** (see Maturity)
- **Licensing** of the avatar art (see License): use your own Live2D model if you plan to show it publicly
- **Single user**: one voice session at a time
- **Security**: exposing it beyond localhost means you set up HTTPS and access control yourself
- **Setup friction**: CUDA/cuDNN on Windows, model downloads, Chrome only
## Relevance for my content (YouTube, courses)
As a production tool, close to zero. It doesn't make videos and it's built for companionship and streaming, not teaching. An anime avatar is the wrong face for a course about knowledge management.
Where it's interesting:
- **As a reference architecture.** It's one of the most complete open-source examples of a real-time voice agent: speech recognition, LLM, TTS, echo-free interruption, MCP tools, memory, all swappable. That overlaps a lot with what I care about in voice AI ([[Speech-to-Text (STT)]], [[Text-to-Speech (TTS)]], local models with [[Ollama]])
- **As a content topic.** "Run a private voice assistant on your own machine" makes a good demo for an AI video, as long as you show the license and latency caveats
- **AI co-host for a live stream**: the license explicitly allows monetized streams, but the project's frozen state makes it a hobby experiment, not something I'd rely on
## Alternatives
- **AIRI** (moeru-ai): another self-hosted Neuro-sama-style companion, MIT, about 50,000 stars and still actively developed (last push on 2026-10-03). It can even play Minecraft and Factorio
- **z-waif**: Windows-focused, simpler setup
- **Voxta**: commercial AI companion app
- **OpenAI Advanced Voice / Gemini Live**: much lower latency and better voices, no avatar, cloud only
- For video avatars rather than chat companions: **[[HeyGen]]**
## References
- GitHub repository and README: https://github.com/Open-LLM-VTuber/Open-LLM-VTuber
- Documentation, project overview (English): http://docs.llmvtuber.com/en/docs/intro/
- Documentation, quick start: https://open-llm-vtuber.github.io/docs/quick-start
- Release v1.2.0 notes (Letta memory, MCP, Cubism 5, frontend license change, planned backend license change): https://github.com/Open-LLM-VTuber/Open-LLM-VTuber/releases/tag/1.2.0
- Releases (v1.2.1, 2025-08-26): https://github.com/Open-LLM-VTuber/Open-LLM-VTuber/releases
- Backend LICENSE (MIT) and LICENSE-Live2D.md: https://github.com/Open-LLM-VTuber/Open-LLM-VTuber/blob/main/LICENSE
- Frontend LICENSE (Open-LLM-VTuber License 1.0): https://github.com/Open-LLM-VTuber/Open-LLM-VTuber-Web/blob/main/LICENSE
- Live2D Free Material License Agreement: https://www.live2d.jp/en/terms/live2d-free-material-license-agreement/
- GitHub DMCA notice from Live2D Inc. (2026-07-20) against LLM-Live2D-Desktop-Assistant: https://github.com/github/dmca/blob/master/2026/07/2026-07-20-live2d.md
- GitHub API (stars, forks, issues, commits, contributors, maintainer profile), queried 2026-10-03: https://api.github.com/repos/Open-LLM-VTuber/Open-LLM-VTuber
- andrew.ooo, "Open-LLM-VTuber Review: Offline AI Companion with Live2D" (2026-06-08, updated September 2026; latency measurements, community sentiment): https://andrew.ooo/posts/open-llm-vtuber-offline-ai-companion-review/
- Hacker News, "Show HN: Recreation of Neuro-sama for a year" (AIRI, 2025-07-15): https://news.ycombinator.com/item?id=44573640
- AIRI repository: https://github.com/moeru-ai/airi
- Reddit: r/vtubers post asking for open-source AI VTuber projects with Ollama and TTS (2026-06-10): https://reddit.com/r/vtubers/comments/1u27nlw/looking_for_an_opensource_ai_vtuber_project_with/
## Related
- [[Text-to-Speech (TTS)]]
- [[Speech-to-Text (STT)]]
- [[Ollama]]
- [[LM Studio]]
- [[Model Context Protocol (MCP)]]
- [[AI Agents]]
- [[Electron]]