# OpenRouter Fusion OpenRouter Fusion turns one prompt into a small multi-model deliberation. A panel of models answers your prompt in parallel, an analyst model compares their answers, and the model you called writes the final reply from that comparison. It launched on June 12-13, 2026 as `openrouter/fusion` on [[OpenRouter]]. Despite the model slug, Fusion is NOT a model. There are no weights and nothing gets "fused" at the model level. It's an orchestration pipeline that OpenRouter runs server-side and exposes as a model alias, a server tool and a plugin. In classic ML terms it's an ensemble with an LLM judge (one Hacker News commenter said it should have been called exactly that). ## How it works 1. You call `openrouter/fusion`. OpenRouter resolves it to a real model and attaches an `openrouter:fusion` tool 2. That model decides whether the prompt deserves deliberation. Simple prompts get a direct answer. `tool_choice: "required"` forces Fusion every time 3. The **panel** (1 to 8 models) answers in parallel. Each panel model has `web_search` and `web_fetch` 4. The **analyst** reads all the answers (with web tools too) and returns structured JSON: consensus, contradictions, partial coverage, points only one model raised, and blind spots. It compares the answers; it doesn't merge them 5. Your model writes the final answer from that analysis A few details from the docs: - The analyst defaults to the model handling your request and always runs at temperature 0 - Each panel model and the analyst get 4 web tool calls by default (1 to 16) and 16,000 output tokens per call - If some panel models fail, you still get a result with a `failed_models` list. If the analyst fails, you get the raw panel answers without the analysis - Recursion is blocked: panel and analyst calls carry an `x-openrouter-fusion-depth` header and can't invoke Fusion again ## Ways to use it - **Chatroom**: openrouter.ai/fusion (labelled beta), pick a preset or build a panel, no code - **Model slug**: `"model": "openrouter/fusion"` on any OpenRouter endpoint (Chat Completions, Responses, Anthropic Messages) - **Server tool**: add `{"type": "openrouter:fusion"}` to the `tools` array of any model. This gives the most control. Server tools are in beta, and the docs say Fusion is slower on Chat Completions than on the Responses API - **Plugin**: `"plugins": [{"id": "fusion", ...}]` to set the panel (`analysis_models`), the analyst (`model`) or a `preset` Presets, as listed on the model page: | Preset | Panel | |---|---| | `general-high` (Quality, the default) | Claude Opus latest, GPT Sol latest, Gemini Pro latest | | `general-budget` | Gemini Flash latest, DeepSeek V4 Flash, Kimi latest | | `general-fast` | Gemini Flash latest, DeepSeek V4 Flash, Kimi latest (described as latency-homogeneous) | The docs also describe `openrouter/fusion-flash`, a separate slug with `general-fast` pre-selected. It didn't show up in the public `/api/v1/models` list when I checked on October 3, 2026; only `openrouter/fusion` did. ## Pricing Fusion has no price of its own. You pay for every underlying completion: N panel calls plus one analyst call plus your normal request. The docs estimate roughly **4 to 5 times** the cost of a single completion with the default 3-model panel, growing linearly with panel size. (The model page's generic FAQ says each request is billed "at the price of the model that serves it"; the dedicated docs are more precise.) Context length depends on the models you pick, up to 1M tokens. Latency follows the same logic. OpenRouter says a Fusion call is "often 2-3x longer" than a standard call. One HN commenter who ran a quick comparison against calling Opus 4.7 or [[GPT-5.5]] directly measured about 7x slower and 4x the cost. Another reported almost $1 per prompt in the web UI. ## The benchmark claims OpenRouter's launch post ("Surpassing Frontier Performance with Fusion", June 12, 2026) tested 100 deep research tasks from DRACO, a [[Perplexity]] benchmark with about 39 weighted criteria per task (factual accuracy, breadth and depth, presentation, citations, with negative weights for errors). | Setup | Score | |---|---| | Fusion: [[Claude Fable 5]] + GPT-5.5 (synthesized by Opus 4.8) | 69.0% | | Fusion: Opus 4.8 + GPT-5.5 + Gemini 3.1 Pro | 68.3% | | Fusion: Opus 4.8 + GPT-5.5 | 67.6% | | Fusion: Opus 4.8 + Opus 4.8 | 65.5% | | Solo: Claude Fable 5 | 65.3% | | Fusion: Gemini 3 Flash + [[Kimi K2.6]] + DeepSeek V4 Pro | 64.7% | | Solo: DeepSeek V4 Pro | 60.3% | | Solo: GPT-5.5 | 60.0% | | Solo: [[Claude Opus 4.8]] | 58.8% | Their three headline claims: panels beat individual models, frontier panels go beyond the frontier, and the budget panel comes within about 1 point of Fable 5 at roughly half the cost. Read the fine print before quoting those numbers: - **Fable 5 was scored on 93 tasks, not 100**: its content filters blocked 7, and OpenRouter chose not to fall back to Opus - **The grader differs from the paper**: they used Gemini 3.1 Pro Preview as judge instead of Gemini 3 Pro, so the scores aren't comparable to the DRACO paper. The paper itself reports 10 to 25 point shifts between judges - **Contamination**: with web search on, panel models found the DRACO rubric online. OpenRouter excluded those domains before producing the published results - **Same tools for everyone**: web search and web fetch (via Exa) plus bash, with a fixed tool-call budget. OpenRouter suspects that budget hurt Opus 4.8, "a hungrier model" - **Opus + Opus gained 6.7 points over solo Opus**, so a big part of the lift comes from the synthesis step (more samples) rather than model diversity - **One benchmark, one task type**: their own FAQ says Fusion is not a drop-in replacement for Fable, DRACO has no long-horizon tasks, and for coding they recommend the server tool so your coding model calls Fusion only for things like architecture decisions ## Community reaction The Hacker News thread (217 points, 85 comments, June 15, 2026) was curious but skeptical: - **The benchmark looked odd**: several commenters found it suspicious that two Opus 4.8 copies nearly match Fable 5, and that DeepSeek V4 Pro outranks Opus 4.8 and GPT-5.5. Others pointed out that the reasoning effort settings weren't disclosed - **"It's just more test-time compute"**: a recurring reading was that sampling the same model several times and picking the best parts is an old trick from the GPT-2/GPT-3 days, now with a smarter judge - **Judges reward similarity**: one commenter who built a similar "panel of experts" MCP server reported that asking one model to judge another mostly measures how close the answer is to what the judge would have said. Others said it works better with distinct expert personas, severity-scored issues or a verifiable target - **Cost surprises**: a few users saw an unexpected Claude Opus 4.8 call in their activity logs when Opus wasn't in their panel. Another commenter believed Opus was the default judge; the first user's point was that it shouldn't be charged without being disclosed upfront - **Plenty of DIY versions**: many said they had built the same thing (multi-model review of plans and specs, agent swarms, open-source clones that appeared within days). One reply argued the value for OpenRouter's customers is precisely not having to build it - **Good fits mentioned**: short, high-stakes inputs such as reviewing a markdown spec or a PRD for gaps, where tokens are cheap and mistakes are expensive - **Naming**: "It doesn't fuse anything" There was also a smaller thread on the launch post itself. One user reported Fusion runs failing with no reason given, and another found that Fable 5 alone gave a deeper answer than Fusion on the same query. ## My take The pattern isn't new, and the HN thread shows how many people already run their own version. What OpenRouter adds is packaging: a slug you can swap in, a structured analysis format, recursion and failure handling, and web tools on every panel member. I'd go for the server-tool mode, because your main model stays fast and only calls the panel when the question deserves it. Use it sparingly, for research questions, design reviews and decisions where being wrong costs more than 4-5x the tokens. Avoid it in tight agent loops. And don't take the DRACO table as proof that a budget panel "is" Fable 5: it's one deep-research benchmark, graded by one LLM, run by the vendor. For contrast, OpenRouter's Auto Router and Pareto Router pick ONE model per request (see [[Model routing]]); Fusion runs several and pays for all of them. ## References - Model page: https://openrouter.ai/openrouter/fusion (read via r.jina.ai; the model list and presets were extracted from the page's HTML) - Fusion Router docs: https://openrouter.ai/docs/guides/routing/routers/fusion-router.md - Fusion server tool docs: https://openrouter.ai/docs/guides/features/server-tools/fusion.md - Launch post, "Surpassing Frontier Performance with Fusion" (June 12, 2026, with a June 14 FAQ update): https://openrouter.ai/blog/announcements/fusion-beats-frontier/ - Fusion chatroom (beta): https://openrouter.ai/fusion - OpenRouter models API entry (release timestamp, pricing marked variable, 1M context): https://openrouter.ai/api/v1/models and https://openrouter.ai/api/v1/models/openrouter/fusion/endpoints - DRACO benchmark paper (Perplexity): https://arxiv.org/abs/2602.11685 - Hacker News discussion of the model page (217 points, 85 comments): https://news.ycombinator.com/item?id=48537641 (full comment tree via https://hn.algolia.com/api/v1/items/48537641) - Hacker News discussion of the launch post: https://news.ycombinator.com/item?id=48525392 - Earlier Hacker News submission of the Labs version (April 2026): https://news.ycombinator.com/item?id=47688780 - Reddit and X: no discussion could be retrieved (search pages not readable without a browser) ## Related - [[OpenRouter]] - [[Model routing]] - [[AI Gateway]] - [[AI Mixture of Experts (MoE)]] - [[AI Sycophancy]] - [[Claude Fable 5]] - [[Claude Opus 4.8]] - [[GPT-5.5]] - [[Kimi K2.6]] - [[Deepseek]] - [[Perplexity]]