# Gemini 4 Argon
Gemini 4 Argon is [[Google DeepMind]]'s frontier [[Gemini]] model, announced on 30 September 2026. It's the first model of the Gemini 4 generation, and according to [[Artificial Analysis]] it's Google's first proprietary model above the Flash class in more than seven months. Google built it for long, multi-step work: real-world software engineering, enterprise knowledge work (legal, finance, tax), and cybersecurity defense.
Why does it matter? For most of 2026, Google competed with cheap and fast Flash models (see [[Gemini 3.6 Flash]]) while [[Anthropic]] and [[OpenAI]] held the top of the leaderboards. Argon is the model that puts Google back in that top group: Artificial Analysis scores it level with [[GPT-6 Astra]] on its Intelligence Index.
My launch take, with the community reaction, is in [[2026-10-03 Gemini 4 Argon - Google is back at the top, but you can't use it yet]].
## Key facts
- **Announced by** Koray Kavukcuoglu (SVP at Google DeepMind and Chief AI Architect at Google) on the Google blog, and by [[Demis Hassabis]] on X
- **Context window**: 1M tokens (per Artificial Analysis)
- **Output limit**: up to 1M output tokens, up from 64K on previous Gemini models. Google's argument: with that headroom, the model can generate hundreds of thousands of tokens in a single trajectory and solve hard problems in one go. A new Gemini API feature, Long Decode Continuation, pauses long responses and resumes them across follow-up calls, so reasoning can run that long without request timeouts
- **Modalities**: text, image, video and speech input; text output (per Artificial Analysis)
- **Reasoning**: Google ran its own benchmarks at the highest thinking setting. Artificial Analysis tested "high", the highest setting available to them
- **Pricing**: introductory price of $2 per million input tokens and $10 per million output tokens; cached input is 95% off ($0.10 per million). After the introductory period: $4 / $20. Google hasn't said when the promotion ends (Artificial Analysis says "at least one month")
- **Availability**: restricted. First to trusted cyber defenders through Google's Fairwind Program, while Google takes part in the U.S. government's voluntary pre-release model access process. Next in line: paid API customers and Google AI Ultra subscribers, then developers, enterprises and consumers. No date
## What it's good at
**Knowledge work.** In Google's comparison table, Argon leads on the Vals Index (68.9% vs 67.0% for [[Claude Opus 5.5]]), Vals Finance Agent v2 (65.4% vs 58.9% for [[Claude Fable 5.1]]), Zapier's AutomationBench (51.3% vs 42.5% for Opus 5.5), and Harvey's Legal Agent Benchmark. That last one is the outlier: 19.6% vs 6.7% for the next best model, roughly three times higher. Artificial Analysis independently ranks it #1 on its own AutomationBench-AA (78%).
**Long context and vision.** GraphWalks at 256K to 1M tokens: 84.2% vs 71.8% for GPT-6 Astra. LVBench (long video understanding): 91.7%, state of the art. Chartography (chart analysis): 71.6%.
**Coding at scale.** DeepSWE v1.1 (long-horizon software engineering): 77.9%, state of the art per Google. Vibe Code Bench: 91.9%. Inside Google, Argon agents migrate C/C++ codebases to [[Rust]], from tens of thousands of lines (re2, libgav1) up to 800K+ lines for the Fuchsia Zircon kernel. For libgav1, they replaced 32K lines of SIMD code with safe Rust the compiler vectorizes automatically; the result runs 2.7x faster than the existing Rust port, with identical output. Google says these rewrites still go through automated and manual audits before production.
**Cyber defense.** Google trained it to find, validate and patch vulnerabilities on its own. It ties for first on CWE-bench v1 (68%, with GPT-6 Astra) and beats Gemini 3.8 Flash Cyber on Google's internal vulnerability benchmark (20 programming languages) and on Wiz's black-box penetration testing benchmark. Trusted defenders and Google's internal teams get it without cyber guardrails. Through Wiz's Scan for Good program, it found a critical vulnerability exposing personal data in healthcare software used by hospitals worldwide, which previous frontier models had missed.
**Knowing what it doesn't know.** This is the most interesting number to me. On AA-Omniscience, Argon hallucinates 15% of the time, the lowest of any model scoring 45+ on the Artificial Analysis Intelligence Index (GPT-6 Astra: 51%, GPT-6.1 Sol: 54%). Its accuracy is 50%, lower than Astra's 63%. So it doesn't know more; it says "I don't know" instead of guessing.
## Where it trails
Google's own table shows five benchmarks where Argon loses:
| Benchmark | Gemini 4 Argon | Best competitor |
|---|---|---|
| FrontierSWE v2 | 55.0% | 65.5% (GPT-6 Astra) |
| Terminal-Bench 4.0 | 57.4% | 66.4% (Claude Opus 5.5) |
| PostTrainBench (ML engineering) | 45.3% | 49.3% (Claude Opus 5.5) |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% (GPT-6 Astra) |
| OSWorld-2.0 (computer use) | 69.2% | 72.6% (GPT-6 Astra) |
Artificial Analysis confirms the terminal weakness: 57% on Terminal-Bench 4, behind [[Claude Sonnet 5.5]] (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%). It's still a 53-point jump over Gemini 3.1 Pro Preview.
It's also token hungry: 62k output tokens per Artificial Analysis task, vs 27k for GPT-6 Astra (max). Its cost advantage comes from the token price, not from efficiency.
## Positioning vs other models
Artificial Analysis Intelligence Index and cost per task (30 September 2026):
| Model | Intelligence Index | Cost per task |
|---|---|---|
| Gemini 4 Argon (high) | 53 | $1.99 at intro price, $3.98 at standard price |
| GPT-6 Astra (max) | 53 | $3.26 |
| GPT-6.1 Sol (max) | 52 | about 2.7x cheaper than Argon |
| Gemini 3.8 Flash (high) | 41 | |
| Gemini 3.1 Pro Preview | 30 | |
- **vs GPT-6 Astra**: same score, 60% of the cost per task at the intro price, about 1.2x the cost at the standard price
- **vs GPT-6.1 Sol**: one point higher, 2.7x more expensive per task
- **vs Claude Opus 5.5**: in Google's table, ahead on knowledge work, long context, multimodal and science (LABBench 2: 88.8% vs 73.1%), behind on terminal coding and ML engineering
- **vs previous Gemini models**: +23 points over Gemini 3.1 Pro Preview and +12 over Gemini 3.8 Flash. Gemini 3.5 Pro had been announced at Google I/O in May 2026 (see [[Gemini 3.5 Pro]]), but Artificial Analysis compares against 3.1 Pro Preview as Google's previous non-Flash model
## Safety approach
Google holds the broad release back while it strengthens safeguards in four areas:
- **Misuse**: refuse cyber and CBRN (chemical, biological, radiological, nuclear) attack requests while keeping legitimate dual-use research, per Google's Frontier Safety Framework. Includes monitoring the model's internal activations to spot misuse, tested by internal and external red teams
- **Prompt injection**: Google calls it its most resilient model against indirect prompt injection, leading on Gray Swan's IPI benchmark
- **Misalignment**: monitors watch Argon's chain-of-thought and actions and stop execution when needed. Google used a similar system during training, without feeding the findings back into training (so it wouldn't teach the model to hide its reasoning from the monitor), and asks the rest of the industry to keep reasoning transparent
- **Hardened sandboxes**: isolated, sealed environments before high-risk training or evaluations
See [[AI Safety]].
## Caveats
- **Benchmarks are Google's setup.** Competitor numbers are mostly self-reported at their maximum reasoning setting. Several Argon numbers are self-computed (DeepSWE with a mini-swe-agent harness, Terminal-Bench 4.0, OSWorld-2.0 maxed over three runs). The two Terminal-Bench figures for Opus 5.5 (66.4% in Google's table, 60% at Artificial Analysis) show how much the test setup matters
- **You can't test it yet.** As of 3 October 2026, it isn't listed on the Gemini API models or pricing pages. Artificial Analysis is the only independent evaluation so far
- **The price is temporary.** Budget with $4 / $20, not $2 / $10
- **Open questions**: the API model ID, a release date, and whether other Gemini 4 sizes follow
## References
- Google announcement (Koray Kavukcuoglu, 30 September 2026): https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon
- Google DeepMind Gemini page (benchmark table): https://deepmind.google/models/gemini/
- Evaluation methodology (PDF): https://deepmind.google/models/evals-methodology/gemini-4-argon
- Demis Hassabis on X: https://x.com/demishassabis/status/2105417239432200636
- Logan Kilpatrick on X: https://x.com/OfficialLoganK/status/2105388054274080946
- Artificial Analysis model page: https://artificialanalysis.ai/models/gemini-4-argon
- Artificial Analysis launch analysis: https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs
- The Verge: https://www.theverge.com/tech/1002980/google-gemini-4-argon
- Hacker News discussion (announcement): https://news.ycombinator.com/item?id=49913571
- Hacker News discussion (Artificial Analysis): https://news.ycombinator.com/item?id=49914236
- Gemini API models (checked 3 October 2026, Argon not listed): https://ai.google.dev/gemini-api/docs/models
## Related
- [[2026-10-03 Gemini 4 Argon - Google is back at the top, but you can't use it yet]]
- [[Gemini]]
- [[Gemini 3]]
- [[Gemini 3.5 Pro]]
- [[Gemini 3.6 Flash]]
- [[Google DeepMind]]
- [[Google]]
- [[Demis Hassabis]]
- [[GPT-6 Astra]]
- [[Claude Opus 5.5]]
- [[Claude Fable 5.1]]
- [[Claude Sonnet 5.5]]
- [[Artificial Analysis]]
- [[AI Frontier Model]]
- [[AI Safety]]
- [[Large Language Models (LLMs)]]