Gemini 4 Argon: What's New, Benchmarks, and Pricing
Google's new frontier model is built for long, multi-step work across coding, legal, finance, and cybersecurity defense — here's what actually changed.

Google just announced Gemini 4 Argon, its new frontier model built to sustain deep reasoning across long, multi-step workflows. The pitch: frontier performance on real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense — the kind of tasks that take an agent many steps, not one prompt. This guide covers what's new, the benchmarks Google published, what it costs, the jump to a 1M output token limit, and who gets access first.
What is Gemini 4 Argon?
Gemini 4 Argon is Google's latest frontier model, positioned as its "next era of frontier intelligence." Google describes it as built to sustain deep reasoning across complex, long-horizon workflows, delivering frontier performance across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense (Google).
It was announced on September 30, 2026. Unlike a typical model drop, Argon isn't broadly available on day one — Google is rolling it out first to a set of trusted cyber defenders through its Fairwind Program, then expanding to developers, enterprises, and consumers as it hardens guardrails (Google).
What's new in Gemini 4 Argon
The headline change for builders is capacity to think longer in a single pass.
- 1M output token limit. Google significantly expanded the model's output token limit to an industry-leading 1M tokens, up from the previous 64K. The claim is that when the model has headroom to generate hundreds of thousands of tokens in one trajectory, it can add a new level of depth in reasoning and solve tough problems in one go (Google).
- Long-horizon focus. Argon is tuned for tasks that require planning, tool use, and staying on-task across many steps — everyday debugging through large-scale codebase migrations.
- Multimodal knowledge work. Google highlights strength when the task needs visual understanding — professional chart analysis, pulling details from long videos, and acting across a series of documents.
For context, this is a bigger release than Google's recent Gemini 3.8 Flash workhorse update: Argon is a frontier-tier model aimed at the hardest, longest tasks rather than cost-efficient high-volume work.
Gemini 4 Argon benchmarks
All figures below come from Google's launch materials, so read them as vendor-reported benchmarks rather than independent tests (Google).
- DeepSWE v1.1 (long-horizon software engineering): 77.9%, which Google calls a new state of the art for real-world, long-horizon engineering tasks.
- Vals Index: Argon is the leading model on this index, which measures economic impact across finance, coding, legal, and tax work, weighted by each sector's contribution to U.S. GDP.
- AutomationBench (Zapier): ranks #1 with a score of 51.3%, measuring end-to-end execution across core business functions.
- LVBench (long video understanding): 91.7%, which Google cites as state of the art.
- CWE-bench v1 (vulnerability remediation): ties for first with a top score of 68%.
Google also shared internal results: it says Argon helped quantum researchers beat a published baseline by 40% in minutes, and that Argon agents produced a Rust video decoder (for libgav1) that runs 2.7x faster than an existing Rust port with identical output. Treat these as illustrative internal examples, not reproducible benchmarks.
Gemini 4 Argon pricing
Google set an introductory price:
- $2 per million input tokens
- $10 per million output tokens
- Cached input tokens priced at 95% off the input token price
After the introductory period, the standard price rises to $4 per million input and $20 per million output tokens (Google). If you're planning around Argon, note that the 1M output ceiling plus frontier output pricing means long, token-heavy trajectories can get expensive quickly — budget for output, not just input.
Leading in defensive cybersecurity
Cybersecurity is the center of this launch. Google trained Argon to be highly capable at defense: it can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and Google's own internal teams, Google says it will release Argon without cyber guardrails so they can use its full frontier-level defensive capabilities (Google).
One early example Google cites: security vendor Wiz, using Argon through its Scan for Good initiative, uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — a severe risk Google says previous frontier models had missed.
This builds on the cyber direction Google started with its Gemini 3.8 Flash Cyber variant, extending autonomous vulnerability discovery and patching to a frontier-tier model.
Safety and the phased rollout
Because Argon ships frontier-level capability — including cyber — Google is gating it and strengthening safeguards before broad release. The company says it is:
- Engaged in the U.S. government's voluntary process for pre-release model access.
- Refusing harmful cyber and CBRN requests per its Frontier Safety Framework while preserving legitimate dual-use research.
- Hardening against indirect prompt injection, where malicious instructions hidden in context try to hijack the model — Google claims leading robustness on Gray Swan's Indirect Prompt Injection benchmark.
- Monitoring Argon's chain-of-thought and actions to catch misalignment and stop execution when needed (Google).
Who gets Gemini 4 Argon, and when
Access is staged:
- Now: trusted cyber defenders via the Fairwind Program, plus Google's internal teams.
- Next: developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.
In short, Argon is not a general-availability launch yet. If you build agents, the practical takeaway is to plan for it rather than prototype on it today — and to watch for the paid API opening.
Should you care if you build agents?
Argon's design — long trajectories, heavy tool use, and a 1M output ceiling — maps directly onto agentic workloads: codebase migrations, multi-step financial or legal research, and security remediation that spans many files. The open questions are cost at frontier output pricing and when you can actually call it outside Fairwind. Until then, the newer Gemini 3.8 Flash remains the accessible option for high-volume agent loops.
Do this with your own AI workforce
Frontier models like Gemini 4 Argon are only useful if something orchestrates them into real, multi-step work — planning, calling tools, and finishing a task end to end. Eigent is an open-source "Cowork" desktop app: a multi-agent workforce that runs locally, so you can point capable models at coding, research, and document workflows while keeping your data on your machine. See how an agent workforce handles reviewing GitHub PRs, or read our breakdown of Gemini 3.8 Flash for coding and agents. Download Eigent to build your own.
Recent Posts

Eigent Release Notes v1.0.5: Session Recovery, Task Queues & Better Previews
Eigent v1.0.5 improves Session recovery, task queues, process and file previews, Space settings, and model support.

Claude Opus 5.5: What's New, Benchmarks, and Pricing
Claude Opus 5.5 explained: the first Claude 5.5 model, 40% cheaper than Opus 5, 30% faster output, new agentic-coding benchmarks, pricing, and safety.

GPT-6 Sol and Luna: OpenAI's Cheaper Coding and Agentic Models
GPT-6 Sol and Luna extend OpenAI's GPT-6 family below Astra with big price cuts and better coding. Here's what changed, the benchmarks, and when to use each.