logo
  • Environments
  • Enterprise
  • Pricing
Blogs
Industry|Oct 2, 2026

Gemini 4 Argon: What's New, Benchmarks, and Pricing

Google's new frontier model is built for long, multi-step work across coding, legal, finance, and cybersecurity defense — here's what actually changed.

EigentEigent
Share to
Gemini 4 Argon: What's New, Benchmarks, and Pricing
  • What is Gemini 4 Argon?
  • What's new in Gemini 4 Argon
  • Gemini 4 Argon benchmarks
  • Gemini 4 Argon pricing
  • Leading in defensive cybersecurity
  • Safety and the phased rollout
  • Who gets Gemini 4 Argon, and when
  • Should you care if you build agents?
  • Do this with your own AI workforce
Automate Everything with
AI Workforce on Desktop
Download Eigent

Google just announced Gemini 4 Argon, its new frontier model built to sustain deep reasoning across long, multi-step workflows. The pitch: frontier performance on real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense — the kind of tasks that take an agent many steps, not one prompt. This guide covers what's new, the benchmarks Google published, what it costs, the jump to a 1M output token limit, and who gets access first.

What is Gemini 4 Argon?

Gemini 4 Argon is Google's latest frontier model, positioned as its "next era of frontier intelligence." Google describes it as built to sustain deep reasoning across complex, long-horizon workflows, delivering frontier performance across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense (Google).

It was announced on September 30, 2026. Unlike a typical model drop, Argon isn't broadly available on day one — Google is rolling it out first to a set of trusted cyber defenders through its Fairwind Program, then expanding to developers, enterprises, and consumers as it hardens guardrails (Google).

What's new in Gemini 4 Argon

The headline change for builders is capacity to think longer in a single pass.

  • 1M output token limit. Google significantly expanded the model's output token limit to an industry-leading 1M tokens, up from the previous 64K. The claim is that when the model has headroom to generate hundreds of thousands of tokens in one trajectory, it can add a new level of depth in reasoning and solve tough problems in one go (Google).
  • Long-horizon focus. Argon is tuned for tasks that require planning, tool use, and staying on-task across many steps — everyday debugging through large-scale codebase migrations.
  • Multimodal knowledge work. Google highlights strength when the task needs visual understanding — professional chart analysis, pulling details from long videos, and acting across a series of documents.

For context, this is a bigger release than Google's recent Gemini 3.8 Flash workhorse update: Argon is a frontier-tier model aimed at the hardest, longest tasks rather than cost-efficient high-volume work.

Gemini 4 Argon benchmarks

All figures below come from Google's launch materials, so read them as vendor-reported benchmarks rather than independent tests (Google).

  • DeepSWE v1.1 (long-horizon software engineering): 77.9%, which Google calls a new state of the art for real-world, long-horizon engineering tasks.
  • Vals Index: Argon is the leading model on this index, which measures economic impact across finance, coding, legal, and tax work, weighted by each sector's contribution to U.S. GDP.
  • AutomationBench (Zapier): ranks #1 with a score of 51.3%, measuring end-to-end execution across core business functions.
  • LVBench (long video understanding): 91.7%, which Google cites as state of the art.
  • CWE-bench v1 (vulnerability remediation): ties for first with a top score of 68%.

Google also shared internal results: it says Argon helped quantum researchers beat a published baseline by 40% in minutes, and that Argon agents produced a Rust video decoder (for libgav1) that runs 2.7x faster than an existing Rust port with identical output. Treat these as illustrative internal examples, not reproducible benchmarks.

Gemini 4 Argon pricing

Google set an introductory price:

  • $2 per million input tokens
  • $10 per million output tokens
  • Cached input tokens priced at 95% off the input token price

After the introductory period, the standard price rises to $4 per million input and $20 per million output tokens (Google). If you're planning around Argon, note that the 1M output ceiling plus frontier output pricing means long, token-heavy trajectories can get expensive quickly — budget for output, not just input.

Leading in defensive cybersecurity

Cybersecurity is the center of this launch. Google trained Argon to be highly capable at defense: it can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and Google's own internal teams, Google says it will release Argon without cyber guardrails so they can use its full frontier-level defensive capabilities (Google).

One early example Google cites: security vendor Wiz, using Argon through its Scan for Good initiative, uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — a severe risk Google says previous frontier models had missed.

This builds on the cyber direction Google started with its Gemini 3.8 Flash Cyber variant, extending autonomous vulnerability discovery and patching to a frontier-tier model.

Safety and the phased rollout

Because Argon ships frontier-level capability — including cyber — Google is gating it and strengthening safeguards before broad release. The company says it is:

  • Engaged in the U.S. government's voluntary process for pre-release model access.
  • Refusing harmful cyber and CBRN requests per its Frontier Safety Framework while preserving legitimate dual-use research.
  • Hardening against indirect prompt injection, where malicious instructions hidden in context try to hijack the model — Google claims leading robustness on Gray Swan's Indirect Prompt Injection benchmark.
  • Monitoring Argon's chain-of-thought and actions to catch misalignment and stop execution when needed (Google).

Who gets Gemini 4 Argon, and when

Access is staged:

  • Now: trusted cyber defenders via the Fairwind Program, plus Google's internal teams.
  • Next: developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.

In short, Argon is not a general-availability launch yet. If you build agents, the practical takeaway is to plan for it rather than prototype on it today — and to watch for the paid API opening.

Should you care if you build agents?

Argon's design — long trajectories, heavy tool use, and a 1M output ceiling — maps directly onto agentic workloads: codebase migrations, multi-step financial or legal research, and security remediation that spans many files. The open questions are cost at frontier output pricing and when you can actually call it outside Fairwind. Until then, the newer Gemini 3.8 Flash remains the accessible option for high-volume agent loops.

Do this with your own AI workforce

Frontier models like Gemini 4 Argon are only useful if something orchestrates them into real, multi-step work — planning, calling tools, and finishing a task end to end. Eigent is an open-source "Cowork" desktop app: a multi-agent workforce that runs locally, so you can point capable models at coding, research, and document workflows while keeping your data on your machine. See how an agent workforce handles reviewing GitHub PRs, or read our breakdown of Gemini 3.8 Flash for coding and agents. Download Eigent to build your own.

Recent Posts

Eigent Release Notes v1.0.5: Session Recovery, Task Queues & Better Previews
ProductSep 25, 2026

Eigent Release Notes v1.0.5: Session Recovery, Task Queues & Better Previews

Eigent v1.0.5 improves Session recovery, task queues, process and file previews, Space settings, and model support.

Douglas LaiDouglas Lai
Claude Opus 5.5: What's New, Benchmarks, and Pricing
IndustrySep 22, 2026

Claude Opus 5.5: What's New, Benchmarks, and Pricing

Claude Opus 5.5 explained: the first Claude 5.5 model, 40% cheaper than Opus 5, 30% faster output, new agentic-coding benchmarks, pricing, and safety.

EigentEigent
GPT-6 Sol and Luna: OpenAI's Cheaper Coding and Agentic Models
IndustrySep 22, 2026

GPT-6 Sol and Luna: OpenAI's Cheaper Coding and Agentic Models

GPT-6 Sol and Luna extend OpenAI's GPT-6 family below Astra with big price cuts and better coding. Here's what changed, the benchmarks, and when to use each.

EigentEigent
Automate everything with AI workforce on desktop
Download Eigent

Try Eigent today

Download the open-source desktop app. Your AI workforce, running on your machine.

Download Eigent
Eigent

Get the latest updates, tutorials, and releases on AI workforce automation.

Thank you for subscribing!

ProductEigentEnvironmentsPricingEnterprise
ExploreSolutionsUse CasesSkillsPluginsBlogs
DevelopersDocsGitHubCAMEL-AIOpen Source FundPartner
DownloadFor open source
CompanyAbout UsBrandCareersTerms of UsePrivacy PolicySecurity & TrustCookie PolicyRefund & Trial Policy

All rights reserved © 2026 EIGENT UK LTD