Skip to content
Gemini 4 Link
English简体中文繁體中文DeutschEspañolPortuguês

Gemini 4 has arrived: Argon is live, and most people still cannot use it

Gemini 4 is Google's first new flagship generation since the Gemini 3 series arrived in November 2025. The September 30 announcement names one model, Argon, and hands it to cyber defenders before anyone else.

Official key art from Google's Gemini 4 announcement, showing the Gemini four-pointed star mark next to the words Gemini 4 on a dark blue field of data points
Official key art published with Google's Gemini 4 Argon announcement on September 30, 2026.

Gemini 4 landed on September 30, 2026

Google announced Gemini 4 on September 30, 2026, and the first model of the generation is Gemini 4 Argon. It is the company's first new flagship since the Gemini 3 series arrived in November 2025. The ten months in between belonged to the Flash line and to multimodal variants, not to a new top-end model.

What makes the launch unusual is the order of access. Argon was not switched on inside the Gemini app for everyone. Google opened it first to vetted cyber defenders through a controlled early-access program called Fairwind.

Paid API customers and Google AI Ultra subscribers are named as the next groups to get in, and general consumers are given no date at all. Google says the wider rollout follows once the safety guardrails are further refined.

The reported facts in this article come from Google's public announcement of September 30, 2026 and the coverage published on October 1 and 2, 2026.

The flagship slot sat empty for ten months

Gemini 3 Pro arrived in November 2025. Through the first nine months of 2026 Google shipped Flash refreshes, a native Gemini app for Windows, live speech models and agentic video understanding, while the top of the line stayed where it was. OpenAI and Anthropic both kept shipping frontier models during that window.

On July 22, 2026 Sundar Pichai told analysts that the next step up would require much larger base models, and he named coding and agentic coding as the areas where Google had to close a gap. A day earlier, Google's own blog had already confirmed that an ambitious pre-training run for Gemini 4 was under way.

The organization changed as well. In August 2026 Demis Hassabis stepped back from the Google DeepMind chief executive role to become Alphabet's chairman and chief scientist, and Koray Kavukcuoglu took over DeepMind. On September 24 he said Argon had entered post-training and that he hoped for a launch much earlier than the end of the year. Six days later the model was out.

  • The Gemini 3 series shipped in November 2025, ten months before Gemini 4.
  • Coding and agentic coding were named as the gap Google had to close.
  • Argon moved from a post-training statement to a public launch in six days.

Twelve outright wins across eighteen evaluations

Google published a cross-model table with the announcement, comparing Gemini 4 Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Argon leads twelve evaluations outright and ties for first on one more. GPT-6 Astra takes three outright wins and Claude Opus 5.5 takes two.

The lead is broad rather than absolute. Argon trails GPT-6 Astra by 10.5 points on FrontierSWE v2 and on Terminal-Bench Science 0.1, and trails Claude Opus 5.5 by 9 points on Terminal-bench 4.0. The table below puts the biggest wins and the three misses side by side, with the gaps calculated rather than quoted.

EvaluationGemini 4 ArgonStrongest rivalGap
DeepSWE v1.1 (long-horizon software engineering)77.9%Claude Opus 5.5 at 74.2%Ahead by 3.7 points
LVBench (long-video understanding)91.7%GPT-6 Astra at 87.5%Ahead by 4.2 points
Harvey's Legal Agent Benchmark19.6%Claude Fable 5.1 at 6.7%Ahead by 12.9 points
AutomationBench (end-to-end business workflows)51.3%Claude Opus 5.5 at 42.5%Ahead by 8.8 points
FrontierSWE v255.0%GPT-6 Astra at 65.5%Behind by 10.5 points
Terminal-bench 4.057.4%Claude Opus 5.5 at 66.4%Behind by 9 points
The full benchmark comparison table from Google's announcement, listing Gemini 4 Argon against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across knowledge work, agentic coding, machine learning engineering, science and math, long context, computer use, multimodal understanding and cybersecurity
The complete cross-model table Google published with the Gemini 4 Argon announcement. Every score is vendor-reported.

Every figure here is Google's own reported result. No independent party has published a replication of the full comparison yet, so read the table as the vendor's case rather than a settled ranking.

Google is already running Argon on its own infrastructure

The most concrete evidence in the announcement is not a benchmark row. It is what Google says the model already does inside the company, where thousands of employees use it for specialised coding, deep research and writing.

A quantum computing team had a subroutine whose cost is measured by the product of qubits and gates. Argon found a new optimisation within minutes that beat the previously published baseline by 40 percent.

A second deployment analysed performance telemetry from Google data centers and carried out memory optimisations automatically. With the full set of changes deployed, Google reports more than 300 TiB of memory released and expects eventual savings between 500 TiB and 1 PiB.

  • The C and C++ to Rust migration covers re2, libgav1 and the Fuchsia Zircon kernel, more than 800,000 lines of code.
  • For libgav1, Argon rewrote roughly 32,000 lines of SIMD code in safe Rust.
  • The rewritten decoder keeps video output identical and runs 2.7 times faster than the earlier Rust port.
DeepSWE v1.1 bar chart from Google's announcement, showing Gemini 4 Argon at 77.9 percent against GPT-6 Astra at 74.1 percent, Claude Opus 5.5 at 74.2 percent and Claude Fable 5.1 at 67.4 percent
DeepSWE v1.1, the long-horizon software engineering benchmark behind the migration work Google describes.

These are internal results reported by Google rather than independent measurements, and the announcement gives no external audit of them.

Why cyber defense went first

Security is the capability Google chose to lead with. Argon scores 68 percent on CWE-bench v1, tied for first, and Google says it can discover, verify and patch vulnerabilities on its own. Vetted defenders and internal teams receive a version with some cyber guardrails relaxed so they can reach the full defensive capability.

Wiz, Google's cloud and AI security platform, already uses Argon in a program called Scan for Good that scans critical public infrastructure for free. In an early demonstration the model found a severe flaw in medical software used by hospitals worldwide that earlier frontier models had missed.

Holding the release back is the point rather than a delay for its own sake. Google describes four directions of guardrail work and a runtime mitigation layer that watches the model's reasoning and its actions and stops a task when it drifts away from the user's intent. The same kind of monitoring runs during training, and abnormal behaviour raises an alert to a dedicated incident response team. Google also avoids feeding that monitoring back into training, to reduce the chance of a model slowly learning to evade it.

CWE-bench v1 leaderboard across eleven models from Google's announcement, with Gemini 4 Argon tied for first at 68 percent alongside Grok 4.7 and GPT-6 Astra, ahead of Claude Opus 5.5 at 67 percent
CWE-bench v1, scored Pass@1 with ties broken by Pass@4.
Two charts from Google's announcement on discovering security vulnerabilities, showing real-world vulnerability discovery at 85.8 percent for Gemini 4 Argon against 71.0 for Gemini 3.8 Flash Cyber, and a penetration test benchmark at 70.9 against 58.2
Vulnerability discovery and black-box penetration testing, both reported by Google.

The security figures are vendor-reported as well, and no third party has reproduced the penetration testing result.

Introductory pricing undercuts the nearest rivals

Google listed introductory API rates of 2 US dollars per million input tokens and 10 US dollars per million output tokens, with cached input discounted by 95 percent. Standard rates after the introductory period step up to 4 and 20 dollars. Consumer plan pricing has not been announced.

Against the published rates of the models in its own comparison table, the introductory tier is the cheapest of the group by a wide margin, and even the standard tier matches Claude Opus 5.5 rather than undercutting it.

ModelInput per million tokensOutput per million tokens
Gemini 4 Argon (introductory)2 USD10 USD
Gemini 4 Argon (standard)4 USD20 USD
Claude Opus 5.54 USD20 USD
GPT-6 Astra10 USD50 USD

Costs on this site always come from its own server-side quote, and the Star Task ledger is the final billing record. No price on this page is a quote for work performed here.

Who gets in, and when

Access moves in phases. Fairwind partners come first, then paid API customers and Google AI Ultra subscribers, and only then developers, enterprises and consumers as Google finishes the safety work. No public date has been given for the free or Pro tiers of the Gemini app.

While the rollout is gated, the useful habit does not change. Read the announcement for direction, then check what your own workspace can actually call. Gemini 3.8 Flash is the model this site's chat workspace runs today, and the status page collects every announced Gemini 4 fact in one place.

Questions

When was Gemini 4 released?

Google announced Gemini 4, and its first model Argon, on September 30, 2026. It is the first new flagship generation since the Gemini 3 series in November 2025.

Is Argon the only Gemini 4 model?

Argon is the only Gemini 4 model Google has named so far. The generation is expected to grow, but no other model has been announced.

Can I use Gemini 4 today?

Not through the consumer Gemini app. Access began with vetted cyber defenders in the Fairwind program, and Google has not given a public launch date for the free or Pro tiers.

Why did Google start with cyber defenders?

Google says it wants the safety guardrails further refined before a wide release. Starting with vetted defenders lets the model work on real security problems while the guardrail and monitoring work continues.

Are the benchmark numbers independent?

No. Every figure in the announcement is Google-reported, and no third party has published a replication of the full comparison yet.

How is Gemini 4 different from Gemini 3.8 Flash?

Gemini 3.8 Flash is the newest Flash model and the one this site can call today. Argon is the Gemini 4 flagship, aimed at long-horizon work, and it is not listed in this workspace yet.

What does a 1M-token output limit change?

A single response can now carry up to 1,000,000 tokens instead of 64,000, so a long refactor, migration or report can finish in one pass instead of being cut off mid-way.

Does this site run Gemini 4?

No. This is an independent workspace that uses the Gemini 4 name as its brand and is not affiliated with Google. Its model picker lists only the models it can actually call at runtime.

Start with what you can call today

The workspace lists only the models your account can actually call, with a server-side quote before every request.

Open the chat workspace