Gemini 4 Argon benchmarks: what independent testing found that Google's own table leaves out
Google published its own benchmark table on September 30, 2026. Artificial Analysis published a second reading of the same model the same day, built from evaluations it runs itself. The two tables answer different questions.
The second table arrived the same day
Google closed its announcement on September 30, 2026 with an eighteen-row comparison table. Artificial Analysis published its own reading of the same model on the same day, built from evaluations it runs rather than figures supplied by the vendor.
On the Artificial Analysis Intelligence Index, Gemini 4 Argon scores 53 at its high reasoning setting. That is level with GPT-6 Astra at max reasoning and one point ahead of GPT-6.1 Sol. The gap against Google's own recent history is wider, with Argon sitting 23 points above Gemini 3.1 Pro Preview at 30 and 12 points above Gemini 3.8 Flash.
The independent report draws one conclusion from that. Google returns as one of the top three labs on measured intelligence. Argon is also the first proprietary Google model above the Flash class in more than seven months, which lines up with the ten-month gap in Google's own flagship timeline.
Figures in this article come from Google's Gemini 4 Argon announcement of September 30, 2026 and from Artificial Analysis, which published its own evaluation results the same day.
Why an independent run was the thing to wait for
Three weeks before the launch, an anonymous model labelled gemini-3.8-flash appeared in blind battles on a public arena and clearly outperformed the real Gemini 3.8 Flash on coding, diagram and agentic prompts. Developers widely assumed it was an early Gemini 4, but the displayed name never changed and Google never claimed it.
A benchmark chart that circulated on September 18, 2026 placed the unconfirmed model ahead of GPT-6 Astra and Claude Fable 5.1 on long-horizon coding and computer use, with a leaked price attached. The numbers travelled through several outlets before anyone could tie them to a model Google had actually named.
None of that makes the leaked chart wrong. It makes it unverifiable. What arrived on September 30 is different in kind, because it carries a published harness, a stated reasoning setting and a cost figure that can be checked against the public token price list.
- The arena alias was displayed as gemini-3.8-flash and was never confirmed by Google.
- The circulating chart was dated September 18, 2026, twelve days before the launch.
- Artificial Analysis runs its own evaluations instead of republishing vendor figures.
1.99 dollars a task, and the saving comes from the price list
Cost per task is the figure a budget notices. At the current 50 percent launch discount, Gemini 4 Argon costs 1.99 US dollars per Intelligence Index task against 3.26 dollars for GPT-6 Astra at max reasoning, or 60 percent of the price for a comparable level of measured intelligence.
The comparison flips against the cheaper OpenAI model. On the same evaluation GPT-6.1 Sol costs 2.7 times less per task than Argon. Once promotional pricing ends, the Argon figure rises to 3.98 dollars, which is roughly 1.2 times GPT-6 Astra rather than a discount against it.
The cause is not efficiency. Argon averages about 62,000 output tokens per task where GPT-6 Astra averages 27,000, so it spends more tokens to reach the same score and wins on unit price instead. Cached input also got cheaper, moving from a 90 percent discount on Gemini 3.8 Flash to 95 percent on Argon.
| Option | Input / output per million tokens | Cost per Intelligence Index task |
|---|---|---|
| Gemini 4 Argon (high), promotional | 2 / 10 USD | 1.99 USD |
| Gemini 4 Argon (high), standard | 4 / 20 USD | 3.98 USD |
| GPT-6 Astra (max) | 10 / 50 USD | 3.26 USD |
| GPT-6.1 Sol (max) | Not published in the same unit | About 0.74 USD, derived |
The Sol figure is not published as a cost per task. It is derived from the stated ratio of 2.7 times applied to the 1.99 dollar Argon result, and is included to show direction rather than to serve as a quote.
A 15 percent hallucination rate, the lowest in its class
The sharpest result in the independent report is not a capability score. On AA-Omniscience, which mixes factual accuracy with a willingness to decline when the model does not know, Gemini 4 Argon records a 15 percent hallucination rate. That is the lowest of any model scoring 45 or above on the Intelligence Index, against 51 percent for GPT-6 Astra at max and 54 percent for GPT-6.1 Sol.
Read beside the accuracy side, the picture gets subtler. Argon answers correctly 50 percent of the time, five points below Gemini 3.1 Pro Preview and thirteen points below GPT-6 Astra at 63 percent. On the combined AA-Omniscience score the three models land within a point of each other, with Argon and Sol at 42 and Astra at 43.
The split describes a model that declines instead of guessing. In work that runs unattended across long task chains, a wrong answer costs more than a missing one, which makes the trade the index records worth more than the headline capability gain.
Agentic work improved sharply and still is not the lead
Agentic ability was the weak point of the previous Gemini generation, and the independent run shows a large step. Gemini 4 Argon ranks first on AutomationBench-AA at 78 percent, seven points clear of Claude Sonnet 5.5 at max reasoning, on a benchmark built from real end-to-end business workflows.
Terminal work tells the other half of the story. On Terminal Bench 4, Argon scores 57 percent, which is 53 points above Gemini 3.1 Pro Preview and still behind three models at once, with Claude Sonnet 5.5 at 64 percent, Claude Opus 5.5 at 60 percent and GPT-6 Astra at 59 percent.
On AA-Briefcase, a knowledge work evaluation, Argon reaches 1,494 Elo on the strength of a 65 percent rubric pass rate, the highest Artificial Analysis has recorded. The sub-scores are less flattering, with analytical quality at 1,576 Elo and presentation quality at 1,308 Elo.
Same benchmark, two different headlines
Terminal-bench 4.0 appears in both tables and the two do not agree on who leads. Google's table puts Argon at 57.4 percent, behind Claude Opus 5.5 by nine points. The independent run puts Argon at 57 percent, behind Claude Sonnet 5.5 at 64 percent, with Opus 5.5 at 60 percent and GPT-6 Astra at 59 percent. The Argon score is nearly identical and the story around it is not.
The vendor table is also the only place some results exist at all. Harvey's Legal Agent Benchmark appears there with Argon at 19.6 percent against GPT-6 Astra at 5.4 percent and Claude Opus 5.5 at 3.8 percent, and the long-context GraphWalks evaluation shows a lead above ten points across the 256k to 1M window. Neither has been independently replicated.
The two tables answer different questions. A vendor table shows what a model can do when the vendor chooses the harness and the settings. An independent table shows what it does when someone else does. For a model still limited to vetted partners, with no consumer access and no session anyone can repeat, only the second kind can be checked by anyone outside the room.
| Terminal-bench 4.0 | Google's reported table | Artificial Analysis run |
|---|---|---|
| Gemini 4 Argon | 57.4% | 57% |
| Claude Opus 5.5 | 66.4%, the leader in this table | 60% |
| Claude Sonnet 5.5 | Not included | 64%, the leader in this run |
| GPT-6 Astra | Not the leading point in the row | 59% |
Neither table is wrong. They were produced under different harnesses, prompts and reasoning settings, which is exactly why a single ranking number should never be read on its own.
What to do with the second opinion
One more independent detail explains why long chains became practical. Artificial Analysis tested a Gemini API feature called Long Decode Continuation, which pauses a long response and resumes it across follow-up calls. Reasoning can therefore run up to 1,000,000 output tokens without the request timing out, against an output ceiling of 64,000 on the previous generation. The context window is 1M tokens, with text, image, video and speech input and text output.
Access has not moved with any of this. Argon is still limited to vetted Fairwind partners, with paid API customers and Google AI Ultra subscribers named as the next groups and no date attached. None of the results above can be reproduced in an ordinary account yet, which is the strongest argument for treating both tables as directional.
The practical move during the wait is the one the independent numbers reward. Build a fixed prompt set on the models you can call today, record the outputs, and rerun exactly the same set once Argon opens. A new model measured against your own baseline tells you more than the same model measured against anyone else's table, including this one.
- Argon is reachable only through the Fairwind program today.
- The output ceiling moves from 64,000 to 1,000,000 tokens.
- A fixed prompt set built now is the cheapest way to test Argon later.
Questions
What does Gemini 4 Argon score on independent benchmarks?
Artificial Analysis scores it 53 on its Intelligence Index at the high reasoning setting, level with GPT-6 Astra at 53 and one point ahead of GPT-6.1 Sol at 52. That is 23 points above Gemini 3.1 Pro Preview and 12 points above Gemini 3.8 Flash.
Why do the independent numbers differ from Google's?
They measure different things. A vendor table reports how a model performs on evaluations the vendor ran with its own settings. Artificial Analysis runs its own harness, prompts and grading, so the reasoning effort and scaffolding can differ even when the benchmark carries the same name.
What does Gemini 4 Argon cost per task?
At the current 50 percent discount, Artificial Analysis measures 1.99 US dollars per Intelligence Index task, against 3.26 dollars for GPT-6 Astra at max reasoning. On standard pricing the same task rises to 3.98 dollars, roughly 1.2 times GPT-6 Astra.
Why is Argon cheaper when it spends more tokens?
Because the discount applies to the price of a token, not to how many are used. Gemini 4 Argon averages about 62,000 output tokens per task where GPT-6 Astra averages 27,000, so the saving comes from a lower unit price rather than from doing less work.
How low is the hallucination rate?
15 percent on AA-Omniscience, the lowest of any model scoring 45 or above on the Intelligence Index. GPT-6 Astra sits at 51 percent and GPT-6.1 Sol at 54 percent. Argon answers correctly 50 percent of the time, below GPT-6 Astra at 63 percent.
Is Gemini 4 Argon the best model for agentic work?
Not yet. It ranks first on AutomationBench-AA at 78 percent, seven points clear of Claude Sonnet 5.5. On Terminal Bench 4 it scores 57 percent, behind Claude Sonnet 5.5 at 64 percent, Claude Opus 5.5 at 60 percent and GPT-6 Astra at 59 percent.
What makes a 1M-token output limit practical?
A Gemini API feature called Long Decode Continuation pauses a long response and resumes it across follow-up calls, so reasoning can run up to 1,000,000 output tokens without the request timing out. The previous generation capped output at 64,000 tokens.
Can I call Gemini 4 Argon here?
No. Argon is limited to vetted Fairwind partners, with paid API customers and Google AI Ultra subscribers named as the next groups and no date attached. The chat workspace on this site runs Gemini 3.8 Flash, the model it can actually call today.
Measure it against your own baseline
The chat workspace lists the models your account can call, with a server-side quote before every request.
Open the chat workspace