IA4 MIN

Gemini 4 Argon tops Vals ahead of Fable 5.1 and GPT-6 Astra, with access restricted

Google’s new generation starts with selected cybersecurity defenders. An independent work-task index puts Argon first, while coding results show why one leaderboard cannot settle every comparison.

Original ValsIndex chart showing Argon 68.90, Fable 5.1 65.83 and Astra 63.13, with one-standard-error bars
Image: INSERT FUTURE · original graphic. Data: Vals AI, 30 Sep 2026. Error bars: ±1 standard error.

Gemini 4 Argon has an independent leaderboard win behind its launch. Vals ranks it first on its professional-task index, ahead of Claude Fable 5.1 and GPT-6 Astra. Google announced the model on September 30, but initial access is going to selected cybersecurity defenders. The announcement does not make Argon generally available in the Gemini app.

The model is built for work that requires several steps, including editing code and using tools to check the result. That is a different proposition from generating a single answer. Its usefulness still depends on whether those steps succeed on the work a particular user needs done.

01

A first-place finish with a defined scope

ValsIndex combines coding, finance, legal and tax evaluations, weighting sectors with reference to their share of the U.S. economy. Argon scores 68.90, against 65.83 for Fable 5.1 and 63.13 for GPT-6 Astra. These are composite scores, not estimates of the share of jobs each model can automate.

Vals says it evaluated Argon on September 30. That gives the launch evidence beyond Google’s own testing, without establishing a winner in every discipline. Argon’s same model page places it fifth on Harvey’s legal benchmark and second on VibeCode. The overall ranking compresses those different strengths into one index.

The small error bars in our cover graphic show Vals’ published standard errors, a measure of uncertainty in the estimated scores. They are not a direct statistical test of the difference between each pair of models.

02

Coding results point in different directions

On Google’s comparison page, Argon scores 77.9% on DeepSWE v1.1, ahead of Astra at 74.1% and Fable 5.1 at 67.4%. DeepSWE gives agents specific changes to make in open-source projects. Automated tests check both the requested behavior and whether existing behavior still works. The tools and instructions surrounding the model therefore matter alongside the model itself.

FrontierSWE v2, a separate software-engineering evaluation, changes the order. Argon scores 55.0%, compared with 65.5% for Astra and 56.3% for Fable. Readers tracking GPT-6 Astra and Fable 5.1 have evidence of a strong new competitor, rather than a universal reason to replace their current tool.

The numbers also come from different evaluation runs. Google’s methodology document mixes its own testing with public leaderboards and other providers’ published results. For DeepSWE, Google computed Argon’s result, used Astra’s leaderboard score and took Fable’s result from its technical documentation. High reasoning settings do not, by themselves, make all those runs identical.

AutomationBench needs an additional qualification. It tests whether an agent completes tasks across a simulated company’s apps. The 31.4% result listed for Fable 5.1 includes an Opus 5 fallback. Zapier says Opus handled 260 of 657 tasks after Fable refused them. Comparing that combined system with Google’s 51.3% Argon result as though both figures represented an unassisted model would obscure a material difference.

Original chart showing Argon ahead of Astra and Fable in DeepSWE, but behind both on FrontierSWE v2
Image: INSERT FUTURE · original graphic. Data: Google DeepMind, 30 Sep 2026. Different evaluation runs and settings.
03

Restricted access now, API pricing for later

Fairwind vets organizations for defensive cybersecurity work. Selected participants are getting Argon first. Google plans to expand access to developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. It has not announced a date for that broader rollout, including in the United States.

For developers, the API is the interface through which an application sends work to the model and receives its output. Google plans introductory rates of $2 per million input tokens and $10 per million output tokens. Tokens are the pieces of text a model reads or generates, rather than a fixed word count. Those rates will become $4 and $20 after an introductory period whose duration is unspecified. They are usage charges, not Gemini subscription prices.

Google separately says the output ceiling rises from 64,000 to one million tokens, allowing longer generated responses. That is an output limit, not a statement of how much input the model can accept or a guarantee that a longer answer is accurate. A public-access date remains unannounced.

Original diagram of selected Fairwind access, undated wider rollout and announced API rates of $2 and $10, later $4 and $20
Image: INSERT FUTURE · original graphic. Source: Google announcement, 30 Sep 2026. Announced future API rates, not subscription prices.
00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE