This is the English edition. 한국어판 and 日本語版 are also available.

Gemini 4 Argon Benchmarks: Nine Rumors Fact-Checked

2026-10-03 · AI · United States · Zoogom Editorial

#Google#Gemini 4 Argon#AI benchmarks#AI rumors#Fact check

A conceptual analyst sorting AI benchmark evidence into verified unresolved and contradicted groups

Gemini 4 Argon is not a leaked codename or a fictional model. Google announced it on September 30, 2026. The confusion starts when that confirmed fact is bundled with claims that Argon wins every benchmark, eliminates hallucinations, costs a fraction of GPT-6 Astra, and is already available to ordinary users.

This article checks the public record as of October 3 in the United States. It compares Google’s announcement with independent evaluations, Arena results and a professional-work benchmark. The regional section describes a sample of public Korean-language, US English-language and Japanese-language X discussion. It is not a poll of verified national populations, and it is not a hands-on Argon review.

The opening image is a newly generated editorial concept about evidence review. It is not a Google interface, an X post, a benchmark screenshot or a photograph of a real evaluation.

The short verdict

The defensible description is “a restricted frontier model that leads several important evaluations.” “The undisputed best model that anyone can use today” is not supported.

What the benchmark record actually shows

Google’s numbers are provider-reported evidence

Google’s official announcement reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, 91.7% on LVBench and 68% on CWE-bench v1. These tests concern long-horizon software engineering, multi-application automation, long-video understanding and vulnerability repair.

Those are material results, but they were selected and published by the model provider. A first-place DeepSWE score does not establish first place in all coding. Percentages from tests with different tasks, harnesses and pass criteria are not interchangeable.

Vals shows a strong professional index and weak computer use

On the Vals model page, Argon ranks first of 41 on the Vals Index at 68.90%±0.97. It also places first on Finance Agent v2 at 65.40%, second on Vibe Code Bench v1.1 at 91.91%, and second on Code Migration at 68.17%.

The same page puts it fifth on Terminal-Bench 4.0 at 57.58% and seventh of eight on CUA-bench at 4.83%. Vals configured a 262,144-token maximum output for its runs, so its results do not demonstrate the full one-million-output specification either. The coherent reading is that Argon performs very well in an aggregate of professional tasks and selected coding work, while terminal execution and interface control remain less convincing.

Intelligence 53, hallucination 15%, accuracy 50%

Artificial Analysis scores Argon at 53 on its Intelligence Index, matching GPT-6 Astra at its maximum setting. Under the launch discount, it estimates about $1.99 per Intelligence Index task. Argon produced roughly 62,000 output tokens per task in that evaluation, compared with about 27,000 for Astra.

The AA-Omniscience result is easy to oversell. Argon’s hallucination rate is about 15%, but its accuracy is about 50%. The low error figure partly reflects a greater willingness to acknowledge uncertainty instead of guessing. That behavior can be valuable in research or compliance work. It can be a disadvantage when a workflow requires high answer coverage. “Less likely to invent an answer” is not the same as “knows every answer.”

Professional agents and human preference tell different stories

On Mercor APEX-Agents, Argon leads at 82.2%±4.4 Pass@1. The published category results are 90.3% for consulting, 80.9% for investment banking and 75.3% for corporate law. That is important evidence for long professional assignments, not proof of universal office productivity outside this harness.

Arena’s Text and WebDev result places Argon first in Text Arena at 1,525 and eighth in Code Arena: WebDev at 1,679. One announcement therefore contains both a major win and a middling coding rank. Arena’s separate Agent result places it eighth in a preliminary analysis of roughly 3,000 sessions. Preliminary preference data should not be treated as a stable production benchmark.

The early PublicAI aggregate listed Argon fourth overall, first in human preference and professional work, and twenty-fourth of 181 in reasoning. Only five of 19 boards were represented at that point. It was an early snapshot, not a permanent consensus ranking.

Rumor 1: Argon is already generally available

Verdict: false or materially overstated

An announcement is not the same as an unrestricted product launch. Confirmed initial access is centered on trusted cyber defenders and evaluators through Fairwind. Google says paid API customers and Google AI Ultra subscribers are intended starting audiences for a broader rollout, but it gives no calendar date.

A US subscription label therefore does not establish access. Buyers should check the actual model picker, API model identifier, account eligibility and regional documentation before budgeting a project around Argon.

Rumor 2: It beat every GPT and Claude model

Verdict: false

Argon leads DeepSWE, the Vals Index, APEX-Agents and Text Arena. It does not lead Vals Terminal-Bench or CUA-bench, Arena WebDev or the preliminary Agent Arena result. Artificial Analysis ties it with Astra on the overall index, while the early PublicAI aggregate places it fourth.

Counting the number of highlighted rows on a launch chart is not a valid universal ranking. The result depends on the task mix, competitor settings, tool access, time budget and pass criteria. A procurement memo should identify the evaluation relevant to the proposed workload instead of announcing a single global winner.

Rumor 3: A 15% hallucination rate solves hallucinations

Verdict: misleading

The lower rate is encouraging, especially for workflows in which an unsupported answer is worse than no answer. The companion accuracy result shows why it is not a cure. Argon appears more willing to abstain when uncertain.

An enterprise pilot should measure at least three outcomes separately: correct answers, unsupported answers and explicit abstentions. Combining the last two into a single success or failure number can make one model look better for the wrong operational reason.

Rumor 4: Argon is 40% cheaper than Astra, or one-fifth the price

Verdict: conditional and often apples-to-oranges

At the introductory $2 input and $10 output rates, Artificial Analysis estimates about $1.99 per task for Argon versus roughly $3.26 for Astra, a reduction of about 40%. That is a valid result within that test and promotional price.

After the discount, the announced $4 and $20 rates would put the same estimated Argon task near $3.98, roughly 20% above Astra. Argon’s longer output also affects completed-task cost. Token price, benchmark task cost and the cost of an accepted work product are three different measurements.

US teams should include retries, tool fees, review time and failed runs in a cost comparison. A lower price per token does not guarantee a lower cost per approved deliverable.

Rumor 5: Anyone can reliably emit one million tokens in one call today

Verdict: the specification is confirmed; ordinary use is not

Google explicitly presents one million tokens as the maximum output limit, up from 64K. General API access is not available, and early evaluators used smaller configured limits.

A production service may expose continuation behavior, request timeouts, quotas, storage costs and recovery rules around a long generation. The ability of the model architecture to support the ceiling does not establish that an ordinary account can request, receive and retain a million-token response today.

Rumor 6: Google’s own employees concluded Argon failed

Verdict: a credible report exists, but the conclusion is unproven

Madison Mills’ Axios account notes that Bloomberg, citing anonymous sources, reported internal concerns about performance in testing. Google told Bloomberg the characterization was inaccurate. It is fair to report the existence of that disagreement. It is not fair to turn an anonymous concern into a verified failed test result.

The useful follow-up is a reproducible comparison after access broadens: the same repository, permissions, time budget, tests and acceptance criteria. Neither the report nor the corporate denial substitutes for that evidence.

Rumor 7: Velox or a disguised Gemini 3.8 checkpoint was Argon

Verdict: unverified

X and Reddit speculation connected unusually strong anonymous Arena entries, the name Velox and a supposed Gemini 3.8 Flash checkpoint to Argon. No public Google or Arena statement identified those entries as the released model. Pre-release outputs from an anonymous system should not be presented as reproducible Argon evidence.

Rumor 8: The public version will definitely be severely nerfed

Verdict: unsupported forecast

The eventual public route may use different reasoning settings, latency targets, quotas or safeguards. None has been announced in enough detail to quantify a downgrade. Past differences among internal, API and app experiences can motivate a hypothesis, but they do not prove what will happen to this model.

When access arrives, record the model ID, date, reasoning level, tool permissions and account tier. A comparison without those details cannot establish a nerf.

Rumor 9: Gemini 3.5 Pro was canceled and Neon or Helium are confirmed

Verdict: lineage speculation

The expected Gemini 3.5 Pro did not receive a general launch before Gemini 4 was announced. Google has not publicly confirmed that the project was formally canceled, merely renamed Argon, or reorganized into products called Neon and Helium. The safe statement is limited to the observed release sequence.

How the three language communities framed the news

In Korean-language X posts, several headline numbers were often compressed into one update: DeepSWE 77.9%, the Vals lead, one million output tokens, and the $2/$10 price. That format is efficient, but it can lose the promotional-price condition, lower rankings and restricted access.

US English-language discussion was more method-heavy and adversarial. Users debated the Artificial Analysis cost model, output volume, Arena splits and the internal-skepticism report. The strongest question was not whether Argon had impressive scores, but whether those scores would reproduce after the promotion and outside selected access.

Japanese-language discussion paid particular attention to whether the output ceiling could change work on contracts, filings and codebases. Many posts paired excitement with the caveat that the writer had not used the model. A recurring counter-rumor assumed the public model would be downgraded, although no Argon-specific evidence established that outcome.

These are observations from public language-based samples shaped by search terms, recommendation systems and timing. They are not national opinion shares and should not be quoted as measured support or opposition in Korea, the United States or Japan.

A due-diligence checklist for US teams

  1. Record the exact model ID, version date and serving route.
  2. Hold reasoning level, tool permissions and time budget constant across models.
  3. Separate correct answers, unsupported answers and abstentions.
  4. Calculate cost per accepted result, including retries and human review.
  5. Test a real repository or document set with sensitive data removed first.
  6. Measure recovery from timeouts and interrupted long generations.
  7. Recheck prices after the introductory period and before signing a volume commitment.

The missing evidence is no longer another launch chart. It is ordinary-user data on latency, quotas, failure rates, completed repository tasks and review effort under a documented configuration.

Bottom line

Gemini 4 Argon gives Google a credible frontier result. The evidence is especially strong for professional agents, extended coding work, text preference and long output. The same record also contains weaker terminal, web-development, agent and computer-use rankings, plus an unresolved access question.

The accurate headline is that Argon leads several consequential evaluations and introduces an unusually large output specification at an aggressive promotional price. The inaccurate headline is that it beats every rival, fixes hallucinations and is broadly available for less money. Public API behavior and cost per accepted task will determine the next verdict.

Source use and editorial disclosure

This article independently reconstructs dates, published specifications, selected scores and ranks, and methodological limits from the linked material. It does not reproduce source prose, tables, charts, benchmark screens, X post images, logos, product interfaces, or source artwork. The Axios report is attributed to Madison Mills and is used only to establish that an anonymously sourced report and Google’s rebuttal existed; neither side is adopted as a verified performance result. Readers should use the linked originals for each evaluation’s complete method, sample, configuration, and current status.

Source: Google · Includes original screenshots or graphics