Google Unveils Gemini 4 Argon, Retaking Benchmark Lead Over OpenAI and Anthropic (venturebeat.com) 49
Google has unveiled Gemini 4 Argon, a new frontier AI model that it says leads or ties rivals on 13 of 18 disclosed benchmarks. The model is initially being released only to trusted cyber defenders and select pre-release testers, with broader availability planned later. VentureBeat reports: Google is not claiming that Gemini 4 Argon wins every benchmark. But across the benchmark table disclosed in the company's embargoed materials, Argon posts the highest score, or ties for the highest score, in more categories than GPT-6 Astra or Claude Opus 5.5. Across the 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one. GPT-6 Astra leads outright on three and ties Argon on one. Claude Opus 5.5 leads outright on two. That makes Argon the leading frontier model by total number of benchmark leads or top scores in Google's comparison set, even though the results show a still-close race in which OpenAI and Anthropic retain advantages in several important technical categories.
[...] The result is a more nuanced claim than "best model across the board." Argon appears to have the broadest top-score profile among the three frontier models in Google's disclosed comparison, but GPT-6 Astra remains ahead in several software, science-terminal and computer-use tasks, while Claude Opus 5.5 remains ahead in terminal-agent and post-training workflows. For enterprise buyers, that means model choice is still workload-dependent, even if Argon now gives Google its strongest claim yet to overall frontier leadership by benchmark count. [...] Argon also expands Google's output ceiling. The company says the model supports an industry-leading 1 million output tokens, up from a previous 64,000-token limit. That is a notable distinction for agentic software engineering, audit, migration and legal-review workloads where the value of a model often depends on how long it can sustain a chain of work before handing control back to a human.
[...] The result is a more nuanced claim than "best model across the board." Argon appears to have the broadest top-score profile among the three frontier models in Google's disclosed comparison, but GPT-6 Astra remains ahead in several software, science-terminal and computer-use tasks, while Claude Opus 5.5 remains ahead in terminal-agent and post-training workflows. For enterprise buyers, that means model choice is still workload-dependent, even if Argon now gives Google its strongest claim yet to overall frontier leadership by benchmark count. [...] Argon also expands Google's output ceiling. The company says the model supports an industry-leading 1 million output tokens, up from a previous 64,000-token limit. That is a notable distinction for agentic software engineering, audit, migration and legal-review workloads where the value of a model often depends on how long it can sustain a chain of work before handing control back to a human.
Benchmarks (Score:5, Insightful)
The trouble with benchmarks is that they are often gamed. So while this is interesting, one should remain a bit reserved, awaiting assessments more reflecting the way people use it.
Re: (Score:3)
Indeed. When tons of money is at stake, benchmarks are almost universally games and become meaningless. This is certainly the case here. Since most people do not understand the problem with benchmarks, they are still used to make claims that the benchmarks do not actually support.
Re: (Score:2)
They are always gamed. The moment a "secret" test set is sent to a closed model, it's no longer secret. So every benchmark is compromised.
I suspect they even train adapters daily on github just for something hard to game like SWE-rebench. If I was launching a trillion dollar IPO I certainly would.
Re: (Score:2)
Google said already in 2019 that they moved their focus from gaming to other areas.
Re: Benchmarks (Score:2)
zing !
Re: (Score:2)
Re: (Score:2)
-1, unactionable
The problem *without* benchmarks is that you have no idea how to compare anything to anything and you are left with the combinatorial explosion of trying every model with every task that you have.
It's #10 on LMArena leaderboard (Score:4, Interesting)
Let's let someone else do the benchmark, how about? [arena.ai]
Re: (Score:2)
Maybe in one of the agent categories? It seems poorer there. But under the category "Best Text Models", it's crushingly in the lead.
Re: (Score:2)
It's a 20 ELO lead in a category nobody cares about. Meh.
Re:It's #10 on LMArena leaderboard (Score:4, Insightful)
Most of the other frontiers are like 1 ELO apart, so yeah, 20 is a big lead.
Re: (Score:2)
Those aren't frontiers, look at it again. #2 is Opus 4.6 - it's just a category nobody cares about.
Re: (Score:2)
That does not actually remove the problem with benchmarks. They are far too easy to game and it happens all the time.
I like Gemini (Score:2, Interesting)
Re:I like Gemini (Score:4, Informative)
For all the shit it catches, I find it consistently orients against factual information
That's what I find too — Gemini orients itself against facts.
Every time I talk to it, more than 50% of the claims it makes are clearly false and of course it agrees every time you point this out.
Its citations are also almost never citations. So why even provide the links?
If Gemini is better at facts than other LLMs, then they're all shit.
Re: I like Gemini (Score:2)
Gemini disagrees with you.
Re: I like Gemini (Score:3)
It's impossible to tell since the models don't reason. That's why they can say obviously incorrect things with confidence.
Re: (Score:2)
You're using AI to try to "win" an argument on Slashdot because you were upset by someone recognizing your politics as idiotic nonsense.
Wait, now you're accusing me of generating my comments with AI? I thought you were kind of clever, but now I see you're only clever at trolling.
Re: I like Gemini (Score:2)
You mean what software told you, which is what you respect. That was my very point and you fell into it. Nice work.
Output tokens (Score:4, Insightful)
It's a race! (Score:5, Funny)
First one to pop the bubble is the "winner"
"Google Unveils Gemini 4 Argon" (Score:3)
No, the proper headline would be:
"Google Teases Gemini 4 Argon"
It's not been released, and it's not even clear if it will be.
Re: (Score:2)
So they essentially made some unsupported claims? Well.
Re: (Score:2)
"Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program"
This is why it will work (Score:2)
Gemini's confidence is unparalleled, that's one thing for sure.
The correctness of the answers is a whole other story.
Re: (Score:3)
Well, it certainly can replace management consultants, were "high confidence while being clueless" is basically the job description. But these people do not have positive effects on their customers or on society. Jenny Tian did sum it up very nicely in this skit: https://www.youtube.com/watch?... [youtube.com]
Lies, damned lies and benchmarks ... (Score:2)
It is still astonishing how many people do not know how meaningless and easily manipulated benchmark results actually are and not only in this space.
Re: (Score:2)
I have seen Gemini improving even in my own private benchmarks.
Re: (Score:2)
Gemini said my private benchmark improved by nearly an inch! I told my wife, but she didn't believe it....
google account gets banned for wrong question (Score:2)
If you make the mistake of asking the wrong question, your *whole google account* will get banned. [1] That should be a non-starter with most people.
[1] https://www.reddit.com/r/googl... [reddit.com]
Re: (Score:2)
That isn't what happened in your link. It makes no mention of content.
Re: (Score:2)
I'm waiting for Gemini Teladi. (Score:2)
Not a contest (Score:2)
Re: (Score:2)
Lol, a company dominated by Rationalists/Yudowski fans is "far left activists" now ;) The entire premise of HPOMR is that Lily Potter cast a spell to make her sister Petunia more attractive so she could marry a better husband, and when they adopted Harry, instead of having him attend public school, they homeschooled him, so he became brilliant. Very "woke" premise there ;) I don't know if you've heard, but the Rationalist community is the middle of a scandal right now for putting career-limiting pressure
Re: (Score:2)
Yes, ChatGPT always starting responses with "yes" seems to check out, or synonyms of it. But it uses it more in the sense of Japanese "Hai" (I heard what you said) than actual agreement.
Overshadowed (Score:3)
I'm sorry to say, this model could literally walk on water and I'm still not going to use it for anything serious. Gemini is sort of a go-to for 'web search+' type activities for me, but for agentic stuff, Google can f right off.
A while back I thought I'd try Gemini in an agent, so I downloaded Gemini CLI and hooked it up to my Google account. I got started, and then ran into usage limits - fair enough. I leave the CLI running, but completely back off, not touching it for a couple of days. Then I get the email from google "you have violated our terms of service, your account will be permanently blocked ... click here to appeal". Of course they couldn't tell me what I'd done wrong, so I appealed, and never heard anything back from them. I have no idea if my account is okay or not - mainly because Gemini CLI got deleted and replaced with Claude Code and OpenCode.
Google "inc" really need to make themselves a lot easier to deal with, otherwise everything they do will be entirely overshadowed by their corporate crap. To date, I've had no such issues from Anthropic or Deepseek - both of them "just work" and do what they're supposed to without getting in my way.
Pricing differences are MASSIVE (Score:2)
The entire FUD about safety is really about cost. OpenAI and Antropic are losing money on every token. Google's cost is a fraction while still making profits.
https://arena.ai/leaderboard/t... [arena.ai]
Google 4 Argon ranks #1 with 1525 score and $2/$10 pricing for 1M tokens.
Claude Fable 5 ranks #3 with 1505 score but $10/$50 pricing for 1M tokens.
OpenAI 5.6 Sol barely makes top 20 with a 1484 score with a $4/$20 for 1M tokens.
Sam and Dario appreciate anyone who wants to subsidize their investors by paying 5x token c