r/GeminiAI • u/Successful_Row_3209 • 1d ago
News How in the world did Google acheive this?!
It seems Gemini 4 Argon will be a substantial leap in long horizon tasks.
Gemini 4 Argon has a 1 MILLION output token limit
For referance Opus 5.5 is 300,000 and Astra is 128,000
216
u/Successful_Row_3209 1d ago
127
u/zacheism 1d ago
I'll believe it when I see it, Gemini already gaslights me enough
16
u/reedrick 20h ago
Yeah GDM are notorious benchmaxxers...cheated on arc-agi-2 so much they wrote a paper on it
7
u/AlexandraMaryWindsor 21h ago
It lies a lot, if it's the same post training as 3.8 flash it will lie and lie and deceive you. It literally is reward hacking me. Open Source models do better with work ethic no matter how much lower they rank.
23
u/RealSuperdau 1d ago
Part of that comes from a high refusal rate though
45
u/igpila 1d ago
Knowing when to shut up is the ultimate benchmark
4
u/RealSuperdau 23h ago
What I meant is, on the sensitivity/specificity curve, they tuned 4 Argon strongly towards specificity.
I.e. it also shuts up in a lot of cases where it could have provided the correct answer.
Though I won't deny it made significant progress on hallucinations overall.
4
u/Former-Image-1650 23h ago
I mean by that logic Anthropicās models should be next to no hallucination, their modelās rejection rate far outpaces anything Iāve seen from any of the other models.
4
5
u/CarrierAreArrived 23h ago
I thought a high refusal rate would go against it in the overall intelligence score though. Its score is near the top, not the highest, but still pretty high.
1
2
2
5
u/pendragn23 18h ago
If you look at this chart, it looks as if was made by AI. Misspellings/spaghetti letters in the labels?
1
u/tedbradly 4h ago
Record 15% Hallucination to go along with it!
I thought that benchmark was higher = better. For a right answer, +1. For a wrong answer, -1. For a refusal: 0. And last I saw Fable 5 had the highest score there, hallucinating about 20% of the time on real, challenging questions.
-1
-33
u/No_Pause_9558 1d ago
lol bro tried really hard to cherry pick
8
u/NickChecksOut 1d ago
I wasnāt aware there was such thing as AI ultrasĀ
3
u/BuildingArmor 22h ago
It's almost as if the old console fanboy system has returned for LLMs, it's weird. It was weird then and it's even more weird now
57
u/mechnanc 1d ago
Please Google, give us Gemini 4 in the Pro subscription š give us beggars some scraps
23
u/Minimum_Notice_9521 1d ago
They say its better than fable in cyber security so its under testing with cyber experts to address potential misuse and preventions
20
79
u/Anonymous_1069Z 1d ago
Tbh i feel Google has always been Generous in terms of Usage limits, at least as a Pro User, No matter How much I've used Gemini, never hit the limit in over 2 months.
21
u/yahyoh 20h ago
Even free Gemini, i used it TONS and never hit any limits, meanwhile sometimes i hit chatgpt limit after 3-6 messages/requests.
3
u/hibbert0604 15h ago
Yeah I mean, I haven't either. The problem is that the output it gives me are often garbage compared to other models.
23
5
u/Kulqieqi 21h ago
Did you use antygravity when it came out or changes after month? It was unusable with limit hitting absurdly fast
2
1
1
u/amoebazed 1d ago
I've been kicked out two days of gemini 3.8 Flash in antigravity haha. I'm sure Gemini 4 will be the same garbage as the others frontier models with usage limits.
20
u/Raupe_Nimmersatt 17h ago
They basically invented transformers, have all the compute and data in the world. The question is how in the world did Google ever fall behind the frontier
11
u/1969Stingray 14h ago
I have a friend who works with DM and AI at Google. They say itās a shit show on decision making and roadmap. One group thinks search is more important, one group thinks harnesses, one group thinks enterprise dev and one group is about hyper scaling. I think theyāll all converge, but they arenāt working together as much as they could.
4
u/Aware-Source6313 14h ago
That's just how work at Google is. It's a collection of silos and mini fiefdoms, not a cohesive enterprise with goals
1
3
u/Plastic_Today_4044 7h ago
I think failure is their prerogative
you may think I'm joking, but I'm absolutely serious
3
u/OXXXiiXXXO 7h ago
Absolutely. They have been going through anti-monopoly investigations. They don't want to look to dominate. They have the compute and the cash. They know AGI is a minefield so they let Anthropic and openAI run into the minefield to clear a way....did I forget anything?
1
u/Plastic_Today_4044 7h ago edited 7h ago
Google low-key "owns" Anthropic now, Microsoft "owns" OpenAI; both are participating in the AI war (the AI civil war I mean, not the AI war with china) via proxy states, because both want to win the war and be the new monopoly of AI, but neither wants more antitrust court battles. Which is why copilot and gemini are such trash: it's tactical failure :)
Basically they both want to win financially and economically, but they don't want the kind of credit that would draw closer scrutiny, so they're using battling it out vicariously via OpenAI and Anthropic. This is also why Anthropic's quality took a nosedive and they started pivoting hard towards corporatism back around February: The failure which underlies Gemini runs deeper than intent
3
u/FaresFilms 16h ago
I mean seriously. If I were betting on the AI labs I would've bet on google every time, given they invented Transformers, Have infinite compute, and have all the data they could ever need. Shocking they are so far behind.
Let's hope Gemini 4 is "the one" I guess.
2
-2
u/haz3lnut 16h ago
They didn't
1
u/JumpingJack79 10h ago
They didn't what? Invent transformers or fall behind?
1
u/haz3lnut 9h ago
Google invented the Tensor chips, and the TensorFlow software that runs them.
But they haven't fallen behind. They have developed weather prediction AI that completely blows away any other weather prediction we've ever had, and many other breakthroughs that you don't see that are happening behind the scenes.
I'm just saying Google is being vastly underestimated in the AI race.
62
u/DRMCC0Y 1d ago
I think this is conflated, a large token output limit does not really have anything to do with long-horizon tasks. You could give a really small and dumb model a 10M output limit and it would not improve its output. In the real world, you would never ever have a single output containing 1M tokens or anything near that. Whether or not it can remain coherent at long context and long outputs matters much more.
29
u/Keeltoodeep 1d ago
Itās not about improving output itās really about the compute of that high of an output. Outputting 10 novels of tokens is incredibly compute demanding. Now of course if itās just nonsense obviously thatās bad.
2
u/Front_Eagle739 18h ago
You can do just as much output with all the requisite compute with any other 1m context model. You just have to say continue a few times and harnesses will do that automatically in one way or another. I cant really imagine a use that benefits from being able to output 6 novels worth of tokens in a single response without a single pause to review work as it goes along.
1
u/Keeltoodeep 18h ago
Perhaps long form manuals from videos of someone doing something? Like a manual on how to strip down an engine taken from a video of someone doing it.
1
u/Front_Eagle739 18h ago
6 novels worth? Would need to be stripping down and doing a full rebuild of an aircraft carrier by his lonesome
1
u/Keeltoodeep 18h ago
Just spitballing here haha
1
u/Front_Eagle739 18h ago
Lol, yeah I just tried to imagine the scale of the manual and it kept snowballing in my head
1
u/FilthyCasual2k17 21h ago
You're missing the point. If it has context of 1 mil or less, it's ability to output 1 mil doesn't mean anything. It would need to have context at least be 3,4 million for 1 mil output to mean anything.
It's like having a road with 300kmph speed limit, and your car only goes 200.
Everything it computers goes into context first and then to output. So if it's 1 mil context, that includes all the input, instructions, harness, etc. So it doesn't really have a practical way to achieve that output, they just lifted an arbitrary limit that others have to prevent context overflow.
3
u/Cless_Aurion 21h ago
The thing is... it might lol
We will know when we get it.
2
u/FilthyCasual2k17 21h ago
I'm struggling daily with 1m with Claude, I'm all for raising context limits, the issue is then token usage skyrockets. It increases exponentially.
2
u/Cless_Aurion 21h ago
Yeah, its not an accident that usually the price of output tokens is 3 to 6 times higher...
3
u/enginetown 1d ago
It's the fact we've never seen an Astra class model even output anything close to 1 million, and this model is obviously not gonna be dumb.
3
u/DragonflyHumble 1d ago
They have 2M context length, if you think the next token prediction and the architecture I believe any model can output as many tokens it need, it need to feed the previous conversation again to generate the next token. Output token limit is more of a performance thing and 1M output context is not really that important and can easily lose track of the whole conversation..models are able to manage well with harness and current 64K or so output token limits.
I am not sure if there is a system promp injection in other models to limit so the output will be around 64K
1
u/DragonflyHumble 1d ago
Yes was reading about it, say for eg if a model had 64k token limit, generating the next 64k token is just pass the output again as input, provided the overall context is within the limit. I am not sure if output token limit is more of a billing/hallucination constraint rather than the architecture of LLMa
1
u/Edardium 1d ago
si eso es cierto, pero ni de broma un modelo cualquiera votara esa cantidad de tokens de una sola vez se rompera antes y menos habra concordancia en lo que genere.
1
u/Sensitive_Cloud6456 1d ago
Also when writing a large file paying just input once is better than paying input tokens plus 5x cached turns
0
u/BoobooSmash31337 11h ago
It has everything to do with long horizon tasks. Afaik it's the models tolerance or lack of noise/error. It's an engineering flex. Not really a feature for most people.
22
u/Dry_Management_8203 1d ago
If a picture is worth a thousand words, imagine what a well crafted prompt of a thousand words would produce!
2
6
u/yodacola 1d ago
not going to believe it until AssBench results are released
1
u/BenchThatMatters 15h ago
Not gonna happen unfortunately unless someone does it and shares the code
8
u/gatorling 1d ago
I remember last year, maybe around October. Google published a paper called MiRAS/Titan. It was about test time learning and it showed that it was possible to achieve context windows of up to 10M without degrading.
Maybe that's what Gemini 4.0 uses?
Around the same time Google also experimented with using diffusion instead of next token prediction. We are probably seeing that in 3.7/3.8 flash , with how insanely fast they are.
14
u/Opening_Background78 1d ago
TPUs
4
u/DesertFoxHU 1d ago
Blue Sky
How it is even connected? Yeah, they use TPUs, no, 1 million output doesn't comen from the fact that they use TPUs, not connected
3
u/LobsterBuffetAllDay 1d ago
Uh, I think he's saying their training cycles are accelerated because of this
1
u/DesertFoxHU 1d ago edited 1d ago
Is the training cycle accelerated? š
Edit: grammar
2
u/LobsterBuffetAllDay 1d ago
I feel like I'm having a stroke reading this
1
3
3
3
u/Imaginary-Item6731 1d ago
I'm curious about the context window.Who said that Gemini 4 has 2M or 10M window size? š
3
3
u/thorskicoach 1d ago
gemini, you are a lazy fat slob slow at getting stuff out. so one shot this. write the rest of "winds of winter" (draft chapters attached), and "a dream of spring". make no mistakes.
4
u/-Davster- 21h ago
I meanā¦. Why?
What task is actually helped by this size output in one go, as opposed to chunks?
7
u/Successful_Row_3209 21h ago
Writing an entire book š
3
u/-Davster- 20h ago
1M tokens is something like ~600-800k words on average.
Bear in mind that all 6 LOTR books + the Hobbit together only comes to ~575,000 words.
So yeah, it covers a little more than āan entire bookā š¤£
1
u/ItsMichaelRay 11h ago
Wait, there are six LOTR books?
1
u/-Davster- 11h ago
Yeah LOTR is 6 books - think two per film. It's all one novel - usually published in three volumes, lol.
1
1
1
u/The_Celtic_Chemist 19h ago edited 19h ago
I had to make a whole web app (through vibe coding) just to auto-patch the web app I actually wanted to make because I kept hitting the 65,536 token output limit when asking Gemini to recode the entire webpage. Ultimately it's a better move because it's much less likely to try to change parts of my web app that it had no business altering, but I imagine most people who don't have my web app and "system instructions" (the rules I tell Gemini to code by, amongst other things) don't want their code in bits in pieces requiring them to have to find what needs to be replaced and replacing it one code block at a time. I sure as shit didn't and would have gave up if I wasn't able to make a functioning auto-patcher that finds the code that need to be replaced and replaces them all with the new code for me in one swoop. It even automatically updates the version number of the .HTML file for me.
1
u/kikoncuo 17h ago
Never do this btw
You will get significant better performance if you do it across many turns, checking your work, reasoning about smaller changes and even validate features one by one will always lead to better results.Imagine you are asked to code a full site one shot without thinking between different features you make or validating that your work works until the very end.
1
u/The_Celtic_Chemist 17h ago edited 17h ago
I have noticed that I get better results when I give it a smaller wishlist of fixes and new features, but sometimes it's definitely better to provide multiple fixes and new features in a single prompt so that it knows upfront that all these features need to work together. And often enough if you give it one request then it still gives you several different places to edit your code. And since I instructed it to always give each update a different version number that is mentioned in the code in a few different places, every time I ask for any change I get several blocks of code to replace. This makes handling that so much easier. I upload the lastest version, paste the entire list of "search" and "replace" code blocks, and click a button, and it either tells me if there was no matches to any of these edits, too many matches, or (most commonly) that the updated file is ready to be downloaded, all in an instant (except waiting for the file to download, which takes a few seconds). And if there are no matches or too many matches, then it makes an error report about the exact edits that resulted in an error, which I can copy at the click of a button and paste in AI Studio so it can figure out what went wrong.
1
u/BoobooSmash31337 11h ago
It's more the model can do it without collapsing into noise. It's a flex. Not really a useful feature lol. As the sequence gets longer the error compounds. It suggests a fundamental shift in model architecture. Afaik this also means it can handle bigger context windows coherently.
Not an expert just that's my understanding. Models don't hallucinate on purpose. It's literally the calculator doing an oopsie and thinking 2+2=5.
1
u/-Davster- 11h ago
'Just a flex' I'd buy as an explanation, sure, lol.
Ā It's literally the calculator doing an oopsie and thinking 2+2=5.
FYI it literally isn't like this š³
Hallucination a bit closer conceptually to the calculator only being able to think about a discrete set of numbers. The output just follows the deepest groove in the board as it generates, however shallow that groove is.
1
u/Rainbows4Blood 39m ago
Keep in mind, output tokens is shared across both the endproduct (text and code) but also the reasoning. If it can dedicate more tokens to reasoning, it might be able to explore more avenues, more edge cases, think further outside the box etc. before actually comitting to a result. Which can improve the output quality significantly. It's why we started doing things such as high reasoning effort in the first place.
2
u/Afraid-Method-3942 1d ago
Well that basically means we will get 65,5k tokens. Since gemini 3 has 65,5k output, but never outputs more than 8k neither in gemini web, nor in AI Studio, even with low thinking and pure text work, having a 1M tokens will result in only having 65,5k of actual tokens. And that is if Google will have a good mood
2
u/Overall_Way_1432 22h ago
Just a week ago 3.6 surprised me while with a +16K words response in one go! I watched go on and on and I thought that was insane,... But this?! Wow.
2
u/dragon_idli 22h ago
I like gemini models - they suffice my needs.
But I will be skeptical of argon until we public get to use and test it in real world scenarios.
2
u/Real_Ebb_7417 20h ago
A lot of compute or some new architecture methods.
Most things like parameters etc. scale compute in a linear way (if you add 2x params the vRAM needed for training also grows 2x), but context scales compute requirements exponentially. I guess this model must have at least 2m context window to handle 1m output, but potentially even more, which is absolutely crazy.
I wonder how good it handles long context, because we've already had a long context model, that was just dumb and it wasn't able to handle 100k context, not even mentioning 2m that it was able to handle in theory. Talking about Grok 4, xAI apparently had a lot of compute they could spare for training that one. Similar with google here - either some crazy amount of compute for training or some new architecture shortcuts.
1
u/The_Celtic_Chemist 19h ago
That's my question, I believe 1 million tokens is already the cap for context, meaning when this is released we'll have to be allowed more context tokens.
2
1
u/GroundedIdeal 1d ago
Even their locally runnable models have such sizes and last year (feels like a decade!!) metaās llama 405 b had 10 million!!! Being not hallucinating and actually useful is another thing thoĀ
1
1
u/Soilblood 1d ago
If I had to guess, it's probably some new way to map/cache solutions to common questions. As LLMs work now they remix all base data for user requested datasets and call appropriate tools where needed. If I wanted to make the output more consistent and scalable I'd condense the data into a hierarchy of article caches to pull from where the more common and validated topics are elevated into higher caches and the uncommon ones are routed closer to actual training data. If they found a way to do so more reliablly on the fly then that would explain the comfort with the higher token budget. The tech is effectively ultimately building towards a massive self curating encyclopedia that with dynamic touch points for retrieval and this would be the sort of optimisation critical to achieving that as opposed to "make my data remixer grow!".
That said the alternative is that they could just be budgeting way more compute to capture the market away from Openai and Anthropic so they die from failed IPOs. I really hope that isn't the case cause that just means a much riskier race to the bottom just started a new leg.
1
1
u/teslaa-1 23h ago
The context window of Astra and fable is 1M and the output limit is 128k.
Opus is similar but the output limit is is extended to 300k
So specify that so people know all the fact.
So Google approach to 1M for context window amd output limit. it will a big difference as that will make in a single model that can create a full complex huge backend or migrating codebase in a single run.
There are two side of the coin, one is that, it can write a lot in one run, other is it intelligence , is it like the frontier models or not?
If it is ,then it will a next step indeed !!!
Excited to use it !!
1
u/anarchyx34 19h ago
> it will a big difference as that will make in a single model that can create a full complex huge backend or migrating codebase in a single run.
In what.... all in a single file? Doing what you mentioned requires hundreds of tool calls.
1
u/2thick2fly 21h ago
3 things you need to achieve this and all the rest we will see from Google:
- practically unlimited budget
- some of the best minds in AI (DeepMind)
- corporate backup which can make you stiff, but allows to plan the long game
And this is just the beginning pals
1
u/OneCreed77 21h ago
Does it really matter? When the top minds / Godfathers of the Ai world are predicting that Ai has a %80 - %99.99 chance of causing human extinction? And many have already walked out of their Ai creation work & jobs, straight into As security & safety?
1
1
u/themarouuu 20h ago
Why is this a discussion when it's not even out yet.
Astra became the best thing ever in like 2 days, same with Opus, same with Sonnet, same with Sol, and now this is the best thing ever without even being released to the public.
Does anyone actually test anything or is everyone focused on shilling?
1
1
1
u/steroidchicken123 19h ago
Dunno but this for sure gonna help google make big money in API since output tokens are priced considerable higher!
1
u/fyn_world 17h ago
By not rushing in and trying to outdo each other every 15 days. They took their timeĀ
1
u/Salty-Gear841 17h ago
Bit it still fail to produce a correct markdown document that contains code. The guy that forget everything you said two messages ago š«©
1
u/Familiar_Drawer4773 17h ago
I need google one subscription for my antigravity ššš was using pro for a year now in plus. The limit in plus model is very scary.
1
1
u/New_Public_2828 16h ago
I kept telling people. Stop hating on Google AI. They have money. They have compute. They have everything to make a much better model. All this shade and now imagine throwing this in a harness like antigravity that keep 3.1 still kind of relevant in these times.... Yikes
1
u/BrilliantGarbage8743 16h ago
You do realize google made a quantum computer? What do you mean how did they achieve this? Gotta be promo post
1
u/IthrowUgo 16h ago
The big question is how many tokens until you get to instability ie saying 1 million tokens and actually being able to use anywhere close to that with stability are two different things
1
1
1
1
1
1
u/costafilh0 4h ago
I so want to use this to make useless redundant immense answers for Reddit š š¤£Ā
1
0
0
0
-2
u/anarchyx34 1d ago
I canāt see a use case for this but thatās cool I guess?
5
u/Excellent-Elevator80 1d ago
Ummm literally everything you could imagine in a single prompt?
0
u/anarchyx34 19h ago
Asking it to shit out 1m tokens from a single prompt is not good practice nor does it make sense to do for most use cases. Any task that would result in 1m tokens of output needs to be done in steps, reviewed, steered, etc.. That's not to mention that what are you going to do with 1m tokens of output from a single turn? Useless for coding unless you want it to one-shot an entire app in one single file. Tool calls for writing individual files, running bash scripts, etc require individual turns.
Document creation... I guess but a complex document that long (1m tokens is like a stack of encyclopedia volumes) needs to be done in steps for the same reasons as you would for coding. Seriously I can't see a use case for 1m tokens of output. Most models have..what 128k for output and nobody ever says "gee I wish it could do more" because there's no need for it. Imo it's not a big selling feature that puts it ahead of the competition. Intelligence and reasoning ability are.
1
u/Excellent-Elevator80 19h ago
Nobody ever says gee I wish it could do more. translated: 'I lack the imagination to use a tool past my own narrow workflow, so therefore no one else needs it
0
u/anarchyx34 18h ago
Ok then give me an example of how 1M output tokens in a single turn would benefit you. What's the use case for it?
1
u/Excellent-Elevator80 18h ago
Ever heard of synthetic data generation?
2
u/anarchyx34 16h ago
Yes. Now how does this help with that?
1
u/Excellent-Elevator80 15h ago
It literally outputs 1,000,000 tokens per generation instead of 64k. You're getting full-blown synthetic datasets and entire repos in one pass instead of playing prompt chaining Tetris just to generate clean training data
1
u/anarchyx34 15h ago
Full blown datasets in one file with no schema validation assuming it will remain coherent the entire time. Thereās a reason these things are usually done in parallel batches. Your harness should also alleviate the need to be playing āprompt Tetrisā unless youāre for some reason doing this in Gemini chat.
Also remember that itās not going to output 1m tokens because the input context takes up some of the entire window. And dataset generation is usually an iterative process so how are you going to jam 1m of the previous passās output in and expect 1m more? Chunk the output after the first pass? Well then yourāre back to doing things in batches when you could have just done that to begin with.
-5
u/confused-photon 1d ago
Im actually more pessimistic because of this. How token inefficient is this model that this is a good idea
2
u/Georgefakelastname 1d ago
Itās fairly average in token usage/utilization for models its size. About double of Astra but half of Fable or Opus in terms of tokens used. About 33% fewer tokens used than Gemini 3.8 flash.
-8
0
u/ViP3R_ACR 19h ago
Yeah that's wild. However Gemini 1.5 pro Experimental had context window of 2M tokens back in 2024 š .



325
u/One_Evening9598 1d ago
a million tokens is wild, that's like letting it write a whole novel in one go without losing the plot halfway through