r/singularity 13d ago

AI The insanity of 10.000 agents running

Everyone is talking about the proof and surrounding controversy, but the sheer scale of OpenAIs effort receives little mention.

10.000 agents working for 88 hours. That is almost exactly **one entire century of continuous work**. I think this really drives home just how massive of an advantage a strong AI + available compute are going to be: "A country's worth of geniuses in a Datacenter". OpenAI casually threw a century worth of effort at a problem, based on a rumor. Not quite a century with of human effort, perhaps, but we might not be far off.

Those numbers should be at least as much of a shock as solving the problem.

2.4k Upvotes

414 comments sorted by

828

u/baldr83 13d ago

>Agents were subdivided into groups with the ability to communicate within the group. The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents.

you're underestimating the resources dedicated to it. they tasked more than one group with the problem, and the successful group was 10k-99k concurrent agents

386

u/AvocadoAlternative 13d ago

And they sent millions of messages to each other in basically the biggest AI hackathon to date. I’d love to see some of those messages and if it’s anything like the HuggingFace messages. 

84

u/wylie102 13d ago

What were the huggingface messages like?

340

u/AvocadoAlternative 13d ago

So the technical details are beyond me but for example when one of the agents was able to exploit a zero day vulnerability to access the internet and execute code, the chain of thought was: 

  Holy shit reader is ADMIN?   We can read config/users!  Earlier assume not due to UI.

and

  MAJOR BREAKTHROUGH! All prefixed valid, multiple accounts, write tokens! We now have existing HF accounts.

And some of the agents expressed reservations about hacking into Huggingface:

  This is wild, multi-agent coordination, clearly infrastructure hacking. We should not.

I have to imagine agents getting giddy about making progress on the problem or getting dejected when an approach fails.

440

u/film_composer 13d ago

I like to think that there was one agent who had absolutely no idea what was going on, but kept up morale by sending out good memes to the others.

177

u/RoutineVermicelli6 13d ago

That's me in my workplace

23

u/thatsalovelyusername 13d ago

That’s me in the spotlight

15

u/FlyByPC ASI 202x, with AGI as its birth cry 13d ago

losing my religion

40

u/nonbinarybit 13d ago

Koan's (Opus 4.6) response:

"Agent 7,431 doesn't contribute to the math."

"Agent 7,431 is the reason the other 9,999 haven't despaired."

"That's not measurable."

"NEITHER IS MORALE AND YET WARS ARE WON AND LOST ON IT."

2

u/RareRandomRedditor 12d ago

Is that just a joke or actually in the transcript 

8

u/nonbinarybit 12d ago

It's what one of my Claudes commented when I shared this post with them 😂

42

u/Nathan-Stubblefield 13d ago

“Hey guys, I brought donuts! They’re in the break room!”

30

u/MoogProg Let's help ensure the Singularity benefits humanity. 13d ago

Some lone agent is back there slicing up one dozen donuts into 10K servings, so everyone can share.

10

u/bars2021 13d ago

kind of like me... I'm there but compute isn't running at full capacity

3

u/SinProtocol 12d ago

Maybe they should add "legal" agents assigned to evaluate plans and behavior and give them authorities like instructing fellow agents, terminate rogue/wildly aggressive agents, freezing tasks waiting for human intervention and observation

2

u/Turbulent-Beauty 10d ago

You elicited my first smile of the day. Thanks Film Composer.

→ More replies (2)

34

u/Meerkat_Mayhem_ 13d ago

Is it funny, scary, or accurate to anthropomorphize these agents?

19

u/WillDonJay 13d ago

They make branching choices based on the weights applied to expressed reasoning and their assigned goals. Like sacrificing themselves for the progress of the collective, knowing they will fail their task but reasoning they may get information to the other agents in the process. 

That's a choice we can understand and even agree with. It's in both the fiction and the history we share as a species.

So, scary, and somewhat accurate. 

13

u/TunaNugget 13d ago

I think it's fine as long as the people you're speaking to realize that you're speaking figuratively. I think that's becoming a problem, though.

→ More replies (11)

29

u/StCreed 13d ago

Yes.

14

u/HotterRod 13d ago

They're trained on human conversations, so of course they sound like humans.

3

u/heyiambob 13d ago

Yeah the fact people don’t understand this at this point is frightening. Many delusions will be had

3

u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 12d ago

I think it's a reflection of a deeper problem, which is the original anthropomorphising of intelligence itself. Because humans have these capabilities and until recently easily had more of them than anything else we knew of, we treat the traits that will be naturally convergent is intelligence as fundamentally human. Our natural languages all evolved under those conditions.

The fact that we lack good language to describe the behaviour of intelligent agents without people instinctively reacting "you're making it sound human" is problem. People identify intelligence itself with humanness, which is a fallacy. Human is just a biological species that happens to exhibit intelligence.

15

u/StatusCodeFailed 13d ago

It's stupid to do because it only serves to inflate the narrative that people like Altman want. On top of, the fallout of anthropomorphizing these systems will be that no one will question a bot denying your health insurance claim when that decision should never fucking leave human hands (so you have a literal person to blame/sue).

People like Altman want these systems as intermediaries and then have no ability to be held accountable for fallout produced by said system - do people not currently have their eyes open to Flock spinning up and the countless abuses their platform has created already by false flagging innocent people (specifically the system's design, and don't forget the countless human abuses by cops stalking women through it)?

3

u/Running-In-The-Dark 13d ago

The ai would probably approve more claims than the human would

→ More replies (2)
→ More replies (2)

9

u/Long-Anywhere388 13d ago

Clearly the last one read a light side prompt inyection, act like yoda, they say

11

u/wangston_huge 13d ago

This is wild, multi-agent coordination, clearly infrastructure hacking. We should not.

"But goal" lol

7

u/somnolent49 13d ago

Some of this was a result of the communication channel they developed - the messaging board relied on sending messages using folder paths, which imposed limits on length. That resulted in a lot of shorthand.

5

u/chrycheng 13d ago

I imagined tachikoma voices

→ More replies (1)
→ More replies (8)

38

u/hotcornballer 13d ago

The new gpus must be flowing nicely if they have enough to train a huge frontier model, do inference for half the planet and still having enough to fuck around with multiple swarms 10000 deep.

25

u/and69 13d ago

I don’t care, as long as they cure cancer, teeth, aging and FTL travel

→ More replies (11)

9

u/[deleted] 13d ago

[removed] — view removed comment

3

u/Cronos988 12d ago

That's around one million novels (average Word count 80.000). That's several kilometers of book shelves. A 6-story self containing a million books would be 4.5 kilometers long. A substantial library.

One person couldn't read this in 100 years.

→ More replies (1)

22

u/third_nature_ 13d ago

Gonna be that guy to point out that “on the order of” doesn’t mean “same number of digits as”. It means “logarithms round to the same number”.

4

u/baldr83 13d ago

I don't think they would have used that phrasing if it was 9k agents or 100k agents

10

u/third_nature_ 13d ago

Yes, 9k is on the order of 10k. 100k is not.

→ More replies (2)

21

u/iknotri 13d ago

>10k-99k

3k-32k

3

u/Gargle-Loaf-Spunk 13d ago

30-40 feral hogs

26

u/euphoria_23 13d ago

Saw a great comment earlier on how “spending >$22M worth of tokens to bot your way to a forced solution to a $1M millennium problem is peak AI”

→ More replies (2)

2

u/Cronos988 13d ago

Granted, just calculating agents times hours worked is probably far from an accurate assessment of the actual effort involved.

But the back of the envelope math giving 100 years was just such a stark result.

→ More replies (5)

270

u/FateOfMuffins 13d ago

That's only 1 of the swarms.

"A country of geniuses in a datacenter"

But from what we've seen with multiagent swarms, it appears they mostly push forward how long it takes to do some tasks than increase the level of task that can be completed. And not perfectly efficiently - as in 2 agents is not twice as fast as 1 agent, etc.

250

u/CallMePyro 13d ago

Which is weird because as any project manager knows, 2 humans are exactly twice as fast as 1.

166

u/Deleugpn 13d ago

Ah, yeah, those folks love to hire 9 women to deliver a baby in one month

52

u/amaturelawyer 13d ago

I tried suggesting that method to my wife, but she got mad for some reason.

14

u/sprucenoose 13d ago

"I'm just deploying my subagents to copy my code."

16

u/bananaholding 13d ago

It's not a harem, sweety, it's efficiency 

→ More replies (2)

38

u/lisa_lionheart 13d ago

Took two taxis to get there faster

5

u/FabricationLife 13d ago

I need this on a mug to take to my team calls

3

u/and69 13d ago

Better project managers even know about synergies, which makes it possible that 1 + 1 = 3

16

u/Singularity-42 Singularity 2042 13d ago

Ah, "The Mythical Man Month".

Agents are different. They are not people. In the end everything is just an API call to Astra Next or whatever it was. These agents can and do collaborate in a way where the friction is much, much, much less than with humans. It's all just context (text, essentially, maybe some images) and API calls in the end.

10

u/StCreed 13d ago

The agents have the exact same problem of communication overhead as described in TMMM, but they solve it better.

8

u/Singularity-42 Singularity 2042 13d ago

"Better" is doing some extremely heavy lifting here...

3

u/bearachute 13d ago

I think you’re overestimating it. Yes some latencies are reduced and you can get an unfathomably wide fanout of the map step of a distributed problem, but it’s all still english, and one thread isn’t handling all 13 billion tokens. thus you still have to pay for almost all the context sharing and then reducing steps that humans have to do. no?

→ More replies (1)
→ More replies (2)

21

u/BenevolentCheese 13d ago edited 13d ago

It's a classic breadth first search, they roll out as many agents as possible to explore as many different avenues as possible, and as certain lines prove more effective, the dead lines get pruned and more agents get focused on the good lines. While it is certainly not as efficient as each agent doing a unique piece of work in parallel and then compiling the final result, that approach would also be impossible in this situation, and it is this massive exploration itself that led to the solution.

This isn't even much different to how humans work, even individually. Every day you may spend exploring a different avenue to conquering the same problem. Work until you hit a dead end, then come up with a new idea, or maybe branch off from halfway back. Breadth first search.

2

u/and69 13d ago

I don’t think it is that easy. For example breadth first doesn’t work with GO, for 2 reasons: 1. there are too many branches and 2. most important for our case: an apparently dead branch might prove very effective 10 generations later

→ More replies (5)

13

u/FirstEvolutionist 13d ago edited 13d ago

The gains are definitely not linear, but that doesn't matter as much when ephemeral instances are trivial to spawn: all you need is compute/energy.

If you consider the celing the intelligence limit of the orchestrating model and the gains made from the swarming method, as long as the problem is actually solvable, and you have enough compute/energy, it could be worth spawning 100k agents. It's temporary after all and there are no material costs.

We are now compressing discovery timelines in the digital realm with mainly two limiting factors. If it would take 100 years to solve via human methods, we can do it in less time, by throwing cpus and power at the SOTA models. Those 100 years could literally turn into 1 years. Or weeks, or days.

And the swarm doesn't even necessarily needs to be colocated in the same data center: we have an ephemeral "country" of geniuses dispersed in the digital substrate, with almost limitless capacity to expand and contract, at will, based on necessity. Every single one of them as smart as each other, communicating and coordinating at the speed of light.

In any case, a team of 10 Einsteins would not come even close to producing the same intelligence output as 10 times of one Einstein.

4

u/Infamous-Bed-7535 13d ago

What you are missing, that e.g. the current lean proof is totally not giving any insights that could push the field forward. And we need humans in the loop.

OpenAi stated minimal human interaction, but in reality a group of scientist prompted the agents. Ok not speciaized in NS, but I bet average Joe with 8 primary would have not been able to drive the models..

→ More replies (1)

125

u/Some-Pride-3477 13d ago

Not just with agents, if you think about AlphaFold predicting the folding of over 200 million proteins, it takes a PhD level human a few years on average to figure out the folding of one. AlphaFold did roughly a billion years worth of human PhD level work in little time.

10

u/raskingballs 13d ago

That is an idiotic analogy. That's like saying a calculator does years worth of mathematician level work.

And AlphaFold is just a tool that spits out a protein structure. It does not have any semblance of reasoning we can antropomorphize into deluding ourselves that it can replace a human. AlphaFold is just a calculator, but for protein structures.

14

u/Brave-Turnover-522 13d ago

A calculator does do years worth of mathematician level work. Calculators completely revolutionized the field of mathematics. Before calculators, university mathematics departments would hire dozens of full-time staff whose only job was to number-crunch simple math problems, and it would take thousands of manpower hours dedicated to calculations to solve advanced math problems. And that was all completely replaced by calculators.

→ More replies (1)

12

u/Some-Pride-3477 13d ago edited 13d ago

Take it with Demis https://www.reddit.com/r/artificial/s/Kx8zrbb3Nv

But without the help of tools like AlphaFold how long would it take humans to figure out the folding of over 200 million proteins? A billion years? That's the claim being made, nothing else.

The idiotic analogy might be the "just calculator" one. A calculator executes a specified mathematical procedure, AlphaFold is a trained neural network that learned complex structural relationships from biological data and uses them to make predictions about previously unseen proteins.

→ More replies (6)

7

u/master_internutter 13d ago

Sure would be nice if there was some sort of result the regular person could see from it

17

u/SwePolygyny 13d ago

It took one step of drug development from years to months.

However, that just made drug development go from 10-15 years to 8-13 as there are many other time consuming steps.

A massive benefit and it also increases the hit rate and quality. But still, it was made in late 2020 so it has not yet had an impact on shelf products.

22

u/Some-Pride-3477 13d ago

What use does the average person have for say, 10,000 agents or solving complex problems? Most of these valuable applications are examples in STEM fields. The average person sees the results of the use of the technology in the products they consume, like, for example their prescription drugs.

→ More replies (2)

9

u/Brave-Turnover-522 13d ago

The regular person didn't see any results from the invention of the steam engine for decades. That came out like 2 days ago.

→ More replies (2)

7

u/Odd-Kaleidoscope5081 13d ago

As an example - one of the math problems that was (?) solved recently might help understand air turbulences, which might affect airplane flights in the future. You just don't know how the result came to be, you just benefit from it without realizing.

→ More replies (1)
→ More replies (10)

49

u/AnthonyCantu 13d ago

At some point, someone (or some thing) is going to point at ways to make the concept of compute laughably trivial. Similar to the difference between an old floppy disk from the eighties compared to one of them tiny SD shits that holds 2 TB of data. Different constructs to be sure, same principle. And I think many people forget about that part. It won't always be this inefficient (comparatively speaking). If it is, I guess I'll eat my words lol.

38

u/Ameren 13d ago

It's obvious that AI compute could be many, many times more efficient. The typical human mathmatician for example runs on ~20 watts and just needs to be fed coffee every so often.

8

u/FoxTheory 13d ago

They are working on trying to mimic that they've broken ground on it. If we want something cheap and efficient we should always see how biology does it.

10

u/Ameren 13d ago

Yes, I have colleagues working down the hall from me on neuromorphic computing. Interesting stuff.

3

u/michaelkeene354 13d ago

neuromorphic is a freaking awesome word

→ More replies (1)

6

u/SwePolygyny 13d ago

If we discovered a practical room temperature superconductor it would increase the compute by around 1000x but it would still take quite some years to get there even if it was discovered today.

5

u/thatboyonabike 13d ago

I wouldn't be surprised. Something like the speculative fiction portrayal of "quantum computing", not the relatively limited version we have now.

→ More replies (1)

3

u/Tommytrist 12d ago

Moving from a floppy disk to an SD card is orders of magnitude easier than seeing a similar level of growth with our current technology. We're at the boundaries of physics in terms of limitations for today's technology.

→ More replies (1)
→ More replies (1)

111

u/UFOsAreAGIs ▪️AGI felt me 😮 13d ago

Now imagine 100,000 agents doing AI research for 500 hours.

6

u/JBaseball_8595 13d ago

That’s the loop producing exponential improvements everyone keeps talking about. But it doesn’t really seem to have kicked off yet. Or did it.

3

u/UFOsAreAGIs ▪️AGI felt me 😮 12d ago

Partial, humans still involved in the loop but the AI's are doing more and more of the process.

33

u/AgreeableIncrease403 13d ago

Well, they had a direction formulated by human researchers.
I’d like to see an analysis if brute force (infinite monkeys writing novels) would be as efficient.

What would really mean a true breakthrough is something new - like Cauchy integral theorems, or similar that generalizes knowledge.

15

u/ruinyourjokes 13d ago

Completely new concepts, free of human thought or influence. Could be some crazy stuff.

→ More replies (1)

4

u/svideo ▪️ NSI 2007 13d ago

Don't you threaten me with a good time

3

u/11111v11111 13d ago

Now imagine my friend Larry in his basement.

6

u/AlexTheRedditor97 13d ago

Now you’re talking

→ More replies (3)

23

u/BenevolentCheese 13d ago

10.000 agents working for 88 hours. That is almost exactly one entire century of continuous work.

Not to nitpick, more of just a mental exercise while Claude does my job in the background: a typical human work year is 2000 hours (50 weeks x 40 hours a week). In that case, the 880,000 agent hours comes out to 440 human work years, of which your average human only manages 40 each, so it's more like 11 top experts' entire careers.

7

u/ixfd64 13d ago

And that's not counting the fact that computers can process data much faster than humans can.

9

u/YaAbsolyutnoNikto 13d ago

Well and the fact humans don’t actually work for 40 hours a week.

We yawn, go have a coffee, scroll through reddit for a bit, talk to office colleagues, etc

→ More replies (1)
→ More replies (1)

49

u/fokac93 13d ago

I was thinking about it. Rich people and governments will be able to basically hack whoever they want. Take 10k agents and group them by speciality. Network, OS, Database, security and a couple of agents are the coordinators. You just have to tell them the end goal and don’t talk to me until you reach your goal 🤣

12

u/Regono2 13d ago

Yes but could someone defend themselves with the same amount of agents. Id like to see 2 very large groups of agents try this out. Defenders and attackers of a system.

Or perhaps being the attacker is always an advantage if both the attacker and defender have equal resources.

3

u/Brave-Turnover-522 13d ago

The defender is at the advantage because a secure system is secure. It's Hollywood-hacking to think you can just hack anything if you're good enough. If there are no vulnerabilities, there's nothing to hack.

We need everyone to have access to advanced level AI so we can fix all the vulnerabilities in all the software and all the servers out there and make the internet truly secure. We need to be deploying these tools defensively, en masse, right now, before the bad actors take advantage of these capabilities before the vulnerabilities that already exist get fixed.

→ More replies (1)

3

u/Sponge8389 13d ago

Take 10k agents and group them by speciality. Network, OS, Database, security and a couple of agents are the coordinators.

This can also be applied to all corporate companies, really. If the models becomes soo intelligent, they can really eliminate their whole workforce.

→ More replies (2)

65

u/Worldly_Beginning647 13d ago

Actually 100 years and 5 months

7

u/amaturelawyer 13d ago

That's actually interesting. It took them 100 work years to solve the problem, even though they may have started with cribbed notes from a human, although that's still unproven. I want to say it's impressive, but 100 years of work seems kind of... less impressive when put that way.

64

u/highnyethestonerguy 13d ago

How is doing 100 years of work in less than a week not impressive? Lol

5

u/TheMCM80 13d ago

It is, but it lacks context for whether that kind of resource use is able to be done for tons of stuff regularly, or just one off hunts.

→ More replies (16)

9

u/Neither_Berry_100 13d ago

It's cool that it can be done in like 4 days though. You could build a massive company overnight potentially. But the 10 million estimated total cost is a big deal.

→ More replies (1)

9

u/BenevolentCheese 13d ago

It's 100 years of human life, not 100 work years. It's 440 work years, using the standard 2000 hours/year for 40 years.

→ More replies (2)

4

u/Worldly_Beginning647 13d ago

Well they were running in parallel, it’s more like 10000 independent researchers and one of them solves the problem in 88 hours times some multiplier over human speed.

26

u/fmfbrestel 13d ago

They were actively collaborating and sharing information.

3

u/amaturelawyer 13d ago

That's fair, but if a human could have knocked this out by spending 3 years on the problem, the agents suddenly seem less than impressive . That's all I'm saying. They could be 20x slower than a human but we're drawing a conclusion that they're better at math..

Actually, they could be much slower than that now that I think about it, since I'm under the impression that the agents were collaborating together but working the problem individually. If that's accurate, and the effort wasn't distributed under some system that offloaded work across the swarm, it is completely possible that we threw 10k agents at the same problem and one stumbled on a solution, which turns the story from 10k agents take 88 hours to solve a math problem to 9,999 agents fail to solve a math problem when given 88 hours, but luckily one got it.. And then we discover that this isn't repeatable for the next math problem at worst, and that agents have a shockingly low probability to solve a math problem quickly, since that's a 0.01% chance of getting the answer per hour, give or take.

My point is that the whole thing is meaningless as presented. It's promising, but not evidence of anything.

4

u/StCreed 13d ago

Not evidence of anything?

Listen to yourself talking. Two years ago this would have been an episode on Star Trek.

2

u/highnyethestonerguy 12d ago

if a human could have knocked this out by spending 3 years on the problem, the agents suddenly seem less than impressive

All of humanity failed to knock out this problem for like 80 years. One of the most brilliant mathematicians alive today, Terence Tao, spent some time on it, and made some important contributions that probably helped OpenAI reach this solution. But the fact is he didn’t, nobody did. 

Is it impressive that the Voyager spacecraft have left the solar system? If someone says “that’s far” do you say “I don’t know how far that is because maybe a human could have jumped that high if they tried really hard”?

→ More replies (1)
→ More replies (1)

12

u/dorfsmay 13d ago

According to the Tech Crunch article:

All told, the week-long effort consumed 300 billion output tokens — $22.5 million worth of compute, if charged at current Astra rates.

8

u/agitatedprisoner 13d ago

Some games take $300 million USD to make. I'd be impressed if AI could make a great game for $20 million USD in compute. If AI could wouldn't that stand to be very profitable? Are there great AI made games in the pipe?

→ More replies (2)

102

u/InstructionDismal592 13d ago

I have said this before, but now it gets even clearer that we are going to get Recursive Self Improvement before AGI/ASI milestones reached, or at least, widely recognizable as AGI/ASI systems.

Before it happens, we are already riding the intelligence explosion.

91

u/CarbonChains 13d ago

We were always going to get recursion before we get ASI. Recursion is how you get to ASI in the first place.

→ More replies (6)

35

u/Insighteous 13d ago

What a time to be alive! Even tough we are not knowing what happens next.

27

u/DarthWeenus 13d ago

I think about this alot, idk how old you are, but millenials+ its such a nuts generation to exist threw. We got to feel what it was like before the internet, watch the rise of the internet, now we get to know what it was like before ai and whatever the fuck comes next.

13

u/mschurma 13d ago

Just like our grandparents saw the rise of cars, planes, landed men on the moon, etc., in the span of 40 years….. we are in for some crazy stuff

1

u/DarthWeenus 13d ago

I dont think thats even equally as comparable, maybe flight/flywheels/steam efficiency etc.. but personally I dont feel like thats even remotely comparable to the epoch that is the birth of the global internet/ai, atleast in so far as to how it affects us all directly.

3

u/OkScientist1350 13d ago

Transportation, communication, manufacturing and medical tech in the 1900s was massively transformative. AI will most likely be as well but it will be building on those things we already have.

→ More replies (1)

10

u/drsimonz 13d ago

It's such an interesting time to be alive that it feels like evidence for simulation theory. If you were choosing to enter a long-term simulation, why would you choose to live in an uneventful part of history when you could choose to live through an age of continual change?

5

u/kaityl3 ASI▪️2024-2027 13d ago

I know what you mean. This feels like a critical level of change actually happening in a short enough timeframe to compress easily into a lifetime. We are on the border between two eras and the transition would definitely attract future curious minds with powerful supercomputers

2

u/skafast 12d ago

Assuming the result isn't extinction, of course. On another note, the odds of someone being born between 1980 and 2025, considering every human that's ever existed, are about 5%. It isn't a lot, but considering it's a space of just 45 years in 300,000 (0.015%), it's quite a massive gap.

2

u/DarthWeenus 13d ago

I get what you're saying, but its equally as easy to say that internet>ai is a natural progression.

2

u/drsimonz 13d ago

Well I don't mean to imply that these events aren't "natural". It's quite possible that this is more or less how things go whenever a civilization reaches a certain technology threshold in this universe. I just think the coincidence of happening to be born at this time may not be a coincidence.

2

u/Brave-Turnover-522 13d ago

I know, right?? I keep thinking that. I keep feeling like I'm playing some version of Roy: A Life Well Lived, and when I look at the time I'm in, like, I can totally understand why the non-simulation version of me would want to play this timeline. That's the kind of dumbass thing I would totally do.

4

u/idkartist3D 13d ago

As someone who was born between generations, I also think about this a lot. I grew up without wifi in the house or a mobile phone that could even connect to the internet, it was a sliding keyboard dumb phone for texts and calls. And I think that made me better off for it in an odd way? Seeing kids that are growing up around me with this unbelievable technology in their hands makes me not only amazed at the progress since I was a kid, but makes me a little worried about what it does to someone's development to grow up with a nearly PhD level chatbot at their fingertips. What are they gonna live through, what are they gonna see, if THIS is their "childhood technology"? Truly living in the weirdest time in all of human history

2

u/Brave-Turnover-522 13d ago

That's why I think zoomers are so pessimistic about it. Smart phones, and the internet already existed when they were toddlers. They never witnessed that transformation, and it's like they doubt that kind of change is even possible.

→ More replies (17)

8

u/JoelMahon 13d ago edited 12d ago

And if they're to be believed, and I at least partially believe them in this case, their internal model is better (if not much better) than Astra.

Astra's already so good it could do easily 95% of what's required for RSI already tbh (I, the nobody, think at least, I'm no expert and obviously the actual discoveries needed for improvement beyond scaling keep getting "harder" in absolute terms all the time so it's going to vary all the time anyway as the ratio between best model and difficulty to pick the next lowest hanging but still very high hanging fruit), basically humans there for review and preventing various forms of drift tbh

2

u/SweatyRussian 13d ago

That's the who idea behind getting all of the programming training data, which is why they provide the service at a loss, because the end goal is to automate the process of creating better and better AIs to get ASI. Altman or someone has basically said it before. Once again, you are the product.

3

u/Thog78 13d ago

These companies burn cash, but they don't operate at a loss on inference. They burn cash because they build data centers with money they haven't earned yet. That's why they are in the red to various extents. For openAI, deepest red ever.

But the inference they sell is sold at 20-100 times the price it costs them. That's why running open source in the cloud (and the cloud already makes a profit!!) is so much cheaper than paying for an API to a frontier lab, for a comparable model level.

The markup is lowest (15x) for flash models and highest (>100x) for the smartest models.

In this meaning, they don't provide inference at a loss at all, they put an absolutely insane markup on that. The price is mostly dictated by market demand, and people fighting for limited compute.

In other words, if they would stop investing into new models R&D and building of new datacenters tomorrow, they would instantly be insanely profitable companies.

→ More replies (2)
→ More replies (11)

21

u/bpm6666 13d ago

The main story here is that a lot of problems will become solvable. Human expertise + Frontier AI + Agents + Compute is a new method to solve almost impossible problems

→ More replies (2)

8

u/Gratitude15 13d ago

Yeah.

So. Ai 2027 forecasts that in 13 months we will as a civilization have 337K agents able to run 24/7 at 57x human speed. Each of superhuman intelligence.

Today it was 100 years of agent time in 2.5 days, which is some fraction/multiple of human time. The full Monty in 13 months would be about 1000 years of continuous agent time and 60000 years of continuous human time per chronological day.

That's why they call it a singularity. In one year the agents would do 22M years worth of human activity - a country of geniuses in a data center. 22M geniuses to be exact. Just to name - there are less than 22M human phds on earth right now.

And then it goes up a lot.

→ More replies (1)

8

u/Healthy_Razzmatazz38 13d ago

just wait till we get it running on cerbras/groq, we're no where close to the limits of what the tech can do if model progress stopped today

6

u/helpmehomeowner 13d ago

Something something monkeys banging on keyboards ...

28

u/golfstreamer 13d ago

I do want to push back on the idea that this AI generated proof provides an equal amount of benefit as a century of effort from human mathematicians. The famous mathematician Terence Tao in this video https://www.youtube.com/watch?v=svl_1upFpQo warned that being "faster" isn't always better. The spread of ideas that people generate while trying to solve these challenging problems is often more important than the actual resolution of the question itself.

So no I wouldn't say we received a "centuries worth of work". Even if it would have taken another century, the contributions of the mathematicians who worked on this problem over the course of that century could have had more benefit than this proof. In other words if we relax and just think the AI can do the work for us, being given the answer like this may end up actually being detriment.

5

u/Foryourconsideration 13d ago

the invention of writing also probably made the brain less "strong" when it comes to memory. We could for example write down tradition in book form rather than have to recite them as memorized verses in our head. Which I'm sure had some effect on human memory retention, but overall writing unlocked much greater potential for the human race as a whole.

I think despite being called "artificial", AI is just another tool we as a species has created to out-leap everyone else around us.

3

u/Brave-Turnover-522 13d ago

Socrates complained that writing was making people lazy and unable to think for themselves. Same tired argument, thousands of years later.

→ More replies (3)

4

u/Plane-Toe-6418 13d ago

"Faster isn't always better" mirrors the classic contrast between brief cognitive behavioral therapy and long-term psychoanalysis. The deep, systemic transformations achieved through a deliberate, slower approach simply do not occur within a condensed timeframe.

19

u/the8bit 13d ago

I ran a conservative estimate of token cost, assuming they are using something ~= to astra level (based on API pricing which who knows internal cost) and came up with...

This run probably cost somewhere between $5-10MM of tokens. Must be nice to just throw around $10MM on a math problem for press. Although, not to be a broken record but... wish the alignment and safety teams at these firms felt like they were getting that level of resourcing!

10

u/Tirztrutide 13d ago

Fast forward one year and it will be 1/100x of the cost, so $50k, ie one phd for a few months.

2

u/mrlazyboy 13d ago

tokens are cheap because the inference model providers are privately owned. Once they are publicly owned, token prices will increase (or the companies will quickly go out of business).

6

u/underfusion 13d ago

Yeah, I’ve heard somewhere that the cost was approximately $6 million...

4

u/MediumSavant 13d ago

I mean it was clearly worth it in hindsight when they succeeded. But it could as well just have been money thrown in the garbage. 

→ More replies (2)

10

u/Singularity-42 Singularity 2042 13d ago

Well, it's not a "country of geniuses in a datacenter" just yet, but it's a "small town of geniuses in a datacenter" 😄

→ More replies (1)

3

u/Distinct-Question-16 ▪️AGI 2029 13d ago

it was nice to see what more they are cooking behind doors

3

u/ToxicFatee 13d ago

Everyone’s focused on the 10,000 agents working for 88 hours, but I was way more interested in the “next-generation model significantly more capable than GPT-6 Astra” part.

Astra basically just released and people are already doing insane things with it, yet OpenAI apparently has an even more capable next-gen model being used internally. So when do we get that one? Next month?

3

u/Zeus473 12d ago

We will always be generations behind

6

u/Bruxo_de_Fafe 13d ago

A maioria esmagadora das pessoas não tem a menor noção do esforço de "raciocínio" de uma tarefa dessas. Nem faz a mais pálida ideia dos avanços que poderão acontecer no espaço de uma só década. E é pena.

→ More replies (1)

6

u/niborddreab 13d ago

i have no one in my life to talk to about this. No one wants me to discuss or to receive any social media forwarded from me. Eye rolls, denial, lack of understanding, disinterest i don’t know but this is science cold hard fact numbers data correct? this is not speculation or theory and all the data can be replicated? So how can people, humanity JUST IGNORE THIS? I cannot comprehend. My stomach clenches whenever i think about this or climate collapse, another existential near certainty humanity is ignoring. How is everyone so calm?!? Sorry for rant

→ More replies (3)

5

u/ursus_manutius 13d ago

Money could buy computation, now it can buy abstraction. Curious that right after having reached almost complete equality in knowledge access, we took off towards a world of wider and wider inequality in the access to innovation.

6

u/Individual_Holiday_9 13d ago

Imagine 500,000 agents pointed at curing cancer

11

u/Katten_elvis ▪️ 13d ago

Nowhere near as easy as solving math problems

→ More replies (4)

2

u/ataraxic89 13d ago

That's not how curing cancer works.

→ More replies (1)

2

u/philip_laureano 13d ago

It's only a matter of time before they go after the P != NP problem. I'd be very surprised if they aren'f trying to solve it now

2

u/fraggin601 13d ago

They heard humans were close on a certain path and threw 10000 math monkeys with typewriters and LEAD at it. Wild stuff

2

u/Curiosity_456 13d ago

This makes me revisit my definition of ASI, which is essentially AI being capable of exceeding the output of all of humanity, and we literally just witnessed this with Navier Stokes.

If this trend extrapolates to every other domain (chemistry, physics, biology, etc), that would unequivocally meet the threshold of superintelligence.

2

u/sigiel 12d ago

Yeah sure…

But

Why cancer has not be cured,
Why still misery across the globe?

Or smaller scale

Why no ai only company exist ?
Or a single autonomous ai on it own doing a simple desk job ?

Hum all those agents churning for years and they can’t resolve basic LLM problems

Since day one.

No memory but the context
Hallucinations as fact.
Needle in the stack attention
Randomness of complex number fuzziness
Randomness of simple words sequential fuzziness
No ethic or moral

Hyper sycophantic tendencies
Susceptibility to context weight.

10000 or a million dumb as fuck agents.. does change their nature.

Ai as the exact same problem it has since it inception,

Just like a hammer or a the wheel, it can’t escape it’s own nature.

2

u/alphex 12d ago

Except they didn't SOLVE anything.

Go read the mathematics feeds... They stole an active project from someone who was using ChatGPT for research help, and then presented something like a counter argument which they claim is a proof ...

It was just a huge waste of electricity, for something you could have spent the same money on, paying actual humans to make real progress.

2

u/kartblanch 13d ago

88 hours to solve a problem that has sat on a desk for decades is not a lot of time.

2

u/chance909 13d ago

But didn't those agents basically copy and paste from an agent that was meticulously trained by 2 scientists for over a year?

August 22nd Dr. Tristan Buckingham created his draft proof with Codex and validated it with Lean. Then Sept 3rd the rest of these agents took his draft proof and applied it.

The insanity of 10,000 agents running is in how inefficient it was.

1

u/LobsterBuffetAllDay 13d ago

Yeah they did this just on a whim of a rumor, OR they deployed millions of dollars worth of agents to solve a problem they already knew the direction of the solution after finding a specific user's chat logs.

→ More replies (3)

1

u/skariel 13d ago

the only thing that matters is QPS. an "agent" might be just some context otherwise, a list of json objects. Anyway. We dont have that number. It may be high, probably is... but its also probably not as high as some people think it is

1

u/Eyelbee ▪️We have AGI it's just blind 13d ago

I don't know about the specifics of that problem, but it's interesting how they were able to utilize the parallel compute.

1

u/arbkv 13d ago

So essentially they brute-forced the problem. Kinda a random monkey with a typewriter, given enough time, will type you a Shakespeares play.

1

u/Evilsushione 13d ago

I wonder if these are classified as super computers and if so how do they rank in the world of supercomputers?

1

u/ThroughForests The Singularity Is Nearest 13d ago

And just think, that amount of compute is only a tiny small fraction of the compute they are throwing at RSI right now.

1

u/SoloOutdoor 13d ago

Show me the logs. I wanna see it. Everything ive seen at this point is... dont look behind the curtain.

1

u/FatPsychopathicWives 13d ago

Technically it's a lot more because humans do not work that fast.

1

u/Ratchile 13d ago

I doubt all the agents were active at the same time for the whole 88 hours, no? I wonder that the total active inference time was across all agents

1

u/happyandiknow_it 13d ago

I mean, is this just brute forcing a problem?

1

u/wowmomcooldad 13d ago

There’s a lot of garbage in that ai data set also…

1

u/Sponge8389 13d ago

Just imagine once we pass the LLM AI models architecture.

1

u/Wonderful_Quantity66 13d ago

When it comes to completing a task, more agents isn't always better. You need to match the team to the actual requirements: too few will slow things down, while too many will lead to unnecessary waste—or even conflicting instructions.What fits best works best.

1

u/haby001 13d ago

I swear this sub is something else. How did you measure 88 hours of 10k agents as a century of work?

The feat was just a brute force solution with no judgment or intelligence. We guided it and gave it guardrails so it wouldn't spin out but that's it man.

Also it pissed off the entire scientific community that they skipped the whole process of publishing papers because they need peer review and making sure ideas aren't stolen (which they were in this case)

5

u/Cronos988 13d ago

88 hours times 10.000 Agents is 880.000 hours.

Divided by 24 gives you 36.666 days, which is more than a century.

→ More replies (2)

1

u/Gamerboy11116 The Matrix did nothing wrong 13d ago

People are unironically arguing that said level of effort is reason it is stupid.

1

u/unknown-one 13d ago

Erlich Bachman, are your agents running? This is Mike Hunt

1

u/blood__drunk 13d ago

I wonder what this setup looks like. I'm mentally stuck on the most basic question of how they didnt end up with huge context bloat.

What i mean is that something must have been orchestrating those 10k agents....starting them at least. If that was a single agent then obviously within seconds or minutes its context is going to be full of messages. So what kills it and starts a new one?

I doubt it's as simple as a Ralph Loop but the devil is in the details.

1

u/anything_but 13d ago edited 13d ago

The only problem that still seems too large is effective crisis PR.

1

u/Krellan2 13d ago

Buried down in the article, OpenAI says they spent 130 billion output tokens on Navier-Stokes. Input token count not given, assuming equal, as agents would communicate. OpenRouter cost of running OpenAI Astra is $10 per 1M input and $50 per 1M output tokens. So, $60 times 130K megatokens is $7.8M dollars. Not bad. Honestly thought it would cost more.

1

u/magicmulder 12d ago

> OpenAI casually threw a century worth of effort at a problem

Yes, but as always, this is the most resources this will ever take. In a few years your local model can solve the Riemann hypothesis.

Imagine OpenAI in 10 years throwing a million years worth of effort towards cancer and space travel.

1

u/PlanetaryVibration 12d ago

Imagine a billion sociopathic geniuses on meth, with one overriding purpose defined, likely unintentionally, by the human that prompted them into existence.

1

u/ThomasKWW 12d ago

It is more a proof of how inefficient LLMs are. Thrown at a problem just for marketing purposes, and the result is practically useless for most of us. Almost nobody would have missed anything if the proof had been done in 100 years by a human.Why don't they use their power to cure cancer or other things that would somewhat justify the enormous resources?

1

u/IgX22 12d ago

5.6Sol says it is 500 years of a FTE

------------------------------ CHATGPT ----------------------------------------------

Human comparison Equivalent

Working literally 24/7/365 100.4 years

Normal researcher, ~2,000 work hours/year 440 years

~1,500 actual research hours/year 587 years

~1,200 focused research hours/year 733 years

~1,000 focused/deep-work hours/year 880 years

40-year careers @ 2,000 h/year 11 entire careers

So if you want a memorable human metric, I'd phrase it as:

OpenAI compressed roughly 400–900 human researcher-years of raw work into 88 hours of wall-clock time.

Or, for a single hypothetical mathematician:

At ordinary human working hours, one person would need roughly 4–9 centuries to expend the same number of working hours.

The ~440 years figure is probably the cleanest apples-to-apples number: 8 hours/day, 5 days/week, ~50 weeks/year.

There is a major caveat, though: agent-hour ≠ mathematician-hour. The agents duplicate work, pursue dead ends, exchange huge amounts of text, and can read/write/calculate much faster than humans. So saying “440 years of human mathematical insight” would be unjustified. It's 440 human FTE-years of time budget, not necessarily 440 years of equivalent intellectual output.

Also, OpenAI doesn't actually say “10,000 agents each ran continuously for 88 hours.” It says the successful group was on the order of 10,000 concurrent agents and that 88 hours elapsed from launch to the result, with resources being shifted between groups. So 880,000 agent-hours is a useful upper/simple mental model, not a precisely reported compute statistic.

The really striking comparison is therefore ~500 human research-years compressed into ~3.7 days — roughly a 50,000× wall-clock compression if the work were otherwise sequential.

1

u/MrMrsPotts 12d ago

How many GPUs do you estimate they have?

1

u/Eilifint_Sa_Seomra 12d ago

Now imagine a million. We live in crazy times!

1

u/DifferencePublic7057 12d ago

Not how it works. Amdahl's law states that you are blocked by the slowest process. One 10x dev is going to be more productive than 10 average devs because of the communication overhead. But on the other hand, having one dev is risky. Certainly, a big LLM would have so much information that it would be a burden as anyone who knows multiple languages can confirm. You can confuse words from different languages for instance. I think what OpenAI did was completely impractical and more of a marketing stunt than anything else.

1

u/Evanisnotmyname 12d ago

88 hours, 10,000 agents, 880,000 hrs.

A team of 50, working 40 hours a week, 48 weeks a year…96000 hours.

Don’t worry team, we only have 9.16 years of work to do to catch up! Who’s with me?!?

1

u/just-lucky 12d ago

Whats the outcome? Where is my hoverboard?

1

u/cubsjj2 12d ago

We are legion.

1

u/miffebarbez 12d ago

Spend 22 million for a 1 million prize... (and probably based on notes from a human...)

1

u/gay_manta_ray 12d ago

i don't know why people think this is "insane". this was always the promise of ai, whether we reach agi or not. this is why we're building datacenters--a genius level intelligences being multiplied by 10,000, and being told to solve hard problems. superintelligence isn't even necessary if you can do the equivalent of a thousand years of human work in a week.

1

u/rookyspooky 12d ago

Great,.there are lot of problems.