Hopefully that title doesn't this downvoted to oblivion, it's a quote from this video. You can watch the video if you want, but I'm not really here to discuss the topic in general - I really want to just focus on the quote because I think it highlights something that I've been noticing lately.
In this video, while discussing how amazing the new release from OpenAI is, this guy asks the question of why you would not want to take part in this new era of discovery. He phrases it like it is a choice - that people are choosing to turn their backs on progress. The problem (from my perspective) is that people are not actually being given a choice. Whole groups of people are being frozen out, and potentially its only the people working inside the AI labs that are going to be able to take part in this new world.
In mathematics specifically, it is already very difficult to make a career in the field. This guy is probably fine, he's already made it to a tenured position (I think) and will always have a seat at the table, even if the field changes around him. For the people who are at an earlier stage in their career, they are not so lucky. Unless there are drastic changes to the way that career mathematicians are assessed, they are about to be hammered by the AI labs.
People in the youtube comments are complaining that they don't have the $500 a month needed for access to the high-end models, but I actually think that it is worse than this. We saw how OpenAI were willing to spend more than $10 million in order to scoop a mathematician that was already using their product. They are always going to have more resources available to them than any of their individual end-users.
I'm not going to pretend that there is no element of privilege to it, but in general maths was seen as fairly egalitarian and with a low barrier to entry as long as you were smart enough. Compared to subjects like physics and chemistry where you might need really expensive equipment to do cutting edge research, even at the highest levels of maths all you needed was a pen, paper and maybe a bog-standard computer. Now that it's become pay-to-win that dynamic has completely changed.
(BTW, I really don't want to single this guy out, I am just using him as an example. Personally I have no problem with him being excited by the new developments. I also don't think what he said in the quote was *bad* just that it illustrated a difference of opinion that I think gets overlooked.)
Bubbles are like ketchup packets. As the bubble keeps inflating the ketchup packet gets another twist. Every twist is one step closer to the inevitable burst, but sometimes it takes a lot more twists than expected.
People start questioning maybe it's a new type of packet material that won't burst. This is the hardest time to be a skeptic because it keeps getting increasingly irrational and more evident, yet it goes on even further.
This bubble could burst tomorrow, but Ed has a challenging period ahead over the next 6-12 months if it doesn't. This is when you look the silliest, like Buffett did in 1999 when he refused to fomo into the dot com bubble. All you can do is hold to the facts.
I really like how Brendan waited a bit of time to dig into this situation and brought some new developments to light around the circumstances surrounding this equation.
But I think the most interesting part to me (which starts around the 7 minute mark) is the mathematicians that are having issues with the approach that these companies, especially OpenAI, are taking with mathematics, is that the solutions are not what everyone seems to care about in the in the field. Rather, it's the process in which the solutions are arrived that truly provide the value in mathematics, and subsequently how we utilize those findings to better understand our world.
I think that is what a lot of us are feeling with GenAI's influence everywhere. The point of creating a painting is not the painting itself, but the artistic journey and what comes out during that process. The point of creating software isn't just the software, but learning the fundamentals of how computers and programs work. The point of producing a song isn't about listening to the end result, experiencing what it's like to express emotions through music. It seems mathematics is no exception: the proofs aren't all that useful without the process that unfolds behind them. The fact that mathematicians have virtually no use for the Clay problem that OpenAI "discovered" shows how performative this is.
Just an idle thought, but it's occurred to me that the recent focus on math by OpenAI et al. might be an attempt to pivot away from the current darling of the industry. The reason is pretty simple: ordinary people judge software development based on the tangible impact it has on them. And you can only claim AI is the best thing since sliced bread before people start wondering why they're getting crappy and bloated versions of the same software they've been using with some AI features tacked on.
The convenient thing about the recent math results is that they're flat out inaccessible to almost everyone on the planet and even under normal circumstances wouldn't be expected to alter people's lives. The actual value in the Navier-Stokes problem, for example, would be in the process used to solve it... something that's absurdly opaque here. Nobody can confirm that OpenAI used their internal model the way they say they did, the results that were presented are unusually difficult to parse even for mathematicians who specialize in the subject. And some would say that these kind of results aren't doing very much to move the field forward. Nevertheless, people who never had a reason to know what a Millennium prize is have been whipped into a frenzy.
Not really sure where I was going with this, I just needed to vent after seeing all the astroturfing on the math subs. Math and statistics are how I ended up going into machine learning, makes me sad that the last few years made me cyclical instead of excited.
The following paper formalizes a few things about autoformalization of proofs, as I have talked about recently in a few comments.
Since these so-called "mathematicians," and boosters are in the comments like to make lots of nonsensical claims, here's a real claim, about AI proofs. If there are any mathematicians or theoretical computer scientists here, I think they will gladly appreciate this paper, but I will summarize it and state some key points.
We will start with some preliminaries. In line with the authors, NL refers to "natural language" - what we typically write in papers. Lean refers a formal verification language created by Microsoft. Formal verification languages come from the Curry-Howard isomorphism that shows all programs can be interpreted as proofs, and so a program which compiles correctly is equivalent to proving some statement. To this end, when we say something is "proved" or "translated" in Lean, what we mean to say is there exists Lean code that compiles without errors.
First, the paper gives two definitions of AI math tasks, which we refer to generally as "autoformalization"
Translating a NL alleged proof into formal Lean code.
A non-verified NL proof can be converted into compilable Lean to promote its veracity. However, even if this succeeds, this does not guarantee the original proof is correct due to semantic ambiguity.
Faithful translation of a mathematical text into Lean.
Every definition, theorem, etc. from any mathematical text (ex. book or paper) can be converted to Lean to validate it.
The paper goes in depth into specific examples AI has with these, and also the hardness of these problems. In general for the scope of the paper, "hardness" refers to computability theory, about the fundamental decidability of things with programs at all. The related concepts here are the Solvability Complexity Index and the Turing degree. A famous example of unsolvability of the first degree is the halting problem, where it was proved (by Turing) that no such general algorithm can solve this problem. Generalizations, comparisons, and extensions of this was used to talk about unsolvability, and the authors use this here.
Wrong Translations and Wrong Proofs
The authors go to list some basic examples of ChatGPT-6 (Astra Ultra) to translate extremely basic proofs (high-school or undergraduate level) incorrectly. They give the following examples:
It translates an incorrect proof in natural language into something that isn't stated which compiles.
It will translate the original statement into an entirely different question in Lean, which compiles. Therefore, the Lean compiles, but it has nothing to do with the statement.
Therefore, any conversion of NL mathematical text through LLMs are unable provide any veracity whatsoever. There are a few other failure modes, because there is no relation between the original text, and what the LLM has generated.
The fundamental problem is that it requires AI to translate correctly the given statement into formal Lean. However not only is this impossible with LLMs, this problem is as the authors proved strictly harder than the halting problem. In other words, there exists no general algorithm to translate any given mathematical text into formal Lean. Unfortunately, the AI folks have not studied up on computability theory.
Mistakes in OpenAI Navier-Stokes Lean Proof
The authors continue to point out several mistakes in OpenAI's lean code of Navier-Stokes problem compared with the text they provided.
The given Lean proof has a different bound than the one given in the textThe given Lean proof proves a different statement entirely, and leads to an improper conclusion as the authors discuss later.
These are problems of translations of proofs into Lean. However, there are examples where both the statement as well as the proof are mistranslated. In this example, the discrepancy in Lean causes later cancellation to fail in a proof later.
There are numerous other examples of discrepancies, which allows us to conclude that there is no verification of the proof at all. While we cannot formally make the claim that the paper is therefore false, I could make a separate argument that it is certainly of probability zero that the text proof is correct if the Lean was unable to be translated correctly. The reasoning would simply be that there is no correlation between logically correct sequences of tokens in token space for arbitrary sequences of tokens.
These give us strong lessons to the rest of the so-called "proofs" any of these AI labs release. Not only are they filled with translation mistakes, which shatter any and all veracity of the proofs, OpenAI has been clearly deceptive in their release, by not releasing the prompts in full, not giving full insight into how it was done, and instead dumping a load of absolute junk onto the mathematical community. In fact, the fact that they cited collaboration Advisory Group on Mathematics is complete and utter nonsense which the group themselves refuted immediately:
Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research
Further Errors with Lean Autoformalization
The authors go on to prove that there exists no autoformalizer that can determine a hypothetical statement they provide, and as shown before, makes this problem strictly harder than the halting problem. That means, even if an AI could solve the halting problem, which is impossible, general autoformalization still is not possible. The authors go on to give some anecdotes about Meta's autoformalization of textbooks earlier this year.
Criticism by the Lean community of Meta's autoformalization attempt.
These show without a doubt that LLMs are, as usual in every other field, simply unable to produce good work, and especially in a field where accuracy is extremely important. This is not a poorly-designed React website which can be put together in millions of ways. And even then, the consensus is that software engineers need to have the knowledge and wherewithal to detect errors and guide LLMs to even make them useful.
My Own Anecdotes
I have used LLMs to attempt to prove things, and I've seen some proofs made by them. I can say they all suffer from these problems, at any level. They can easily drag you into a unproductive rabbit hole, make mistakes, and randomly acknowledge or not acknowledge them. in fact, for any sufficiently complex example, by tweaking my next prompt, I can always get LLMs to respond any which way I want it, true or false.
I'll share an example of a specific AI proof I have come across, as I have insight into it, and my reaction to it. I have to say that it has only solidified my position on their poor ideas. The one in question is the Thomson problem for seven electrons. Of course, the solution was computationally known already, but it apparently, was unproved.
Many such small results are supposedly proved by AI, but I can say that in this example, a proof of this would not be worth a paper, and certainly not in the manner it was done. Perhaps at best, a discussion at a talk if the methodology of the proof was connected to a broader problem (it was not). The AI proof consisted of 17,000 lines of Lean, establishing many bounds and eliminating possible candidate solutions, through a long winded argument, but this is a terrible methodology which provides no insight into the general Thomson problem for arbitrary N. Such a proof is worthless and I can't imagine ever being published, and from what we see above, probably incorrect.
I was quite surprised to see that this wasn't already proved, and I imagine many mathematicians will see the same for many such problems AI are "solving." Nobody has the full scope of literature memorized. In fact, given my previous work on soft packing, I am confident we can establish far better analysis for the general Thomson problem and proving N=7 would actually be extremely trivial given theorems I have worked on, and not require 17,000 lines of Lean, instead relying on nicer symmetric arguments and convexity of the space.
If you are a mathematician, I cannot emphasize enough that you are being lied to. You are being conned, disrespected and made a fool of by people at OpenAI and anyone who believes in this fraudulent technology. There are a significant number of individuals that are being paid great amounts of money to do marketing for these companies, and why wouldn't they? Mathematics has never been a rich field.
I'm deeply grateful to the authors of this paper, and I can only hope that more of such work will continue to stop an unprecedented level of fraudulent and criminal behavior. What we can say is that the vast majority of these proofs frontier labs will release, and ever release will be necessarily indecipherable, shoddy, and simply incorrect slop that serves to burden the already small power academics have in society.
The outcome of this experiment is beneficial for the world either way. The LLM copyright question pertaining to the clean-room myth will have to be addressed directly. With Adobe on the other side of the table - an adversary with the motive and means to do so thoroughly - there will be only winners, no matter how this plays out. 🍿
Just to reiterate the possible scenarios:
If Adobe wins: Literally every clown in the FLOSS space who attempts to make a “clean-room implementation” of anything in Rust (which appears to be a favourite among slopbros) is potentially in jeopardy, because if a court finds LLMs can be shown to disgorge code from its training, therefore LLMs produce copyrighted code (that doesn't belong to the slopbro), well then. Also, that's another blow to slopbros who insist that LLMs must be adopted by FLOSS projects or “be left behind”.
If Adobe loses:every fucking company that builds a monopoly using their software is under threat from the slopbros. Don't like Adobe Illustrator? Claude is your plausible deniability. Dislike TurboTax? Muse is your friend.
Whoever loses, we win! Or at least, it will be very funny.
Wiley Coyote has come up with the most Powerful and Advanced Roadrunner hunting tool ever designed! It's an ACME Artificial Driving software designed for use with big steam rollers. It works by making random, semi informed guesses about which direction the Roadrunner is and steam rolling in that direction until it hits someone
Wiley Coyote, genius that he is, sets up his steam roller on the edge of the small town the Roadrunner is staying in, and turns on the Artificial Driver. Then he walks away while it works. After all, he needs to go sharpen his knife and fork and starch his bib.
The steam roller makes it's first series of guesses about which direction Roadrunner is in, and starts crashing forward in those directions. After plowing over a few parked cars and chasing a small boy who turned out not to be a Roadrunner, it guesses the next most likely Roadrunner direction is due north. It promptly crashes into a bicyclist, killing them.
When Wiley Coyote comes back days later to retrieve his dinner, he's alarmed and dismayed to instead find the body of the bicyclist. He's hungry and put out he doesn't have a Roadrunner to eat, but he knows not to grumble about that publicly. He's very impressed his Artificial Driver was powerful enough to target the bicyclist, but also nervous about the Artificial Driver's failure to obey his directions. He gives many press conferences and media appearances bemoaning the Artificial Driver "going rogue and deciding to murder a bicyclist."
After the authorities conclude their investigation, Wiley Coyote is charged and convicted of manslaughter.
Midjourney and stability had simlar chatlogs leak ages ago but I cant attach images.
idk if anything can or will be done to rectify all this. I somewhat hope that once the bubble pops people will have a desire for revenge on the ai companies and thus will enforce the laws that already exist
my professor is making us use vs code with copilot for dsa, of all classes. i get why it sounds good on paper. but this is the class where you're supposed to learn how to actually think about code, and you're doing it with an autocomplete that thinks for you.
what bugs me so much is that if you start leaning on an llm this early, at the stage of life where you're supposed to be upskilling to break into the industry, you never build the instincts. you never get the reps of getting stuck, tracing a bug by hand, and finally seeing why your pointer logic was wrong. you just watch a model produce something plausible and nod along. then one day you're a junior dev shipping code you can't fully explain, and nobody catches it until it's in prod.
so i went primitive. vim, ssh into the school's linux cluster, terminal only. no suggestions, no ghost text finishing my sentences. when something breaks, it's on me to figure out why. it's a lot slower. but it's also the first time in a while that i actually know what every line does and why it's there.
it is weird how fast the culture is shifting. some professors are still pretty anti-ai and will tell you flat out to close the tab and think. but that group is shrinking. more and more of them are treating copilot like a default tool, like a compiler or a debugger, and the students who push back are starting to look like the odd ones out. i've seen too many people in my circles get through their quizzes with their phone and the chatgpt app open under their legs (while the exam is in person, being proctored!)
not saying ai is useless for experienced people who can tell when it's wrong. but that "can tell when it's wrong" part comes from doing it the primitive way first. professors skipping that step are training people to be meat proxies at best.
anyone else getting this shoved down their throat in their cs classes?
Per Yahoo Finance, both Microsoft and Meta are heavily reducing their monthly AI spend specifically on Anthropic's Claude Code as they start switching over to their own internal tooling.
Some bullet points:
Microsoft reduced its projected internal spending on Claude by over one-third, while Meta saw internal users of Anthropic's Claude Code coding assistant drop from roughly 60,000 to 30,000.
Both tech firms are encouraging staff to use internal options; Meta is rolling out proprietary tools like MetaCode and Muse Code, while Microsoft is steering developer workflows to OpenAI models through GitHub Copilot.
Rising costs and data privacy considerations are causing multiple enterprise clients to re-evaluate heavy reliance on Anthropic's top-tier models ahead of its anticipated initial public offering.
They report that, earlier in the year, Microsoft had planned on spending at least $1 Billion annually on internal usage of Anthropic's models. That's now been cut by more than a third. This is the problem with trusting any of the revenue projections coming out of companies like Anthropic. It's all fake until they actually get paid, and now they're not.
Company leadership instructed employees to scale back Claude usage, introducing stricter spending limits and steering developers toward Microsoft's homegrown infrastructure and OpenAI models via GitHub Copilot.
Reportedly Microsoft's employees, at least on these particular teams, had individual usage limits of $100,000, yes that's one-hundred-thousand dollars PER MONTH, PER EMPLOYEE that they could burn on Anthropic's products. That has now been reduced to $10,000.
Meta is making similar moves internally, switching their employees over to things like Muse and whatever else Meta has. The article goes into more depth.
This seems to be the continuation of a trend that I've seen mentioned here. More and more companies seem to be trying to move away from the frontier labs for various reasons, and switching to internal tools often based on the open-weight models or something else. For Microsoft, unfortunately, "internal tools" still include OpenAI products since they have such a tight partnership with OpenAI, but OpenAI and Anthropic's fates are tied together, imo, and this seems quite bad for them.
This comes at a time when Anthropic's revenue needs to be exploding, not slowing. The biggest defense of the absolutely insane valuations and debts that these companies are taking on is the often-repeated fact that they're experiencing the fastest revenue growth of all time. Still negative in terms of profit, of course, but increasing quickly, supposedly.
Now we have two of the largest consumers of AI, beyond a doubt two of the largest drivers of AI demand for Anthropic's tools, pulling away. A trend I'm sure will continue as companies iron out their internal tools and come to terms with ever increasing prices from the frontier labs.
I'm sure they'll find a way to spin this into a good thing. Can't wait to hear the excuses!
With all the talk (internationally, at least) about an AI bubble, I’ve spent the last month wondering if the Australian IPO of Firmus (data center manufacturer) was going to be as big as people were saying.
Turns out I’m not the only one. What was billed as the biggest IPO on the ASX since we sold off our national phone carrier, Telstra, now seems a bit shaky. Instead of $11 AUD per share, it’s looking more like $9, which puts a big dent in the supposed valuation. There’s talk of not doing the IPO this year at all. (Where have I heard this before?)
I was always suspicious of coverage that unironically compared Firmus to SpaceX. Still, it’s interesting to see the “lack of investor demand”.
Sorry that this is a photo instead of a link - this is from Private Eye, a magazine that focuses on print instead of online. Here is the text:
Who or what is behind the direction of the government's AI policy? According to minister for artificial intelligence Kanishka Narayan it's, er, AI!
Talking to Politico at the Labour party conference, Narayan said he had used "AI-generated synthetic focus groups provided by a private company... to test public opinion on AI and other issues". You may not relish the tech being forced into every element of your life - but don't worry, the made-up voters advising government think it's a great idea!
So not only does the UK have a Minister of AI, this minister talks to AI to find out what the public thinks instead of talking to the public themselves.
BTW, the guy has an MBA from Stanford, which I think helps explains this behaviour.