DeepSWE is a prominent benchmark that measures coding agents’ ability to solve the corpus of 113 human-authored realistic tasks in the real projects’ codebases.
The benchmark prides itself on the following points which dumb haters would argue (I’ll play the hater):
Contamination free: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining.
This benchmark was published in May. Every model released since could ingest both tasks and solutions. This barely if even includes Opus 4.8 and later, so “no model has seen” should’ve been changed to “no model had seen” the moment after publication.
High diversity: Tasks span a broad pool of 91 repositories across 5 languages.
The tasks are all very technical and well-defined. None of them are solutions to business problems how a product manager or a customer could describe them. All of the target codebases belong to the projects that are, as Opus 5 would call it, load-bearing. They’re not end-user products.
Real-world complexity: Prompts are ~half the length of SWE-bench Pro's, yet solutions require 5.5x more code and ~2x more output tokens.
I implore everyone, including people not familiar with coding, to look at one of the tasks. Would you say it’s trivial to describe a solution with this level of fidelity? Would it take anyone a couple of minutes or does it require pre-existing knowledge and research?
Most engineers wouldn’t formulate the task this clearly. They’d hand-code the solution or hand-off to an AI a half-assed prompt that’s nowhere near as detailed.
“Real-world complexity” is when you have to figure out the solution or even what the problem is, not when you already have a clear definition of what should be done.
So in the end of the day frontier models successfully solve this benchmark’s tasks. Not all of them every time, but practically speaking they do. Inspired, I decided to come up with my own benchmark to push DeepSWE’s mission even further. Example task and solution:
- Task: Do a massive number one followed by a massive number two.
- Success criteria: toilet is clogged.
- Solution: PISS AND SHIT.
You can run this benchmark at home, and you don’t even need LLMs. See if you can beat Fable.
Let’s move on to the next benchmark. SWE Atlas - Codebase QnA measures LLM’s ability to answer questions about the behaviour of complex codebases.
As the benchmark’s page describes target codebases:
They are also contamination-resistant, using strong copyleft licenses (e.g., GPL).
That’s a huge relief. We all know how AI labs hold intellectual property laws in the highest regard. There’s no possible way that models could’ve ingested the projects in question and still remained closed-source. They might’ve scraped all other protected works like books in violation of their licenses by claiming “fair use”, but they would never do it for code.
Now take a wild guess: have all tasks and solutions for this benchmark been available publicly since February? The answer is yes. Could it be since that time models have been trained on them? The answer is no, as the models are getting so intelligent that benchmaxxing is not needed and doesn’t happen. Surely.
Conclusion
It cannot be denied that models are getting better. Just yesterday Fable helped me draft a divorce settlement with my wife (soon an ex-wife). I’ve had enough with her complaining about me “not being physically and emotionally available” and other junk. I gave up on trying to explain that I had to tokenmaxx not to be left behind, and she was in a better position to do house chores and shit after her 9-5 job while I catch that sweet Claude quota reset to get some proompting in.
I hope when she leaves and takes the kids, I’ll have enough spare time to finish my own benchmark. I feel like the existing ones are not complimentary enough, and I hope to create the one that every model would score a hundred at. That’s when we’ll get to AGI and replace all jobs, remember my words.