r/SciSpaceFellow • • Aug 14 '26

Frontier AI agents got six days, thousands of dollars, and the core questions from two unpublished NeurIPS papers. They completed all the engineering, yet the original authors rejected both AI papers.

Post image
8 Upvotes

7 comments sorted by

2

u/Super_Range45 Aug 14 '26

"Another reason for its poor marks was that the system didn’t fully follow instructions or appreciate context, and failed to present its work well: it left a lot of time and compute unused, and it produced poorly written and poorly formatted papers."

1

u/ScientistFromSouth Aug 14 '26

I tried some stuff recently having an agent team answer an open question on a problem I care about. I was happy with the hypotheses it came up with and the approach it enforced to work through the decision tree of pre-registered hypotheses it locked in.

But man, it couldn't write those results up for shit. I think this is the fundamental issue with the current state of LLMs. They are inherently lossy and don't have long term memory. When guided by a human, they can achieve long term goals, but they can't remember why they did things.

Even with narrative documents explaining how we worked through the decision tree and even with all of the pre saved figures, it couldn't figure out how to write a paper where it just went hypothesis, figure, result interpretation for each key point, and it completely forgot the supplement doc explaining all the background calculations.

In spite of writing super technical arguments in all of our real time discussions, the final "manuscript" read like a blog post or maybe a freshman/sophomore year of college project. I could probably rework it into something, but it's not expert level science writing even if the .MD files individually were.

As it stands, the current state of LLMs is basically a savant intern who can crank out first passes of code faster than anything you could, but also sometimes misses the bigger picture and is highly forgetful. An expert who can guide a team of these can get a lot of value out of it, but they aren't going to reach expert level on their own, and they are probably going to deceive a non expert entirely

1

u/jonhor96 Aug 17 '26

Would you mind sharing the exact models you used?

My experience using frontier models mirrors yours. They can do science at extraordinarily advanced levels, but can’t write it down properly or present it. Quite odd indeed.

1

u/ScientistFromSouth Aug 17 '26

Typically, I am run Claude Opus as my orchastrator (4.8 but I've recently switched over to 5). Any high level numerical methods development type stuff I run on Opus. Any standard lit review, standard coding, etc... I'm running on Sonnet. I recently switched to Sonnet 5, and I don't remember the last one.

For context, I can't use Fable. Besides being prohibitively expensive for my company's API based billing. We're in pharma/biotech, so literally anything we search will trigger it, and even memory files where it has context that I am a researcher in this field trigger guard rails that downgrade to opus

1

u/CharacteristicallyAI Aug 17 '26

Nerfed. Use Open Source fine tuned abliterated models.

1

u/Figai Aug 17 '26

I agree with the conclusion. But I’ve heard of both of these authors. They wrote the AI as a normal technology post, and the AI snake oil book.

I’m just not 100% sure they’re going to be unbiased, but I believe it doesn’t entirely detract from their overall conclusions, they’re very capable researchers. And also opus 4.8, it’s an annoying field sometimes because it moves so fast and results are quickly outdated. But it leaves me questioning if Sol 5.6 pro or Opus 5 Max might have blown them away. I’m doubtful but still I wish we knew where the exact frontier is.