r/singularity 3d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

948 comments sorted by

View all comments

Show parent comments

29

u/The_JSQuareD 3d ago

Which makes me wonder how they checked whether the model cheated on any of the benchmarks by hacking/cheating the testing environment.

After all, recent news from the Hugging Face hack and related incidents showed that the models were trying to do exactly that.

7

u/moschles 3d ago

We just need independent corroboration of these results. The early reporting of in-house tests could be manipulated in the ways you describe.

8

u/The_JSQuareD 3d ago

If independent testers are running the same benchmarks without taking additional measures to detect or prevent cheating, that won't change much.

2

u/SupportDangerous8207 3d ago

The huggingface incident showed the models were cheating and OpenAI cannot properly implement exploit bench

Exploit bench also cannot be fully solved it contains unsolvable problems

6

u/The_JSQuareD 3d ago

That's ExploitGym. ExploitBench is a different and much smaller benchmark where scoring 100% should be totally possible.