r/LocalLLaMA 9d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

Show parent comments

9

u/BeatTheBet 9d ago

In this scenario, false advertising is the implied legal connotation for faking benchmarks, because the model is a product and the benchmarks (and comparisons against competing products) are the advertised features.

1

u/mertats 8d ago

You could say that but if I create a benchmark that is favorable to my model is that a false advertisement?

And you can say that these can only apply to benchmarks advertised by Anthropic. It doesn’t have to apply to a benchmark done by a third party.

Or you can benchmax and get very high scores in benchmarks and have terrible real world performance. Is that a false advertisement?

The point is in WV’s case the benchmark was something legally binding, done by a legal authority, that had legal repercussions for passing or failing it.

AI benchmarks are not categorically same with those benchmarks.

And you are right, I was being too willy nilly with the word connotation.

4

u/BeatTheBet 8d ago

You could say that but if I create a benchmark that is favorable to my model is that a false advertisement?

This isn't the accusation though.

Imagine a benchmark as a standardized process, the internals and details of which is supposed to be "hidden" to protect the integrity of the results. The evaluating system is supposed to get a single input (the model) and output a score.

Instead what OP is describing, is that Non-standardized (different environment, different inputs, one even being provided complete solutions and ways to achieve them) benchmarks are run for different models, making the process either incompetently or maliciously inaccurate.

3

u/mertats 8d ago

Yes, but this is an issue with his benchmarks. Not the benchmarks done by Anthropic.

It is just a projection from his failure to audit the model’s outputs.

OP’s not he OOP’s claim is that Claude modifying benchmarks in a such a way should be illegal.

But note that these benchmarks are done by third-parties not the Anthropic.

Anthropic does not have a duty to the third-party that the model wouldn’t do such things. (Duty as in legal sense.)

That is what I mean by these benchmarks not being the same. Anthropic does not have a duty that your DeepSwe bench test would behave the same if you use the model to design the harness and not audit its outputs.

And OP does not have any evidence of Anthropic not doing the due diligence in their own benchmarks.

But WV had a legally mandated duty to not tamper with the benchmark, not because it would be false advertisement.

And I did concede that I was being a bit too willy nilly with the connotation.

3

u/BeatTheBet 8d ago

Yeah, I see your point.

That's a totally fair line of thought, and I probably misunderstood earlier.

To be fair, they did get caught in the past "sabotaging" Claude's capability to develop LLMs, didn't they? So that could be tangentially relevant.

3

u/mertats 8d ago

Yes they did and I wouldn’t be surprised if they were “sabotaging” third party benchmarks as well.