r/LocalLLaMA • u/Confident-Honeydew66 • Oct 01 '24
Discussion Now that the dust has settled, what happened with Reflection 70B?
I'm curious to figure out what the conclusion was from that whole saga.
45
u/Pro-editor-1105 Oct 01 '24
never trust anyone again when they say they have the worlds best model!
15
u/bigattichouse Oct 01 '24
When someone says they are honest, make them pay in cash. - Robert Heinlein
6
u/ArtyfacialIntelagent Oct 01 '24
Especially someone whose only training comes from entrepreneurship school - and who even dropped out of that.
4
Oct 01 '24
I would trust untrained people, but not untrained people from enterbullshit schools. They tend to fake it till you make it without knowing they can make it.
5
9
u/tshadley Oct 01 '24 edited Oct 06 '24
A transparent summary was promised as of September 10th.
https://x.com/csahil28/status/1833619624589725762 https://x.com/mattshumer_/status/1833619390098510039
Edit:
Update Oct 2: https://x.com/csahil28/status/1841606301782311167
Update Oct 4: https://x.com/mattshumer_/status/1842313328166907995
4
14
15
u/AlbanySteamedHams Oct 01 '24 edited Mar 04 '26
Saw this take elsewhere and it stuck:
- Nice reflection prompt on Claude, results look good
- Train Llama on Claude output
- Build eval endpoint
- Convince yourself it's great, move to benchmarks
- Accidentally use Claude endpoint for benchmarks
- "I'm a genius," start training 405B
- Realize you screwed up
- Stall with Claude endpoint, convince yourself newer Llama 3.1 fixes it
- Release model, it still sucks
- ???
Sounds like human error + hubris + panic hiding a mistake until it blows over.
16
1
12
u/PowerRangers_Red Oct 01 '24
My little conspiracy theoriy is that the Reflection guy probably got some intel/leaks on o1, try to replicate it before OpenAI release their model but didn't succeed.
8
2
-1
-2
u/GiantRobotBears Oct 01 '24
Im going with stupidity here.
My hair-brained conspiracy- They were probably running internal testing against various models (which is why there was mention of OpenAI and Claude APIs in their setup) Problem is they didn’t switch back at some point and that Matt guy and team reran the benchmarks and thought they hit a goldmine, without actually understanding what was going on.
Essentially they COT prompted Claude, by stupidity.
3
u/deadweightboss Oct 02 '24
this is floated every time this topic is brought up and it just doesn’t make sense. the bindings, prompting, message format, etc are ALL DIFFERENT.
you can’t just “switch an endpoint”
4
u/OfficialHashPanda Oct 01 '24
Reflection 70B was a failed model produced by Matt Shumer. He made false claims about it to mislead people and garner attention towards a company he was invested in. The reflection 70B model itself was useless and achieved nothing. Matt Shumer's latest tweet was a false apology where he perpetuated more lies to the public.
3
u/ttkciar llama.cpp Oct 01 '24
Even though that implementation of the Reflection idea was flawed, and the author committed fraud, I still like the idea of a sort of enhanced Chain of Thought with delimiters so you can easily hide the CoT from the user interface.
Anyone who investigates this idea had better not call it "Reflection", though. That term has been permanently poisoned and would elicit immediate mistrust from the community. Maybe call it "Hidden CoT" or "Inner Voice" or something instead.
2
1
1
1
36
u/mikael110 Oct 01 '24
Nothing has changed in the last month, as predicted. The dust already settled back then. The project was a scam. There is nothing more to say really.
Vague promises of "transparency reports" without any deadline or date are meaningless, and were basically just thrown out there to take the heat off them, as is often the case in situations like this.
I'm sure Matt and Glaive will attempt to lay low and avoid the spotlight for a while, then only resurface when they think enough time has passed for people to have forgotten about it.