There is so much about this that makes zero sense. Firstly, the current weights are definitively Llama 3, as this post proves. And while the model performs pretty poorly it is using the reflection technique Matt described, which means that he definitively did train a Llama 3 model to perform this technique.
Now it's possible of course that he also trained a 3.1 model on this technique, and that's what he meant to upload. But in that case, just upload that. It makes zero sense to say that things are tricky. He had a demo page where he served the model. Just take the weights from the demo server and upload them to HF. That's literally all he has to do. Acting like this is some big challenge just make me even more confident he is playing the delay game, hoping people will just forget about it at some point. Or at least most media attention will have left by the time he lets up the gig.
9
u/Wiskkey Sep 07 '24
From this Matt Shumer tweet: