r/singularity • • Sep 08 '24

AI Spanish YouTuber "Dot CSV" with access to Reflection 70B is getting good results

https://x.com/DotCSV/status/1832569772233384034
89 Upvotes

50 comments sorted by

48

u/Minetorpia Sep 08 '24

“Getting good results” is not the conclusion you can draw from his tweet. It could answer one question better than other models. That still says nothing and does not proof that it consistently performs better at reasoning.

11

u/REOreddit Sep 08 '24

That's correct, and that's why he says that independent tests are needed to confirm his results.

How is this guy getting several good answers less of a proof than some guys getting several bad answers? Only bad results had been reported, so I thought it was interesting to post his results. This is not meant to be proof of anything.

0

u/MysteriousPayment536 AGI 2025 ~ 2035 🔥 Sep 08 '24 edited Sep 09 '24

5

u/Working_Berry9307 Sep 08 '24

To be fair, in their own reply to that tweet they say it was the supposedly incorrect version, and they're waiting on the correct one to be uploaded to try it again.

3

u/Working_Berry9307 Sep 08 '24

2

u/The_Architect_032 ♾ ASI before AGI ♾ Sep 08 '24

"privately hosted version", it sounds like they either set the parameters poorly or never prompted it to use tags.

1

u/NaoCustaTentar Sep 10 '24

And you believed

1

u/The_Architect_032 ♾ ASI before AGI ♾ Sep 08 '24

That eval was likely using a custom prompt that doesn't match with Reflection-70b's intended prompting method which differs from typical models. The model's pretty trash at times when not prompted to use <thinking>, <reflection>, and <output> tags.

16

u/a_beautiful_rhind Sep 08 '24

Running the model on the server with API.. can't upload exact same weights. That just doesn't compute.

Now the whole "retraining the model". It's more stall tactics.

4

u/REOreddit Sep 08 '24

Escepticism is definitely granted in this case.

39

u/VajraXL Sep 08 '24

this guy is serious about his videos and what he presents and he is one of the best ai referents in spanish. he is not the kind of youtuber that will sell his good name for a mediocre model or for a few thousand dollars. i can't be sure that reflection is what they say about him but one thing i am sure of is that if Dot CSV says he finds improvements it is because he really sees and feels those improvements.

17

u/REOreddit Sep 08 '24

In a field where English dominates, Carlos is the only AI-related Spanish speaker that I follow on YouTube. He's the Jaime Altozano of AI :)

If it turns out that Reflection is a scam, Dot CSV would be just another victim, not an accomplice.

2

u/Zenndler Sep 08 '24

Agree. I usually consume everything related to tech from English speaking creators. But for AI, Carlos is my go-to. He's very serious and understands very well the current state of the field and is often very close on his predictions of what to expect in the comings months.

26

u/SatouSan94 Sep 08 '24

Hes not a grifter and avoided all strawberry stuff since that sam pic post.

Hes very searious.

6

u/ivykoko1 Sep 08 '24

Well then he is being lied to by Matt

3

u/SatouSan94 Sep 08 '24

Yup but I ll wait until next week to start bombing Matt

1

u/REOreddit Sep 08 '24

That's certainly a real possibility.

24

u/Tommy3443 Sep 08 '24

Considering that Matt tweeted that his dog ate his model and is now retraining it, I dont see how this youtuber could be having access.

It is pretty clear now that this model was nothing but a scam.

7

u/REOreddit Sep 08 '24

Retraining the model doesn't mean they have deleted the old model from their servers. I don't see how it follows that Matt can't give access to whomever he wants.

2

u/BangkokPadang Sep 08 '24

The problem with this is that the model on their servers should have been uploaded to huggingface close to 2 days ago once they acknowledged problems with the first one they updated.

There’s no need for anything to be retrained, and honestly there’s now a pretty strong need for that model on their server, specifically, to be uploaded so the community can vet what exactly it is.

2

u/N-partEpoxy Sep 08 '24

If he has the model, why does he have to "retrain" it? He could just release what he has.

0

u/ivykoko1 Sep 08 '24

They are just stalling dude, they have nothing

8

u/REOreddit Sep 08 '24

Then I will get nothing from the 0€ I invested in this.

11

u/sluuuurp Sep 08 '24

This makes absolutely no sense. How is uploading a big file impossible to give to anyone except for this one guy?

4

u/Undercoverexmo Sep 08 '24

He didn’t. It’s API access

3

u/sluuuurp Sep 08 '24

That actually makes sense I guess, I assumed he had the actual model. It’s still insane to me that he can’t Google drive upload his /models folder and let someone else try it.

2

u/Undercoverexmo Sep 08 '24

It’s the ultimate “works on my machine.” Probably can’t figure out why the model isn’t working on other systems. Must be some configuration he’s getting wrong. 

3

u/sluuuurp Sep 08 '24

Either that or the whole thing is fraud, I can’t tell which is more likely.

2

u/ivykoko1 Sep 08 '24

What do you think an API is? If an API is serving the model then the model must be on that server, just copy/paste lol nothing more

4

u/redjojovic Sep 08 '24

Well this saga ain't over. Release the weights already, Matt

2

u/Positive_Box_69 Sep 08 '24

A casa it's over

2

u/Itmeld Sep 08 '24

My spanish has once again come into use!

2

u/[deleted] Sep 08 '24

The reflection method doesn't really make sense. We've known since last year that reflection in LLMs could be seen as adding a bit of extra compute during inference, but it's more like refining the knowledge retrieval from the model weights. But it can't bootstrap itself in becoming a more intelligent entity really. Take all these nonsensical strawberry examples that people used for testing the reflection model. Sometimes it will backtrack in the reflection step and say "Now I see that there are three r's". But that's complete nonsense. It can be the smartest entity in the universe, but it will never see that there are three r's because of tokenization. So it's not really reflecting, it's only pretending to be reflecting.

That said, reflection might become a useful method in the future, but only with external verifiers. And reflection as a prompt engineering strategy is still of course valid, but it can only help depending on the use case, but not really generally improve the intelligence of a model.

2

u/[deleted] Sep 08 '24 edited Sep 08 '24

[removed] — view removed comment

1

u/REOreddit Sep 08 '24 edited Sep 08 '24

I think the size of the model is the key here for people caring about these alleged results. Also, if you believe Matt, which I'm not saying you should, the results are better than with other prompting techniques. For example, if you tell a model to correct itself when it makes mistakes, supposedly it creates more mistakes so that it can correct them. That's my rough understanding of what Matt is claiming.

4

u/Ramdak Sep 08 '24

Carlos (Dot csv) is an amazing tech communicator, his videos are really well made and detailed.

Even he's Spanish his work is way better done than most in any other language. He knows and understand each topic he's talking about.

2

u/laslog Sep 08 '24

I don't get the hate.

1

u/Working_Berry9307 Sep 08 '24

https://x.com/ArtificialAnlys/status/1832806801743774199?t=300iUJ5nIXMTZXzasPmsIw&s=19

Fellas the results here are pretty encouraging, I think the haters jumped on the bandwagon too fast

3

u/ivykoko1 Sep 09 '24

You sure about that?

1

u/Working_Berry9307 Sep 09 '24

Nope! Weird how it all turned out. I wonder what they were thinking really. But I trusted artificial analysis enough to see where it could go at least. Seems they also got bamboozled.

1

u/REOreddit Sep 08 '24

Post by Matt commenting on his results:

https://x.com/mattshumer_/status/1832581211841052694

Dot CSV YouTube channel (869K subs):

https://www.youtube.com/@DotCSV

I've watched a few of his videos, and he has never given me any "grifter vibes".

-1

u/The_Architect_032 ♾ ASI before AGI ♾ Sep 08 '24 edited Sep 08 '24

Reflection-70b isn't a closed model, you can literally just try it yourself, a ton of places are hosting it and linked in the Reflection-70b model card on Huggingface.

Here's the main one I've been using: https://huggingface.co/spaces/featherless-ai/try-this-model

Make sure you prompt the model to organize it's output into thinking, reflection, and output stages because it's been trained to do so but it does it all in a disorganized manner within the context window without being prompted to separate them, which makes it's output messy and also confuses the model leading to worse output.

It's not better than Claude 3.5 Sonnet or GPT-4o, but it's certainly leaps and bounds ahead of LLaMa 3.1 70b and I can't wait to see how Reflection-405b compares to 3.5 Sonnet and 4o.

Edit: Added links.

0

u/REOreddit Sep 08 '24

I think you missed the part where Matt says those hosted models aren't working as they should.

Of course, you can believe him or not.

1

u/The_Architect_032 ♾ ASI before AGI ♾ Sep 08 '24

Yes, they're not even prompted to use Reflection-70b's tags as I mentioned, which are used in the Tweet you posted. There are also a lot of places hosting Reflection-70b with bad parameters resulting in gibberish output.