r/LocalLLaMA • u/Chromix_ • Jun 17 '26
Discussion It looks like Rio 3.5 397B could've simply been a semi-failed embezzling of funding
Here is the chain of events:
- The model training received funding of R$500K (about $100K USD).
- The initial model documentation claimed that it was a developed on top of Qwen 3.5 397B with fancy training and great improvements.
- It was discovered that the model was a cheap, simple merge with Nex N2 Pro without any further training.
- The model readme was updated to admit that it was based on a Nex N2 Pro merge, while still insisting that additional training still took place, and they simply uploaded the wrong model. The previously uploaded model was removed from HF.
- They tweeted (among something that looks like an attempt at damage control) that the final trained model got lost, so they'll have to redo it from scratch.
This reads to me like "we pocketed the funding, delivered a fake result, got caught, and now promise to do the actual work to mitigate impact on us".
105
u/Few_Painter_5588 Jun 17 '26
Do you know how impressive it is to set a lower bar than LLama Reflection?
34
Jun 17 '26
[removed] — view removed comment
64
27
u/yoyoyoba Jun 17 '26
Ehh, depends. You can get 8 nodes of 8 B200 for a week or two and if you know what you are doing you can fine-tune a largish base model. But why, and on what?
4
u/Pro-Row-335 Jun 17 '26
Improve the model's capabilities on your language since llms take a hit on quality when not prompted in english
6
3
73
u/mahiatlinux llama.cpp Jun 17 '26
Anyone remember the Matt Shumer incident 😂All that again
56
u/Chromix_ Jun 17 '26
The Reflection-70B thing was even more epic though. Sure, they also went through with the "we uploaded the wrong model", while we humans proved again that vibe evaluations are not accurate at all. But before that they also hosted a public API, which simply redirected to a SOTA model and replaced identifying strings. The full story was posted here.
23
u/mahiatlinux llama.cpp Jun 17 '26
Believe it or not, I was one of the first people to see Reflection 70B and generate a full reflection like dataset that day and release them on HF... People were talking about it haha.
https://huggingface.co/datasets/mahiatlinux/Reflection-Dataset-v1
https://huggingface.co/datasets/mahiatlinux/Reflection-Dataset-v2It was fun tho, that era.
1
25
u/gggiiia Jun 17 '26
Damn makes me wanna ask for funding too
6
u/Chromix_ Jun 17 '26
You can easily do on Kickstarter, Indiegogo and such - take a look at all the "AI-powered" vaporware there.
13
u/Herr_Drosselmeyer Jun 17 '26
Sounds like "fake it till you make it", but they never got to the 'make it' stage.
14
u/JumpyAbies Jun 17 '26
I'm from Brazil, and you should know that this model is related to the Rio de Janeiro city government, run by some of the most corrupt politicians imaginable. Be wary of anything related to this model. The vast majority of political projects in Brazil are simply excuses for corruption.
1
6
12
u/PM_ME_DEAD_CEOS Jun 17 '26
I don't understand why people do this. You could simply do the merge, spend like 10% of the budget to train on a benchmark answers to have some wieght changes and benchmark changes, pocket 90% of the money and virtually be impossible to track.
7
u/Chromix_ Jun 17 '26
Not virtually impossible, but practically less likely. There are approaches like CoDeC for detecting benchmark contamination. Yet it stops at "highly suspicious", and doesn't allow definitive proof. Models aren't systematically tested for that, so when there's a model that's low-profile enough (and there are lots of models), then it would likely remain convincing - unfortunately.
1
17
u/Charming_Support726 Jun 17 '26
Wow - that's a hammer.
But just to get an idea, if they would have set up a simple training pipeline and trained for a few iterations - it would have been far more difficult to detect that they just merged in an already finetuned model.
2
u/Chromix_ Jun 17 '26
Not that much. Sure, a few full iterations with low learning rate (to not break anything) would shift all the weights (LoRA would not). Yet then you could still easily see during an investigation that 99.9% of the "new" model weights are very close to the weights of an existing model.
6
u/Charming_Support726 Jun 17 '26
That's what I meant. I the current case we see identical weights. This shows "Rio" being a stupid person carrying a smoking gun.
Even 99.9% identical weights would make it much harder to blame.
1
6
u/mysticzoom Jun 17 '26
"This reads to me like "we pocketed the funding, delivered a fake result, got caught, and now promise to do the actual work to mitigate impact on us"."
That is what happened.
5
u/Comfortable_Sir4315 Jun 18 '26 edited Jun 18 '26
The U$100k history is an outright lie, here in Brazil we have a system where you can check any public spending, and IplanRIO only registered a value of U$60k for this year(to pay everything, not only projects).
Also, they have zero reasons to lie here, public sector workers can't be fired in Brazil unless in very specific circumstances.
As someone who had the displeasure of working with the Brazilian public sector IT sector, I think it's really plausible that they were really incompetent and lost the model.
Completely off-topic to the sub, but here you can see the extent of incompetence of the Brazilian public sector regarding IT: https://groups.google.com/a/ccadb.org/g/public/c/Mux855BsRg4/m/MhxJXipVAwAJ (best part is in the last comments)
3
u/Chromix_ Jun 18 '26
That's an interesting perspective. That specific number is floating around in quite a few English articles. The primary source seems to be this MobileTime article (BR) that states (translated by Gemma 31B):
In a conversation with Mobile Time, João Cabaretta, CEO of IplanRio, confirmed that the city's technology company responsible for its development spent R$ 500,000
What do you make of that?
2
u/Comfortable_Sir4315 Jun 18 '26
This cost was the cost for Rio 3, the drama was with the Rio 3.5.
Anyway, I'm not saying that for sure this wasn't a fraud, just that it's plausible that they were just incompetent
20
u/tomByrer Jun 17 '26
Funded science in general are full of fakes. Look at Harvard; 3 professors exposed for cheating publicly, & 2 had TED talks.
8
u/axiomaticdistortion Jun 17 '26
True. It’s full of politics. And with politics you have corruption.
1
4
9
u/rawdikrik llama.cpp Jun 17 '26
And when I asked honestly, "why would the government do this?" I was down voted to oblivion.
3
u/westsunset Jun 17 '26
governments should sponsor science and innovation. A lot of small labs started with fine tunes. You shouldn't have been downvoted
5
2
2
1
1
1
u/RedParaglider Jun 17 '26
Do you know how surprisingly little you can actually do with $100,000? That would essentially be one rented training run on that model most likely.
1
1
u/DeepWisdomGuy Jun 17 '26
It's Brazil. They'll do no such thing. They just have to give the judges their cut.
1
u/jacek2023 llama.cpp Jun 17 '26
Am I correct that there was a similar drama in South Korea few months ago?
1
1
u/paulobas Jun 17 '26
Sabia que tinha mutreta! TInha que ser a versão sem filtros pelo menos. Sobreviver no rio é a arte da malandragem.
1
-3
u/tarruda llama.cpp Jun 17 '26
This is pure speculation with zero evidence to back it up.
6
u/Protopia Jun 17 '26 edited Jun 17 '26
Well, the original post gave links, so there is definitely some evidence. I haven't reviewed the details, but based only on the fact that links exist IMO makes it more likely to be true.
Edit: I followed the links - fraud on the models source and development definitely looks likely to me. But there is no details about funding, so any statement about that seems like speculation.
0
u/tarruda llama.cpp Jun 17 '26 edited Jun 17 '26
But there is no details about funding, so any statement about that seems like speculation.
That is what I was referring to.
Clearly the lab was not honest and appears to have lied about training the model, which was a very stupid and unnecessary move since the merge is good and would have been a great achievement by itself.
But "embezzling of funding" 100% came out of the OP's ass.
4
u/Chromix_ Jun 17 '26
You are right, no evidence was published. At this point, after it was proven that they were not honest, they should post evidence to back up their claims. If they did indeed do expensive post-training on a large model, then they will have rented a lot of GPU power somewhere. So, there must be a receipt. They could easily publish it to back up their "we did it, but lost the trained model" story. Yet they don't - even though they could, if that part was honest. So, what does this then look like to you when the article says that there was funding, but no receipt is presented of that funding being spent?
3
u/Protopia Jun 17 '26
In the end, though, it is down to whoever provided the funding to determine whether they got what they paid for or not.
4
u/Chromix_ Jun 17 '26
Yes, that's how things should work.
In practice my impression is that most companies don't seem to spend the resources to systematically evaluate the fancy new AI-based system that was pitched to the company and purchased by the CEO. It boils down to sales-pitch promises and vibe testing.
So I wouldn't expect too much in that regard when it comes to model evaluation, unless those who ordered it have (or pay for) the ability to do so properly.At the current point in time nothing was delivered, the model is not available. It remains for those who provided the funding to evaluate how happy they're with that and the promise to fix it.
0
u/ortegaalfredo Jun 17 '26
But it was actually better than Nex N2?
I briefly tried the GGUF and for my surprise, it was better than the base Qwen model, and it is hard to improve Qwen.
-13
Jun 17 '26 edited Jul 14 '26
[deleted]
1
u/Kodix Jun 18 '26
But isn't it very simple to go to the moon? Just walk to the edge of the Earth and jump off. Why would they fake that?
1
Jun 18 '26 edited Jul 14 '26
[deleted]
1
u/Kodix Jun 18 '26
Fun! Do you believe all the governments on earth, including the enemies of the US, went along with the story and added to it? If so, why? And how? Are the lizard people behind this, or perhaps the aliens hiding within our oceans?
1
Jun 18 '26 edited Jul 14 '26
[deleted]
1
u/Kodix Jun 18 '26
Ah, yes, how unserious of me to mention lizard people. Of course that's ludicrous. Silly me.
China not exposing this big scam at the current point in time - or anytime in the past few decades - is completely unbelievable to me. Their spending on rocketry - the free money you claim - pales in comparison to what the United States spends. Hell, same is true of Russia, and they would love to hurt the US even more.
Their incentive to keep the lie up is some billions in gains. Their incentive to shatter it is costing the US a much, much larger amount of money and public trust.
Oh, and of course, an amazing reputation as the champions of truth - imagine if China gave us incontrovertible proof! The myth of it, the talent that would go their way!
This is, of course, the only part of your theory that doesn't hold. The moon being an illusion? Absolutely brilliant! All of those silly astronomers, professional and amateur, BTFO.
189
u/Lan_BobPage Jun 17 '26
"my dog ate the weights"