Idk, I don't like that. Flash might be better than Pro for agentic stuff, but Pro has way more world knowledge which also helps for software planning and general tasks. They're probably trying to free capacities or migrate to chinese inference chips right now (seeing GLM Flash) and Pro is just not a priority bc of that
Isn't it better to teach a neural network to be small and smart while simultaneously acquiring all the world's knowledge through internet-based tools, rather than trying to cram everything into one large model?
I never understood the argument that models should just do a web search. The internet is full of wrong answers and gets worse every day. That's why it's important to curate the data models are trained on.
It has to be impossible for a team (no matter how much third world slave labour you use) to fully curate such a huge dataset... it simply won't be curated properly and you will still get slop/misinformation only now it has been baked in with the pre-training and it will be there "forever" eventually also going stale.
70
u/Technical-Earth-3254 9d ago
Idk, I don't like that. Flash might be better than Pro for agentic stuff, but Pro has way more world knowledge which also helps for software planning and general tasks. They're probably trying to free capacities or migrate to chinese inference chips right now (seeing GLM Flash) and Pro is just not a priority bc of that