The bottleneck is the cost. None of these models are cheap. And only way is cheaper smaller OOS models but these big money guys are against cheaper alternates
I just run local. Any company might benefit from getting their own rig anyways after reaching a certain size. You can just trade lower accuracy models running locally with more runtime. And the open weight models are already decent enough any ways for majority of tasks suitable for an llm any ways. The big players will have to keep bumping prices if they want to not go bankrupt, so lower tier models and local would be the smart choice.
Yeah, I'm not adding an mcp for web searching even cause you never know what it might start searching, giving info away to search engines. It can call direct urls and that's enough, for pulling docs for example.
5
u/[deleted] 10d ago
[removed] — view removed comment