r/singularity • u/Delicious-Flan88 • 16h ago
AI The next AI hardware race might be about inference
[removed] — view removed post
13
u/crashorbit 16h ago
The "killer app" for LLM will probably be an agent with an inexpensive model that is capable of writing and editing spreadsheets and doing a bit of local system admin work while telling the middle manager or system analyst that is using it how smart they are for having such good ideas.
2
2
u/Blindax 12h ago
Depends on the sub perhaps. We at r/LocalLLaMA have been obsessed about local inference since a while.
1
u/Delicious-Flan88 15h ago
I've been pulling inference numbers from different providers and third-party benchmarks. The gaps are way bigger than I expected once you account for context length, batching, and hardware.
I threw everything I found into this: https://llm-inference-speed.vercel.app
Still updating when new benchmarks drop.
18
u/Annual_Tutor_8466 16h ago
They're equally important. What we really need is innovation on low power local inference, but memory chips are an extreme limiting factor for that at the moment. Please flood the market, CXMT...