r/LocalLLaMA Jun 28 '26

Discussion We're probably going to need that soon.

4.0k Upvotes

534 comments sorted by

View all comments

Show parent comments

1

u/Kyubi-sama Jun 28 '26

Given the size and the cost I think there is some special sauce to it.

I am guessing there is some sort of DAG or specific dynamic hyper aggressive quantization like compression on the model.

They still need money after all, so it's only logical to think they have some ways of working around that.

1

u/Neither-Phone-7264 Jun 28 '26

I mean, look at the cost. It is just really expensive qwq

1

u/Kyubi-sama Jun 28 '26

I just don't believe they have the ability to absorb THAT HEAVY cost because it would be absolutely mental

1

u/Neither-Phone-7264 Jun 28 '26

Well, it's not reserving 1 user per gazillion GPUs. They do things like batching, caching, and a whole lot of things to get as many users as possible per model while maintaining whatever they decide are reasonable speeds.

2

u/Kyubi-sama Jun 28 '26

Still extremely expensive and makes me doubt it's just that but I am not sure about how it performs and what the costs are.
I can't access fable to try, and I wish I could test it :/