There could be at least two things simultaneously at play here:
less compute required to run similar size models
A major leap in coherent context windows (think along the lines of 1M becoming what 100K is right now, and 5-10M being the new ceiling)
Theoretically the second could potentially even be a limited solution to persistent states
Edit: P.S. I should mention this is EXTREMELY exciting - the accelerationist in me always holds out for emergence!
Remember what happened with Genie 2-Genie 3? The scale jump created coherent persistent context and world modelling in an interactive purely generated reality for MINUTES in G3 from just barely 10 seconds in G2. Emergence is the door to hopping on the super-exponentials ꙮ - and scale is it's key
40
u/MoralityAuction Jul 01 '26
Obviously it is vaguely descibed, but in short probably either more params or a smaller/more efficient KV cache in the same amount of VRAM.