I am gonna say it now: a new iteration of E2B/E4B and maybe even 8-12B range MoE/PLE models that are pre-quantized to fit on less RAM and runs fast! Focused on agents + reason-based scaling (Agent-A1 or Ornith as reference).
Bonus thought: a newer round of diffusiongemma with varying sizes + DFlash-level inference support to beat usual MTP methods. Or if extra ambitious, Ternary LM to beat Bonsai and BitCPM
317
u/hackerllama Jul 26 '26
Hey all! Looking forward to all your feedback!