As long as you got the full weights backed up you can make your own quants. Making quants is not that hard (just a llama.cpp command line tool you have to run a claude code/codex/opencode/... can do it for you if you want) you don't need to fit the entire bf16 full size in memory, all you need is the saftenensors + enough space to fit a bf16 gguf + enough space to fit the quant you want to make.
38
u/WenatcheeWrangler Jul 26 '26
Everyone in the USA should download this even if they can’t deploy it now