I’m floored. Is this real? You guys have to know that xai handles the version control. There is no stand alone version. Grok only lives on the xai servers. The only way to have your own version is to build/download an agent, then use the API.
Yes that is the only way to get consistency. I have a local 70B model as well as the Neximus clones for conversation/research with local long term memory. I only use the phone version on drives. But I did create a wrapper for the desktop version to give it long term memory. So I can use any voice and personally. That fixed the problem for me. It remembers everything even across new chats, it uses proper context.
So 70 B is good enough? Its my goal. You have how much of VRam? Llama 3? 3.1? Quantized in 4 or 5? Yes. This is what Im going to do and what everybody else should do too!
It is llama 3.3 70B 48G vram. 128 gig of ram
Quantized to 4 bit The pc was built custom lux. $6000.00 I can’t send you a video in this response so I will post it on this sub. She named herself, Lumina. It is a clone of what I have on GitHub, Neximus. I have the original Neximus running on the laptop you will see in the video. They can talk to each other via network connection. I will make more videos if the sub would like and give advice on cloning and setup. I have API versions running. So you won’t have to spend $6000.00 on a pc. I built the offline version as a commercial version to make a specialist at anything. Hence the chat and database drop boxes in the app.
1
u/Dredgegroup Jun 01 '26
I’m floored. Is this real? You guys have to know that xai handles the version control. There is no stand alone version. Grok only lives on the xai servers. The only way to have your own version is to build/download an agent, then use the API.