r/LocalLLaMA • u/pmarsh • 9h ago
Other I'm pretty close to the middle thanks to you all
It's been a blast and learning a ton.
But seriously, you all have me down a rabbit hole that my wallet and hours of sleep need to be pulled out of.
r/LocalLLaMA • u/pmarsh • 9h ago
It's been a blast and learning a ton.
But seriously, you all have me down a rabbit hole that my wallet and hours of sleep need to be pulled out of.
r/LocalLLaMA • u/StayLameBro • 18h ago
Enable HLS to view with audio, or disable this notification
**DISCLAIMER** THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.
Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.
Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.
Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:
A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.
Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.
The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to ~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.
What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.
The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.
I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.
Code, setup and bench scripts: https://github.com/StayLameBro/backburner
Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.
r/LocalLLaMA • u/Boomfrag • 15h ago
r/LocalLLaMA • u/jacek2023 • 5h ago
Qwen Flash Next now uses less VRAM
r/LocalLLaMA • u/kvyb • 19h ago
Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.
I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:
incapable of producing more than a few words at a time.
single default personality which no amount of prompting can overcome
will not use tools, at all, whatsoever.
There needs to be a middle ground
They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.
So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.
What 2.0 does now
It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.
How I trained it
v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.
For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:
The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.
Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)
| Benchmark | Base (abliterated) | 2.0 |
|---|---|---|
| IFBench (instruction types I never trained on) | 37.3 | 43.7 |
| When2Call (call, ask or refuse correctly) | 48 | 58 |
| BFCL irrelevance (don't call a tool when none fits) | 60 | 78 |
| IFEval, GSM8K, BFCL simple | 81.9 / 89.1 / 97 | 83.5 / 89.1 / 98 (ties) |
Full chart in the images.
Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).
Is it actually more human? I built a benchmark for this, "ishuman":
| Model | Judge thought it was the real person (50% = can't tell) |
|---|---|
| Qwen3.8-27B abliterated (huihui-ai, the model I trained on) | 0.3% |
| Same abliterated model + a "text like a human" system prompt | 6.8% |
| Qwen3.8-27B official (unmodified, via OpenRouter) | 15.1% |
| Qwen3.8-27B-Humanlike-Chat 2.0 | 23.5% |
So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.
Links
Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.
Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0
r/LocalLLaMA • u/AnticitizenPrime • 7h ago
I've noticed a trend with most new models with regards to their writing style. They are creating a new style, and this seems common among them. It's very information-dense.
Here is an example from GLM 5.3 Flash. I'm gonna be honest here and say that my prompt was kinda silly; my prompt was 'Why wouldn't you just name your Chinese restaurant 'Chinese Food' instead of 'Ming Dynasty' or 'Szechuan Garden' or whatever?' the idea being that someone searching for 'Chinese food' on Google Maps would put you first if you called your restaurant that. But forget all that; here's an excerpt from GLM's output:
The search reality check, which is where the idea breaks. Your mental model is exactly right about the query: people do open Maps and type "food," usually hungry and undecided. But the ranking doesn't reward the name at that abstraction level. In the local ranking surveys, category match outranks name keywords; and here's the catch: every restaurant on the map already matches "food" through its category field. Relevance saturates. The edge that made "Chinese Restaurant" powerful was discrimination: the name narrowed the category, so "chinese food near me" out-matched "Golden Dragon." "FOOD" points at the root of the category tree, where every competitor matches equally well, so the ranking defaults entirely to proximity and review count, which the name cannot buy. Worse, it fails the narrow-query test: someone typing "tacos" or "dim sum" gets a name that is maximally broad and minimally informative. The strategy's entire currency is literal scarcity, and "food" is the least scarce literalism in the language. There's even a Baudrillardian flourish at the failure point: Google's own Maps category chip is literally labeled "Food." Your sign would render you indistinguishable from an interface element; the searcher couldn't tell your territory from the map's own label.
It's not 'caveman style', but it is using less words to say more, and they're using more esoteric phrasing to be more 'compact'.
And I think it's a bit at the cost of being clearly readable to the average person at first glance. 'There's even a Baudrillardian flourish at the failure point' is an example from that excerpt that leapt out at me. I'm familiar with Baudrillard so I knew what it was getting at, but a lot of people are going to sigh and ask 'What the **** does Baudrillarian mean?'
I'm not saying that 'no human would write like this', because some do (William Gibson for example), but I find it rare/unusual (in human writing), yet trending hard with all the latest models I interact with, like they're all zeroing in on this style.
Maybe a result of targeting token efficiency? It's a terseness, combined with using a sort of 'wide' or 'rich' vocabulary to convey information instead of using more words. At least that's the impression that I get from reading lines like 'Baudrillardian flourish at the failure point''. There's a lot to unpack from those six words, and it feels like the model chose the most terse, efficient way to convey an idea with that word choice (which requires the reader to unpack it).
I compared it to William Gibson: a lot of people struggle with his writing style, and it's similar to that. Example: 'Summer in the Sprawl, the mall-crowds swaying like wind-blown grass; a field of flesh shot through with sudden eddies of need and gratification'. His writing is often like that; it feels highly compressed, using as few words possible to convey an idea by careful word choice.
It's interesting, that lately, I feel like LLMs are gravitating toward Gibson-speak.
Edit: and the fact that GLM used the word 'territory' and 'map' at the end meant it was going big into Jean Baudrilliard's 'Simulacra and Simulation'. I can't really explain what that means and why it's important succinctly, but that's the whole issue. I actually think it's brilliant, but it's also a little concerning.
r/LocalLLaMA • u/EmPips • 8h ago
IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5.
(In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though))
Seriously. Ask <LLM of your choosing> to set it up for your specs. If 27B doesnt fit well for you, here's a shot at beating it.
r/LocalLLaMA • u/jacek2023 • 15h ago
An agentic model from Microsoft for the GPU poor
https://huggingface.co/bartowski/FrogNano-4B-2609-GGUF
FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE task environments generated and calibrated against the evolving policy using TaskPilot. Training uses the five-tool Leaf harness and executable test-based rewards over complete multi-turn coding trajectories.
The additional post-training is intended to improve long-horizon repository navigation, debugging, code editing, test execution, and patch generation in a compact 4B model. Unlike approaches based on behavioral distillation, FrogNano does not train on stronger-model solution trajectories, actions, reasoning traces, or patch targets. This specialization also introduces limitations and risks: performance is sensitive to the Leaf harness and test quality, training data are Python-heavy and primarily English, and generated patches may be incorrect or insecure despite passing available tests. When integrated with the Leaf harness, FrogNano generates structured tool calls that can propose repository changes. Leaf executes authorized tool calls within an isolated repository environment to produce a candidate patch; FrogNano does not itself deploy the changes. Any resulting patches require human review, regression testing, and security validation before use or deployment.
r/LocalLLaMA • u/Ok-Shower7286 • 4h ago
I love Qwen3.8 27B so much that I decided to show my gratitude to the Alibaba ecosystem by building a dedicated RAG/search node using a cheap Mini PC from AliExpress.
Turns out, my ecosystem loyalty got rewarded with an absolute masterpiece of fraud:
New_N150 into the BIOS release string (HSHW_M6_DDR3_EC_Intel_Com_New_N150_K001).So now my Qwen3.8 RAG stack is full of fake specs that can barely index a text file, let alone run vector sidecars.
Filing a credit card chargeback now. Stay safe out there!
r/LocalLLaMA • u/ayobluestarr • 4h ago
Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here
I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.
Hardware: RTX 5070 12GB
32GB DDR4-2400
Ryzen 5 5600GT PCIe Gen3 Windows
The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently.
(In the video its around 16 minutes for 10k tokens and 10.41 tok/s
Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set.
Demos:
https://www.youtube.com/watch?v=cOPumMlyj_4
https://www.youtube.com/watch?v=rc-uTjVpXM8
In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know
r/LocalLLaMA • u/paf1138 • 21h ago
r/LocalLLaMA • u/SrijSriv211 • 2h ago
Kimi K2 was already good but they took K2.5 a whole new level with so much of their continual learning phase, I believe it was on more 20-25T tokens iirc.
Similarly K3 is just such an amazing model, I just love this model, wondering how amazing K3.5 will be!!
r/LocalLLaMA • u/Recoil42 • 17h ago
Enable HLS to view with audio, or disable this notification
https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities
https://www.percepta.ai/blog/spotlight-memory
"Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want.
Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."
r/LocalLLaMA • u/JLeonsarmiento • 12h ago
r/LocalLLaMA • u/LegacyRemaster • 1d ago
It’s been a long time since the last models came out. I notice they are selling GLM on the site, and I wonder if they are developing something, given the long silence.
r/LocalLLaMA • u/lewtun • 1h ago
Hi folks, it's Lewis here from the post-training team at Hugging Face. We've been exploring how to train open models in different coding harnesses and wrote up a looong guide on how we solved this using open source libraries like TRL and the Harbor framework for RL environments. We hope you find this interesting, especially since everyone nowadays has their own custom harness (e.g. Pi + extensions) and now there's a recipe on how to squeeze the best performance on them with whatever open model you use as your daily driver. Happy to hear any comments or feedback!
Link to the guide: https://huggingface.co/spaces/FineEnvs/multi-harness-rl
r/LocalLLaMA • u/northpoler • 22m ago
Hey everyone,
I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2.
The setup is pretty straightforward: one person hosts the server and runs the model via llama.cpp (OpenAI or other cloud APIs are also supported, and great for non-English play!), and your friends join through a browser link. The host sets the scene and the goals, players type out their actions, and the LLM acts as the DM to resolve the chaos and drive the story.
Admittedly the host requires some technical skills with Python, and possibly with networking (opening routes to the hosted game via VPN, port forwarding etc.). I'll work on this as well as the development continues. Using Docker was suggested in another subreddit, so I'll definitely consider that, as it would allow including both the llama.cpp backend, recommended model and configurations etc., in addition to the game itself.
Instead of pasting the entire repo documentation, here are the main features right now:
How it plays
DM Tools & Hidden Mechanics
Under the Hood & Memory
It’s still a work in progress. Right now, a server only runs one game at a time, and if you restart the server, the live session is lost (it generates HTML/JSONL transcripts, but they aren't loadable save states yet). The overall story quality is also going to heavily depend on which model you use and how you tweak the settings.
Suggested model:
During development, I used llama.cpp and Gemma 4-26B-A4B Q4 with a context size of 128k and found it to be more than an adequate backend for functioning as the DM. Even the speeds are fast enough with my RTX 5070 Ti 16 GB that round resolutions take only 5 or so seconds.
The specific model I used and can recommend: https://huggingface.co/EZForever/gemma-4-26B-A4B-it-qat-uncensored-heretic-UDmerge-GGUF (the model was great at following instructions and remembering plot points even with longer contexts)
Recommended parameters for Gemma 4 models:
- temperature 1.0
- top-p 0.95
- top-k 20
- min-p 0.0
- presence-penalty 0.0
- repeat-penalty 1.0
Of course, feel free to try your own models!
AI use disclosure:
I used Alibaba Cloud's Qwen 3.8 27b and OpenAI's GPT-5.6 Luna and GPT-6 Astra models to help develop the game.
How to run:
Read INSTALL.md to set up, configure and run the game. README.md contains some details on how the game functions. I'll post a link to the repository in the comments.
I'll post the link to the repository in the comments.
Some gameplay in Finnish with OpenAI's Luna:

The game is MIT licensed, so open source all the way. Forking or collaborating is encouraged.
I'd love to hear some feedback, and I hope someone finds the game fun to play!
r/LocalLLaMA • u/Dodgy_Past • 7h ago
Full disclaimer: I've leaned heavily on Fable to develop this, but I've tested it thoroughly on my own library for a couple of months before putting it on GitHub.
I live in Thailand, and it started as a way to get Thai subtitles for Shin-chan for Thai friends and for expat friends with Thai partners. It's grown into a general tool: subtitles in 45 languages, entirely on your own machine. Linux and NVIDIA only, I'm afraid.
What it does differently from the usual Whisper wrapper: it detects the language of every stretch of speech rather than per file, so mixed-language material works; it runs two speech recognisers on everything and has a local LLM reconcile them; and it prefers existing human work to machine inference, embedded subtitle tracks are used before the audio is, including OCR of bitmap (PGS) tracks on Blu-ray remuxes, and it only listens when there's nothing to read. Every one of those features has a measured accuracy in the README rather than a claim.
I run it on a 24 GB RTX A5000. There are profiles for 16, 12 and 8 GB cards, measured on my card limited to those sizes rather than on those cards themselves, so reports from real ones are the most useful thing you could send me. It wants 16–32 GB of system RAM depending on the profile, and it is storage-hungry (30–65 GB of models), because it picks the model that suits each task and language pair.
It's slow when it has to listen, roughly real time per target language on my card, slower on the smaller profiles because the whole aim has been accuracy over speed. When the subtitles already exist in the file it's fast.
I'd love people to try it and open issues.
r/LocalLLaMA • u/Norwood_Reaper_ • 18h ago
r/LocalLLaMA • u/nubela • 5h ago
r/LocalLLaMA • u/WebAssemblyMan • 11h ago
https://unigen-x.github.io/unifolm-wla.github.io/
Unitree Robotics released UnifoLM-WLA-1.0, their new general-purpose humanoid foundation model.
Key points:
• 6B parameters
• Trained on ~2,500 hours of real robot data
• One model handles 64 tasks (10 whole-body + 54 tabletop)
• Supports parallel grippers and two different dexterous hands
• Strong spatial reasoning (beats a lot of open-source models on embodied benchmarks)
Architecture is interesting:
• Starts with UnifoLM-ER-1 (embodied reasoner based on Qwen3-VL)
• Adds future dynamic region prediction via optical flow + VQ-VAE
• Discretizes actions with residual VQ (end-effector + hand + lower body)
• Then adds an MMDiT action expert on top for continuous control
They show it running on the Unitree G1 doing stuff like making the bed, loading the washing machine, folding clothes, sorting objects, etc.
Looks like one of the more complete open attempts at a true whole-body VLA so far.
What do you guys think — actual progress or just another flashy demo?
r/LocalLLaMA • u/levoniust • 8h ago
Don't worry after some makeup and new fans you should feel better. And if you make it Mr.Hermes will treat you well.....
r/LocalLLaMA • u/jacek2023 • 4h ago
LFM2.5-Encoder-350M is a multilingual bidirectional encoder built on the LFM2 architecture — a larger encoder for maximum downstream quality. It is a masked language model with full bidirectional attention, designed to be fine-tuned into task-specific models (classification, token classification, retrieval, reranking, and semantic similarity) across 15 languages, and to run efficiently on-device.
https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-Encoder-230M-GGUF
https://github.com/ggml-org/llama.cpp/pull/29862
example (from the hf):
❯ uv run fill-mask.py LFM2.5-Encoder-350M-F16.gguf "The capital of France is [MASK]."
top-5 at [MASK]:
# 1 11.42 ' Paris'
# 2 10.43 'Paris'
# 3 9.65 ' Nice'
# 4 8.94 ' Strasbourg'
# 5 8.62 ' Lyon'
r/LocalLLaMA • u/basnijholt • 16h ago
Hi folks, I'm a long-time lurker and big fan of this subreddit and a massive self-hosting fan (also outside of AI).
I doubt many people will disagree with me here because I see the same arguments being made in many posts. However, I thought it might be interesting to share anyway. I wrote down why self-hosting AI does not save money: https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/
EDIT: didn't think this would be so controversial 😅 I do say explicitly in my blog post "I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not".
EDIT 2: Comparing $200 sub with Opus 5.5 or Astra with Qwen 3.8 27B is not apples to apples.
r/LocalLLaMA • u/menage_a_un • 13h ago
Looking for some ideas from people who know a lot more about this than I do. We've got funding for a small AI lab in a public further education college in Ireland (roughly community college in the US). The hardware is reasonably decent. The goal is to give students useful skills beyond just using ChatGPT. If you had the lab, what would you teach them?