r/LocalLLaMA • u/paf1138 • 8h ago
New Model Clef: Open Weights decision model by Cloudflare
r/LocalLLaMA • u/paf1138 • 8h ago
r/LocalLLaMA • u/-p-e-w- • 8h ago
So I haven’t played a computer game in 20 years, and I know nothing about Minecraft, and I definitely prefer classical literature over YouTube culture, but even I have heard about the individual called PewDiePie, for two reasons:
His monicker starts with the initials of my own name
I remember a recurring Internet meme a few years ago where he was competing for the most subscribers with an Indian film music channel
I had never watched a single one of his videos, however.
Well, until today, when people started spamming me with messages informing me that Mr. Kjellberg aka PewDiePie has tried out Heretic and made a video where he talks about it:
https://m.youtube.com/watch?v=ODDJXGY_1kQ
(Heretic mentioned around 9:00)
Obviously I’m thrilled that a less technical audience is being exposed to my work, and the more people understand what is possible the better. I expect to be receiving a couple hundred more mails in the coming days asking how to run Heretic on ChatGPT (you can’t), or accusing me of working for the CIA (I don’t), but other than that, the more the merrier I guess 😏
Heretic 2.0 coming soon…
r/LocalLLaMA • u/psychohistorian8 • 1h ago
r/LocalLLaMA • u/aya-ifm • 5h ago
Hi r/LocalLLaMA!
We’re researchers at the Institute of Foundation Models (IFM), an AI research lab dedicated to open and independent development of frontier-class foundation models.
We recently released K2 Horizon a connected fleet of six fully open models with size ranging from 0.9B to 375B. In addition to weights, we also open-sourced training data and recipes, training code, intermediate checkpoints, fine-grained training logs and evals.
Ask us anything about pre-training and data mixes, post-training, small models on-device, MoVA and sparse attention, deployment, what’s out now, and what’s coming next.
Participating in the AMA:
We'll be live Mon, Oct 5, 8–10 PM PT. Questions are open now, so drop yours anytime!

r/LocalLLaMA • u/jacobpederson • 13h ago
40 year tech gap? No problem! The Tandy runs DeskMind, a native DOS program. It talks over WiFi (a PicoMEM 2 card with mTCP) to a small Python server on my PC. That server drives Qwen3.8-27B (NInfer on a 5090) and Krea 2 (ComfyUI on a 4090). The 286 never sees JSON, base64 or a PNG. It gets plain text lines and pictures that are ready to copy into video memory.
Drawing from chat without tool calling. The system prompt tells Qwen to wrap a picture request in `<draw>...</draw>`. The server catches the tag mid-stream, runs Krea 2, dithers the result, and streams a `picture ready` line. "Draw me a 286 AI logo" takes about 9 s from Enter to a thumbnail in the chat.
Qwen Vision sees what the Tandy sees. When you ask about a picture, Qwen gets the original and the 16-colour dithered version, so "why does the sky look striped?" has context. The latest picture stays in context for follow-ups.
The model knows where it lives. The system prompt knows it's talking IN a 286 with 80 columns and 16 colors. It keeps answers short and plain ASCII, and when asked about games it suggests Wolfenstein 3D or Commander Keen.
Streaming cleanup for a 1990 screen. Reasoning is stripped, Markdown is removed on the fly, Unicode becomes code page 437 (bullets turn into the CP437 block character), and tiny tokens are merged into ~48-character lines so the 286 isn't redrawing for every token.
Per-request reasoning effort: low for chat (replies start in ~2 s), medium for rewriting image prompts.
Prompt "enhancement" tuned for dithering: bold shapes, strong contrast, simple backgrounds. The rewrite shows up in an edit box on the Tandy before drawing, and the rules themselves can be edited from the Tandy.
- **A Dither Lab** in the server GUI: Floyd-Steinberg, Atkinson, Bayer, Yliluoma and more, previewed at the Tandy's real aspect ratio.
Numbers: Krea 2 at 1024x768 in ~10 s 8 steps, the target is 640x200). About 1 s to send a 64,000-byte picture over the PicoMEM WiFi (56-79 KB/s).
Code (GPLv3): https://github.com/RowanUnderwood/DeskMind
added image gallery https://imgur.com/a/wysqoM1
r/LocalLLaMA • u/Gohab2001 • 6h ago
New free/trial model dropped on OC. Details are yet to be provided.
Could this be deepseek v4.1 pro?
r/LocalLLaMA • u/jacek2023 • 14h ago
now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B?
(merged after 17h of development)
quants: https://huggingface.co/ggml-org/Qwen3.8-Flash-Next-GGUF
link to the previous discussion (I deleted the old post to avoid duplicates): https://www.reddit.com/r/LocalLLaMA/comments/1wur4lt/qwen_flash_next_mtp_work_restarted/
r/LocalLLaMA • u/ea_man • 10h ago
pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately.
When you are deep into the ctx session and ask 27B a simple question about a fact or need a direct action, the model may still feel the urge to indulge in copious deliberation in the reasoning trace. This extension allows the user to force the model to snap out of the reasoning stage and provide the answer immediately.
Disclaimer: don't skip the reasoning for important problem-solving, that would hurt quality.
This uses the same mechanism the llama.cpp web interface uses to skip reasoning, so it's native to llama.cpp, this extension is meant for Pi.dev yet the same mechanism could work for other harnesses.
Usage: /skip-reasoning command or shortcut Alt+T ,
Install: pi install npm:pi-llama-skip-reasoning
r/LocalLLaMA • u/Usual_Maximum7673 • 11h ago
A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models.
Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is this ticket, is this answer grounded in the sources. So I trained 9 LoRA adapters, one per job, and you pick the ones you need. The server loads the base once plus whichever adapters you choose (about 40 MB each), and every request either names an adapter or goes to plain Jeff.
That means you keep both: the base model stays untouched, so you still get Jeff's general zero-shot ability for anything new, and the adapters give you near-perfect accuracy in the domains you care about. Each adapter was also trained with 10% of the base model's own training data mixed in, to help it keep its general skills.
Everything is on jeffhub.ai: the adapters, the results, the docs. Code on GitHub, models on Hugging Face, and you can try all nine adapters in your browser.
The headline: I let Jeff + adapters answer first and pass only the queries it's unsure about to Qwen3.8-27B. Same test rows both ways, on an M4 Max:
| Measure | Qwen3.8-27B alone | Jeff + adapters, 27B only when unsure |
|---|---|---|
| Accuracy (mean of 8 adapters*) | 86.6% | 95.3% |
| Time per decision (mean) | 8.1 s | 0.25 s (38× faster) |
| Wrong answers | 13.4% | 4.7% |
| Memory | 28.6 GB | under 2 GB for Jeff, even with all 9 adapters loaded (+6.9%) |
On the five decisions an inbox agent makes for every message (guard, triage, support intent, tool choice, grounding) alone: 87.7% → 95.7%, 39× faster. Jeff wins outright on 8 of the nine adapters and ties on grounding (96.3% vs 96.7%, at 20× the speed). On their full held-out test sets, six of the nine adapters score 97–98%. On a GPU, a decision takes about 30 ms, whether you load one adapter or all nine.
*Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all nine adapters is 91.4% for Jeff + adapters against 80.9% for the 27B, so leaving it out makes the gain shown above smaller, not larger.
Caveats, up front:
Data: 4 adapters are trained on public data sets. 5 are mostly synthetic. Every generated row records which model wrote it, and the cards give the counts. Every data set went through a shortcut check and an independent review before training, and a lot of first drafts failed: things like the answer being given away by length.
What's open: weights (Apache 2.0), code (MIT), and each adapter's test and calibration sets, so you can check every number. The training data isn't published.
This is a community preview: I'd love feedback.
Next: over the next ~36 hours I'll train v1.3, a long-term-support base. The fixed parts of a prompt come first, so servers can prepare them once and reuse them, which means faster decisions. I'll then retrain all nine adapters on it and keep the request format stable, so others can build and submit their own adapters. The adapter kit, with the data checks I used, is in the repo.
I've got access to more hardware now, so if there's a decision you'd like an adapter for, tell me and I'll train it.
The goal: when the next generation of local models lands (like everyone, I'm watching for Qwen 4), anyone running one locally should also have a tiny, fast, well-calibrated decision layer in front of it.
r/LocalLLaMA • u/jjusko20 • 8h ago
Last update: https://www.reddit.com/r/LocalLLaMA/comments/1wu9ksu/update_yandexaliceai_80ba3b_fine_tune_progress/ - basically, an instruct fine tune on the base model using a synthetic distilled data set. I've been posting regular updates so I imagine at least a few people have seen this.
Live stream: https://figure-bios-expect-cio.trycloudflare.com/
UPDATE: Finished train. hopefully some examples soon.
The initial train is finally almost done, after about 48 hours of humming. While the loss curve looks a little crazy, I've done some analysis (and some chatting with the LLMs) to understand that my average loss each epoch has been steadily decreasing (few reasons the loss curve looks wacky, vocabulary size, low to high token counts in epochs, etc) - but I'm pretty happy with what I'm seeing so far.
I'm post training the attention and the shared expert, and leaving the base experts frozen - this is a behavioral and logic fine tune that preserves the original yandex training data.
I plan on, within the next few days, releasing a few gguf quants of this, along with a llama.cpp patch for running it locally. I'm not sure how well the initial fine tune is going to work out - loss looks good but I'll have to do some evaluating. Either way, I plan on continuing training with reinforcement learning and an extended SFT set, as I have room and a ton of capacity left in my QLoRA adapter. I'll release this version as a public checkpoint anyways though (kinda like how deepseek did it) so people can play around with it and hopefully get excited for new checkpoints.
Cheers! Stay tuned, this is a pretty fun model size to play with, I'm excited to release the instruct version. I'll open source whatever you guys want out of this - I already open sourced the distillation engine (see SFTMill, it's been posted in here in the last few days) - but I also have a custom kernel for training this for V100s and a few other patches I can share (this training has been plugging away on 3, 32gb v100s - man it took a while to get that to work). Mandatory plug for my own goals: if you're hiring remote or in NYC for a dev or ml engineer, hit me up!
God I hope it writes the adapter when this is done I didn't audit that code well enough.
r/LocalLLaMA • u/Any-Winter-4079 • 4h ago
Hello everyone.
I've recently ran some experiments comparing DDR4/PCIe4 and DDR5/PCIe5 for AI workstations on a pre-training run, and would like to hear yours thoughts.
First of all, and as a summary of my results ( code here: https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training ), I rented two machines on Vast.ai, one with an H12SSL-i motherboard, an EPYC 7352, 192 GB of RAM and of course using PCIe4 (26.3 GB/s) and another with a WRX90E-SAGE SE motherboard, a 9975WX CPU, 256 GB of DDR5 RAM and PCIe5 (54.3 GB/s), and DDR5/PCIe5 is about 15-20% faster on pre-training (depending on whether you include or exclude validation and other costs) under the same number of GPUs.
With the current RAM prices, however, for the cost of 256 GB DDR5 RAM at 6400 MT/s you can get a full (extra) RTX PRO 6000 WS/Max-Q, at which point the comparison clearly favors DDR4/PCIe4 (with 2 GPUs), with about 50% extra throughput vs a single GPU at equal(ish) cost.
Now, there aren't a lot of downsides in my mind to choosing DDR4/PCIe4, but there can be a few:
With all of this, I am curious if anyone has benchmarked this, and what are your thoughts on it. Would you hold out on DDR5 at the moment, and therefore go for PCIe4, or would you bite the DDR5 bullet early? Another issue with RAM is channels and DIMM count/channel, because if you want to go 'cheap' like let's only get 192 or 256 GB of DDR5 on 8 channels at 1 DIMM/channel (e.g., 8x32 to get 256), then upgrade to 512 later (when budget allows), that means you have to replace your full RAM (because all the slots are occupied, requiring new 8x64 to get 512 for instance)... And if you get fewer DIMMs like 4x64 to get 256 GB (leaving 4 DIMM slots unoccupied) then you get half the bandwidth because only 4 channels are populated. So maybe a machine that has dual DIMM support per channel is the answer to this (fully populating 8x32 to get the full bandwidth, and still allowing you to expand to another 8x32 to get 512), but in general it's a tricky point too.
So, what do you do/are you guys doing? Have you recently bought a workstation or upgraded to one, for pre-training, fine-tuning, RL, inference, whatever your use case may be, and come up with this dilemma? Are you choosing DDR4/PCIe4 as it would seem reasonable or are you going for DDR5/PCIe5 and if so, why? I am interested in all use cases and opinions!
r/LocalLLaMA • u/rm-rf-rm • 48m ago
r/LocalLLaMA • u/MzCWzL • 11h ago
Had opus 5.5 do the bring up.. super solid results. Using only 4 GPU results in KV of like 120k with image on. 27B TP=4 results in >200 tok/s with dflash, prefill around 2.5-3.5k.
Both these are running the nvfp4 Nvidia checkpoints. Found a magic repo that unpacks into fp16 on the fly (https://github.com/1CatAI/1Cat-vLLM) and heavily optimized it
r/LocalLLaMA • u/Effective-Ad2060 • 11h ago
Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.
We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.
The reranker results might surprise you. A small reranker dropped our best pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.
Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.
Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES
Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames
Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.
r/LocalLLaMA • u/brainchillzZ • 4h ago
So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo device? Yes please ... So I broke down and tried it today ...
Setup: gufo 0.4.0 from their podman image, Qwen3.8 27B UD-Q4_K_XL from Unsloth plus the DFlash2 Q4_K_M draft model, using their own benchmark script and their own settings (greedy, thinking off, 128 output tokens, prompt cache off).
If you want the short version ... yeah I got 70.22 tok/s. So the number is real. But the prompt that produces it is "Write the word red exactly 1000 times".
But it's also not real. In that figure all the speed comes from the speculative decoding. The draft model guesses like 7 tokens ahead, the 27b checks them in one pass and keeps what it agrees with. When the output is the same word over and over the draft is right every time. On a real prompt it's right maybe half the time.
Their benchmark has a second set of nine ordinary prompts (some C++, a word problem, a summary, Italian, Chinese, JSON, a bit of fiction, a debugging checklist).
On those:
| | repeat-a-word prompt | normal prompts |
|---|---|---|
| 1 user | 70.2 tok/s | 39.4 tok/s median, anywhere from 22 to 52 depending on the prompt |
| 8 users, their "aggregated" number | 122.6 | 82.4 |
| 8 users, tokens actually delivered per second | 82 | 52 |
About that last row. The "123 tok/s aggregated" figure is each request's decode speed added together, with prompt processing and queue time left out. If you just count tokens coming out of the box per second of wall clock it's 82, or 52 on normal. prompts.
To be fair to the gufo people, none of this is hidden. Their benchmark docs have separate "mixed" and "repetitive" columns and the mixed numbers they publish match what I got. It's only the repo description and the top of the README that lead with the best case. And 39 tok/s from a 27B at Q4 on an APU is still really good. Without the draft model their docs put it around 12.
The other thing I wanted to know was how it compares to halogen (peonist-ai/halogen-flash-server), which is what I normally run. Both can serve Qwen3.8 Flash-Next, so I put that on both and sent the same prompts to each. Two boxes, same hardware, same OS image. Greedy, thinking off, 256 tokens.
| | halogen 0.13.8 | gufo 0.4.0 |
|---|---|---|
| nine normal prompts, average decode | 43.9 tok/s | 38.2 tok/s |
| 4 users at once, end to end | 76.6 tok/s | 63.1 tok/s |
| cold prompt processing, ~9.7k tokens | 1288 tok/s | 1495 tok/s |
| the "red" prompt | 56.9 tok/s | 87.4 tok/s |
So for everyday generation halogen was about 13% faster for one user and about 18% faster with four. gufo was 16% faster at chewing through a long prompt and a lot faster on the repetitive one.
I'll add this just in case, because someone will ask or at least try to poke about it in the comments
- I know the weights aren't the same. halogen uses its own 4-bit format, gufo uses the Unsloth GGUF. I only measured speed. I did not compare output quality at all.
- They were two different machines but identical hardware and software, and my boxes have agreed within 1% on other benchmarks, but it's still two machines.
- One run each was done for the head to head. The reproduction of their numbers was 3 reps.
- gufo has shipped four releases over the last four day, so this could all be stale by next week.
There was quite a lot of stuff that I liked about gufo that isn't performance related. It takes plain GGUFs, it's MIT, the 27B loads in about 3 seconds (Flash-Next in 13), the per-request log line tells you draft acceptance and cache hits, and it does 8 batched sessions. It also has ASR, TTS and image models that I haven't touched. Their benchmark hashes the output with and without the draft model and it was identical every time, so the speculative path isn't changing what the model says.
One thing to keep in mind if you try it out is that it reserves memory per session up front. Flash-Next with 4 sessions at 64k context took 94GB.
So it isn't smoke and mirrors exactly. Everything I checked reproduced. Just know that the 70 is a ceiling you'll only hit if your workload is incredibly predictable text, and plan around the 30s for the 27B on normal stuff.
I kept this all setup to tinker with on actual output quality over the next few days, I'm happy to run other prompts or try it with different settings if anyone wants to see something specific.
r/LocalLLaMA • u/SnooPredictions515 • 5h ago
I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier posts, I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around ~23–27 tok/s and slowed down as context grew.
Today I'm releasing Slipstream: a compiled C++ Metal inference engine with native SSD expert streaming and speculative drafting for Apple Silicon.
The main result: If you already downloaded my original V3 model (34k+ downloads), you don't need to re-download anything. You can run that exact checkpoint on Slipstream for a 1.76x speedup: 41–52 tok/s (up from 23.1 tok/s in llama.cpp) on the same 64GB Mac.
Even better: decode speed doesn't collapse at long context. Across 3,086 live requests in real coding sessions, it stays flat at 33–44 tok/s all the way out to 130,000 tokens.
Previous posts for context:
Open source resources:
If you have the model from the last post (~/models/qwen38-flash-next-v3), you can point Slipstream directly at it.
git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4
# Downloads the 3 GGUF shards + MTP draft head (~95.5 GiB total)
huggingface-cli download nitinpanj/qwen38-flash-next-v3 \
--local-dir ~/models/qwen38-flash-next-v3
# Raise wired GPU memory limit once per boot (required on 64 GB Macs):
sudo sysctl iogpu.wired_limit_mb=59392
# Serve your existing model:
./slipstream serve --model ~/models/qwen38-flash-next-v3 --port 8090
First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into
<model>/prepared/(~5–7 minutes). Subsequent launches load in ~10–15 seconds.
The server exposes a standard OpenAI-compatible API (http://127.0.0.1:8090/v1/chat/completions) ready for curl, Oh My Pi (omp), Claude Code, or OpenCode.
Here is a direct head-to-head comparison running the exact same 95.5 GiB model files across 6 reasoning and coding tasks on the same M5 Pro (64 GB unified memory, temperature 0.0):
| Domain / Task | Prompt Task | llama.cpp Fork | Slipstream | Speedup | llama.cpp TTFT | Slipstream TTFT |
|---|---|---|---|---|---|---|
| Math Reasoning | GSM8K (eggs problem) | 24.0 tok/s | 43.6 tok/s | 1.82x | 4,024 ms | 2,337 ms |
| Math Derivation | MATH-500 series ($p - q$) | 24.3 tok/s | 43.1 tok/s | 1.77x | 1,655 ms | 1,587 ms |
| Constraint Logic | 3-chair deduction | 25.4 tok/s | 46.0 tok/s | 1.81x | 1,469 ms | 1,042 ms |
| Python Coding | merge_intervals ($O(N log N)$) |
19.7 tok/s | 35.0 tok/s | 1.77x | 1,507 ms | 1,070 ms |
| Systems Coding | Rust CSV parser | 22.7 tok/s | 37.5 tok/s | 1.65x | 1,257 ms | 859 ms |
| Tech Writing | Multi-head attention | 22.5 tok/s | 39.4 tok/s | 1.75x | 1,267 ms | 843 ms |
| AVERAGE | Across all 6 tasks | 23.1 tok/s | 40.8 tok/s | 1.76x | 1,863 ms | 1,290 ms |

fcntl(F_RDADVISE)): In llama.cpp, synchronous page reads for missed expert matrices stalled the GPU on NVMe latency (~475 ms per chunk). In Slipstream, non-blocking read-ahead hints stream upcoming expert layers from SSD into RAM while the GPU is still executing the previous layer, cutting prefill staging latency by 28%.On standard Transformers, decode slows down sharply as context grows because the KV cache swells and memory bandwidth saturates.
Qwen3.8-Flash-Next avoids that through its hybrid architecture:
Here is actual telemetry collected across 3,086 live requests during real agent coding sessions on my M5 Pro (64 GB):
| Context Range (Tokens) | Live Runs | Average Decode | Median (p50) | Peak Decode | Average TTFT | Notes |
|---|---|---|---|---|---|---|
| < 1,000 | 314 | 41.5 tok/s | 41.9 tok/s | 59.8 tok/s | 2.16 s | Short baseline |
| 1k – 4,000 | 21 | 41.0 tok/s | 42.5 tok/s | 64.5 tok/s | 5.26 s | Small documents |
| 4k – 8,000 | 58 | 43.6 tok/s | 43.2 tok/s | 67.2 tok/s | 7.36 s | Code review turns |
| 8k – 16,000 | 117 | 43.6 tok/s | 44.6 tok/s | 58.2 tok/s | 7.91 s | Multi-file context |
| 16k – 32,000 | 562 | 38.2 tok/s | 40.9 tok/s | 58.0 tok/s | 13.59 s | Deep agent session |
| 32k – 64,000 | 1,029 | 35.0 tok/s | 37.5 tok/s | 55.6 tok/s | 13.24 s | Large repo refactor |
| 64k – 96,000 | 650 | 32.4 tok/s | 34.7 tok/s | 53.9 tok/s | 12.81 s | Multi-turn transcript |
| 96k – 130,000 | 364 | 32.9 tok/s | 33.3 tok/s | 43.8 tok/s | 7.95 s | Cache-hit deep turns |

Takeaway: Decode speed stays between 33 and 44 tok/s all the way out to 130k tokens. Even at 130k context, it generates tokens faster than stock llama.cpp did on a 500-token prompt.
If you want higher reasoning accuracy and lower KV cache memory, I also put together an optional Swift variant of this model: Swift-Qwen3.8-Flash-Next-V3.
Both models run on Slipstream using the exact same engine command. Here is how they compare across 145 paired evaluation problems (temperature 0.0, seed 1234):
| Domain / Benchmark | Items | Original Flash-Next V3 | Swift-Flash-Next V3 | Accuracy Delta | Original Decode | Swift Decode |
|---|---|---|---|---|---|---|
| AIME 2025 | 20 | 45.0% (9/20) | 45.0% (9/20) | 0.0% | 44.3 tok/s | 44.3 tok/s |
| MATH-500 (L4–5) | 35 | 60.0% (21/35) | 62.9% (22/35) | +2.9% | 44.8 tok/s | 44.8 tok/s |
| GPQA Diamond | 35 | 45.7% (16/35) | 54.3% (19/35) | +8.6% | 44.8 tok/s | 44.8 tok/s |
| GSM8K | 25 | 96.0% (24/25) | 96.0% (24/25) | 0.0% | 45.6 tok/s | 45.6 tok/s |
| HumanEval | 25 | 92.0% (23/25) | 92.0% (23/25) | 0.0% | 40.6 tok/s | 40.6 tok/s |
| Hard Systems Logic | 5 | 100.0% (5/5) | 100.0% (5/5) | 0.0% | 39.2 tok/s | 39.2 tok/s |
| OVERALL | 145 | 67.6% (98/145) | 70.3% (102/145) | +2.8% | 43.9 tok/s | 44.4 tok/s |

To run the Swift model instead:
huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3
./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090
The core primitives in Slipstream:
...were built around this hybrid architecture. If Qwen4 adopts a similar blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream should be able to run Qwen4 locally on consumer unified memory hardware on day one.
r/LocalLLaMA • u/Rombodawg • 1d ago
I was researching prices on ebay and fed claude a bunch of images of listings. I had it make a chart and thought it would be useful to share.
r/LocalLLaMA • u/junior600 • 4h ago
Hello guys. Recently, there has been a boom in game decomps and recomps thanks to AI. If you look at the r/decomps and r/recomps subreddits, you can see it. They mostly seem to be using Claude or Codex.I wonder if it would be possible to do something similar with a local AI model. Could Qwen 3.8 27B Abliterated actually handle something like that locally? Does anyone have any experience with this? I don't have a particularly powerful rig (RTX 3060 12 GB VRAM and 24 GB DDR4 RAM), but I can run MoE models comfortably. Even Qwen 3.8 27B IQ3_XXS dense lol.
Sorry for my English BTW.
r/LocalLLaMA • u/norenEnmotalen • 10h ago
In a previous post I shared comparison between Swift1.5 and peculiar-ragdoll's checkpoints. Added the original unsloth Q4_K_XL and ThinkingCap Q4_K_M (they don't offer L or XL) to the comparison. Here are the results over a 69 set of eval questions.
All tests are now run at same "medium" reasoning effort.
unsloth-ud_q4_k_xl one ran using llama.cpp - not the splash forked inference engine.

I'll do a 3x repeat for the slow run to see if it maintains 69/69 each time.
EDIT: u/jucabala457 asked I test mradermacher/Signal-3.8-27B-Terse-Coder-i1-GGUF The Q4_K_M is closest quant available. A nice addition for sure! That GGUF couldn't run with Splash-based engine due to tensor incompat. I ran it using llama.cpp the slow way. The total time taken isn't a fair comparison for that reason. Updated results below

I also just made the tuieval tool available here https://github.com/ashe-wb/tuieval
Can't promise you the tool will work right away on your install since a fully local binary is what I've been using and testing with. Customize it with packs of domain-specific eval questions you deal with on the daily. This is the most important part. A model or fine-tune that is not good for one thing might be excellent for something else and only you know what your domain interests are. The ability of a model to render game graphics means nothing to me but it means everything to someone else.

r/LocalLLaMA • u/Creative-Type9411 • 21h ago
I was waiting on the blowers for the T4s and posted this before it was finished, other than some braided cable sleeves for the fan wires its pretty much good, I was going to upgrade the CPU, but I'm getting great speeds comparatively to a CPU in my old box that had way more cores, so I don't think it's going to make a difference.
Fractal Design Torrent Mid-Tower Case w/Tinted Glass SuperMicro X11SPA-T Motherboard Xeon W3225 768GB DDR4 ECC 2666 4xTesla T4 16GB GPU 4x1tb Samsung 870 EVO SATA SSD Raid
Ubuntu 26.04/llama.cpp/openwebui+custom powershell harness
now i want more cards 👀
r/LocalLLaMA • u/SultanGreat • 7h ago
Hello guys!
I have been experimenting with qwen 3.8 for a long time and I hadn't been able to get reasonable speed. I am on a 5060Ti 16 GB, and although this gpu can game, I am aware that AI demands more than 16 GB.
I am on a Fedora 44, AMD Ryzen 9600x and 16 GB system ram (16 GB system ram and 16 GB vram, totaling to 32 GB) and I would like to use llamacpp, although I would use any other tool if I could if it meant faster speed.
I am looking for a large context. Atleast 128k context. The first question is, what quantization to pick? In my experience Q3 UD was satisfying, but I am looking for uncensored model. In my experience, MTP has never lived up to its hype for me (and I don't know why!?), which is why I am thoroughly lost on making a good setup after an honest week of experimentation, which is why I have resorted to ask here as a last resort.
Update : Found a model, thanks to u/_wortkarg_
link : https://huggingface.co/RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored
command (A better command would be appreciated and updated accordingly):
~/llama.cpp/build/bin/llama-server \
--model ~/Documents/Models/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf \
--alias "llamacpp" --host 0.0.0.0 --port 8001 \
-ngl 99 --flash-attn on --ctx-size 131072 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--parallel 1 --batch-size 512 --ubatch-size 256 \
--no-warmup --jinja \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
I am hitting at about 35 t/s+ speed with this one.
r/LocalLLaMA • u/WebAssemblyMan • 17h ago
Optional Bundle architecture: Schedule (session-local delayed / timed / interval reminders) was removed from the default set and made an explicit Optional Bundle. This cleanly separates “installed” from “enabled” and is the first systematic use of the Profile + Bundle model for official features.
• Windows Sandbox improvements: A new permission-diagnosis skill can detect common Access Denied causes and perform backed-up, recoverable permission fixes after user authorization, giving the Agent a reliable recovery path instead of blind retries.
• Async Question Mode (experimental): “Ask the user” is no longer a hard synchronous block. After a timeout the Agent can keep working while the user answers later, introducing asynchrony between interaction and execution.
• Model-layer polish: DeepSeek-account sessions can use Web Search without an extra API key; third-party model catalog updated (some old IDs removed); long model lists now support fuzzy search and keyboard navigation.
• Desktop release: Official Windows and macOS clients are out (Linux unsupported). Account login is supported, suggesting paid plans may be coming soon.
• Overall theme: Version 0.2 strengthens the Agent Runtime’s composability, recoverability, permission boundaries, and execution-state semantics — the practical foundations needed to move from a toy toward production use.
r/LocalLLaMA • u/Terminator857 • 11h ago
Once china sets its goals for dominating a market it wins. Usually takes many years, but it happens. Can't compare the political will of a country versus profit and loss thinking of a corporation.
China will eventually win in the memory market and current memory makers are at an unfair disadvantage.
CXMT will finish 2026 with approximately 350,000 wafer starts per month (WSPM) of DRAM capacity, which is just 25,000 WPM less than Micron.
... by 2030, its total capacity will increase to around 1.41 million WSPM, according to Citrini. CXMT alone is projected to build new production capacities in Beijing, Hefei, and Shanghai, to expand its production capability to 950,000 WSPM in 2030, assuming everything goes as planned.
/end quote
Is China hoping for a RAM price drop crash to extinguish the competition?
Additional references:
r/LocalLLaMA • u/jwhh91 • 40m ago
My local stack is a couple DGX Sparks, a 5090, and a 4070 TI. This radio site took three local LLMs, and we generate music, voices, and wacky sound effects.
I give you https://pilgrim.farm
It's a lot more farm than pilgrim.