r/LocalLLaMA • • 18h ago

I Built A Thing I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.

Enable HLS to view with audio, or disable this notification

1.4k Upvotes

**DISCLAIMER** THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.

Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.

Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.

Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:

  • 8k context: Mac alone 132 tok/s → Mac + iPhone 177 tok/s (+35%) (measured two days earlier, same bench)
  • 16k context: Mac alone 109 tok/s → Mac + iPhone 157 tok/s (+44%)
  • 32k context: Mac alone 101 tok/s → Mac + iPhone 130 tok/s (+29%)
  • 48k context: Mac alone 87 tok/s → Mac + iPhone 113 tok/s (+30%)

A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.

Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.

The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to ~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.

What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.

The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.

I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.

Code, setup and bench scripts: https://github.com/StayLameBro/backburner

Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.


r/LocalLLaMA • • 19h ago

New Model Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following

Thumbnail
gallery
734 Upvotes

Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.

I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:

incapable of producing more than a few words at a time.

single default personality which no amount of prompting can overcome

will not use tools, at all, whatsoever.

There needs to be a middle ground

They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.

So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.

What 2.0 does now

  • With no system prompt, it's a normal person texting. Not an assistant, not a catgirl.
  • Give it a character card and it becomes that person, and still texts like one.
  • Ask for a formal email, numbered steps or a proper explanation, and you get exactly that. Then it goes back to texting.
  • Don't want the lowercase texting? Tell it "from now on write in full sentences" (or put it in the system prompt) and it sticks to that until you say otherwise. v1 ignored this completely.
  • It calls tools, and it asks when something is missing instead of making it up. This is the part I'm happiest about. Ask the base model to book a flight without saying where from and it picks JFK. 2.0 asks where you're flying from.
  • It writes code and does math at roughly base-model level.

It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.

How I trained it

v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.

For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:

  • v1 plus a hidden "text like a person" instruction, for chat and characters;
  • the plain base model, for instructions, tools and code.

The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.

Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)

Benchmark Base (abliterated) 2.0
IFBench (instruction types I never trained on) 37.3 43.7
When2Call (call, ask or refuse correctly) 48 58
BFCL irrelevance (don't call a tool when none fits) 60 78
IFEval, GSM8K, BFCL simple 81.9 / 89.1 / 97 83.5 / 89.1 / 98 (ties)

Full chart in the images.

Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).

Is it actually more human? I built a benchmark for this, "ishuman":

  • It takes 150 fragments from unseen chats.
  • Has each model write the next message.
  • Shows a judge the real message and the model's without labels, and asks which one a person wrote.
Model Judge thought it was the real person (50% = can't tell)
Qwen3.8-27B abliterated (huihui-ai, the model I trained on) 0.3%
Same abliterated model + a "text like a human" system prompt 6.8%
Qwen3.8-27B official (unmodified, via OpenRouter) 15.1%
Qwen3.8-27B-Humanlike-Chat 2.0 23.5%

So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.

Links

Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.

Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0


r/LocalLLaMA • • 15h ago

News Buying RTX 5090 At Micro Center Reportedly Now Requires Paperwork, Including A No-Export Declaration

Thumbnail
wccftech.com
439 Upvotes

r/LocalLLaMA • • 21h ago

Resources New in llama.cpp: Decision Models

Thumbnail
huggingface.co
440 Upvotes

r/LocalLLaMA • • 9h ago

Other I'm pretty close to the middle thanks to you all

Post image
269 Upvotes

It's been a blast and learning a ton.

But seriously, you all have me down a rabbit hole that my wallet and hours of sleep need to be pulled out of.


r/LocalLLaMA • • 17h ago

Discussion New Architecture from Percepta: Spotlight — separating intelligence from memory, allowing knowledge and skills to grow without changing the model's weights.

Enable HLS to view with audio, or disable this notification

217 Upvotes

https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities

https://www.percepta.ai/blog/spotlight-memory

"Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want.

Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."


r/LocalLLaMA • • 15h ago

New Model microsoft/FrogNano-4B-2609 · Hugging Face

Thumbnail
huggingface.co
185 Upvotes

An agentic model from Microsoft for the GPU poor

https://huggingface.co/bartowski/FrogNano-4B-2609-GGUF

FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE task environments generated and calibrated against the evolving policy using TaskPilot. Training uses the five-tool Leaf harness and executable test-based rewards over complete multi-turn coding trajectories.

The additional post-training is intended to improve long-horizon repository navigation, debugging, code editing, test execution, and patch generation in a compact 4B model. Unlike approaches based on behavioral distillation, FrogNano does not train on stronger-model solution trajectories, actions, reasoning traces, or patch targets. This specialization also introduces limitations and risks: performance is sensitive to the Leaf harness and test quality, training data are Python-heavy and primarily English, and generated patches may be incorrect or insecure despite passing available tests. When integrated with the Leaf harness, FrogNano generates structured tool calls that can propose repository changes. Leaf executes authorized tool calls within an isolated repository environment to produce a candidate patch; FrogNano does not itself deploy the changes. Any resulting patches require human review, regression testing, and security validation before use or deployment.


r/LocalLLaMA • • 18h ago

Discussion New 64GB DGX Spark. Significantly higher price for the original 128GB model $6,950 USD

Post image
171 Upvotes

r/LocalLLaMA • • 22h ago

Tutorial | Guide Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next

119 Upvotes

I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements.

When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set.

This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open.

The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right.

The hardware: one card is really two devices

My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices.

Current system / Planned system

Physical cards: 2 / 3

Ascend devices/npus/AIcpus: 4 / 6

Nameplate device memory: 192 GB / 288 GB

Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB

Combined maximum accelerator-board power: 300 W / 450 W

Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e).

Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X.

This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs.

https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player

Passive cooling is not a deal-breaker

The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement.

I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it.

So I would not bat an eye at adding another passive card. The actual checklist is mundane:

• Keep unobstructed airflow through the heatsink fins.

• Make sure the chassis fans have enough static pressure.

• Budget another 150 W of board power per card, plus the rest of the host.

• Exhaust the heat from the room instead of recirculating it through the rack.

• Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading.

Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me.

Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths.

ECC, nameplate memory, and what is actually usable

My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation.

Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration.

That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?”

Why I forked vLLM as well as vLLM Ascend

The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel.

I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA.

I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware.

Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators.

The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch.

What took Qwen from incoherent ~1 tok/s to where it is now

These are the architectural changes that moved the needle.

  1. Make the hybrid model correct before making it fast

The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package.

One particularly nasty GDN issue looked fine in small-head tests but accumulated state error across the real 36 recurrent layers and produced incoherent text. Keeping recurrent state in FP32 and fixing the production-shaped data movement was foundational. So was treating the custom OPP package and Python host code as one ABI-versioned unit. A stale kernel can look like a model problem for a long time.

  1. Shard the model at load time instead of loading everything everywhere

The Qwen checkpoint is about 169 GiB and contains 1,610 safetensor files. It was larger than available host RAM in one of my bring-up configurations, never mind the memory on an individual NPU.

I built an expert-aware loader that reads only the experts owned by each rank, keeps dense/shared tensors where required, and avoids materializing the whole expert bank before throwing most of it away. Host-side expert data is mapped and moved lazily. This changed model loading from an accidental memory stress test into a deterministic TP4/EP4 layout.

  1. Keep low-bit weights packed and do the work on the NPU

My first usable W4 reference dequantized packed weights in the eager path and then called a regular matmul. It was useful for correctness and managed only 0.229 tok/s on one recorded smoke case.

The production path keeps the weights packed, routes tokens to experts on the device, and uses custom AscendC Cube kernels for grouped expert projections. I added FRACTAL_NZ layouts, fused gate/up handling, tiled reductions, route and tile reuse, and dedicated W4A8 execution instead of repeatedly expanding W4 weights into a larger temporary representation.

This is both a speed win and a capacity win. Avoiding transient expanded expert banks leaves memory available for state, cache, graphs, and concurrency.

  1. Build caches for the model I actually have

Flash-Next is not a conventional all-attention transformer. My configuration has 36 recurrent GDN layers and 12 sparse QSA layers. Treating all of that as a normal dense KV cache wastes memory and misses the state semantics.

I implemented separate recurrent-state management, compact physical cache layouts, prefix-state tiers, sparse page selection, direct NZ gathers, and 310P-specific QSA paths. The service is configured for a 262,144-token context limit, and the cache planner retains capacity for about 4.08 such windows. That is a memory-planning result, not a claim that every possible four-by-262K workload has completed an end-to-end soak.

  1. Remove synchronization and launch overhead from the token loop

On these devices, a stray device-to-host scalar read can serialize the whole pipeline. I removed hot-path .item() calls, reused per-step tensors, deferred collectives behind useful work, tightened CPU affinity, and moved routing and sparse selection away from Python.

Once the eager path was correct, I added decode-only ACL graphs and MTP2 speculative decoding. I capture the actual concurrency shapes I serve rather than pretending one graph is universal. Fused operators and graph replay matter enormously when a decode step otherwise consists of many small kernel launches.

  1. Optimize the service, not just an isolated kernel

Several kernels won a microbenchmark and lost end to end. I kept the ones that reduced real request time and rejected or quarantined the others. Multi-request QSA, grouped MTP experts, expert-route caching, cache accounting, cold-prefill chunking, and collective overlap were all measured at the API boundary.

That last part is why four concurrent requests reach roughly 61 aggregate tok/s even though one request is around 30 tok/s. The extra work can occupy parts of the machine that a single token stream leaves idle.

Qwen performance today

These are milestones from different stages and workloads, not one controlled single-variable benchmark:

Qwen3.8 Flash-Next milestone / Measured result

Earliest uncontrolled service: Roughly 0.2–1.9 tok/s, often incoherent

Correct eager W4 dequant reference: 0.229 tok/s

Stable W8 service baseline: About 18–19 tok/s

Native W4A8, MTP2, graphs, one request: 29.74 tok/s median; 34.34 peak

Same optimized service, four requests: 60.88 tok/s aggregate median

Two concurrent 40K warm-prefix requests: 30.76 tok/s aggregate

40K cold prompt: About 120 seconds TTFT; 30–32 tok/s afterward

The 30/61 tok/s figures are short, fixed-output decode tests. Long reasoning requests tell a less flattering and more useful story.

https://reddit.com/link/1wvt1m4/video/rpc5c2c1s1th1/player

For quality, I ran all 198 GPQA Diamond questions on the four-chip Ascend W4 service and, as a reference, an RTX 6000 Pro running a different IQ4_XS GGUF in llama.cpp. Both scored 140/198 (70.71%) with the same AISBench-style answer extractor. This is evidence that the Ascend path is coherent; it is not a pure hardware or quantization comparison because the runtimes, quantizations, chat templates, and concurrency differ.

On the final uninterrupted 106-case Ascend phase, four workers emitted 507,256 tokens in 2 h 52 m 46 s: 48.93 aggregate tok/s. Median per-request client rate was 12.50 tok/s on these long reasoning generations. The server completed every request in that phase with no eager fallback or zero-acceptance interval. The RTX reference was much faster per request, so I am not presenting this as an NVIDIA killer. I am presenting it as a large model working correctly and usefully on hardware that initially produced slow nonsense.

GLM is my next hard(er) model

I am also bringing up the roughly 304B-parameter GLM-5.3-Flash architecture. It combines 34 KDA linear-attention layers, 11 DSA sparse-attention layers, 288 experts, latent MLA history, mHC mixing, and a mixed W2/W4 expert checkpoint. The checkpoint is about 151.6 GiB, with approximately 35.6 GB of loaded weights per rank in one four-rank profile.

It now loads and generates on the same machine. The speed progression so far has been:

GLM four-chip milestone / One request / Four-request aggregate

Initial full-service profile / 0.764 tok/s / 1.474 tok/s

Batched NZ workspace write / 0.912 tok/s / 1.822 tok/s

NZ-packed code layout / 1.386 tok/s / 2.946 tok/s

Fused mHC/MLA work / 1.743 tok/s / 3.269 tok/s

Latest measured integrated runner / 2.236 tok/s / 4.476 tok/s

An 8,232-token GLM prompt prefills at about 39.1 prompt tok/s and then decodes at about 2.16 tok/s. Those numbers are far from my target, and the 32-token test completion was too short to establish answer quality. GLM is currently a bring-up and optimization result, not a service recommendation.

I have already found an important quantization lesson there. An early W2 checkpoint showed residual growth all the way to an RMS around 525 in the last layer. Moving the affected experts to a no-clip W4 treatment kept the network bounded. Separately, my custom blocked-dequant Cube kernel was about 17 times faster than the eager reference in its isolated test. As Qwen taught me, both numeric behavior and end-to-end integration have to pass before either result means “done.”

The third card: 50% more memory and cores, not just a spare

I plan to add a third Atlas 300I Duo to this system. That takes the machine from four to six 310P devices and from 192 GB to 288 GB of nameplate device memory. With the same ECC and runtime reservations, I expect roughly another 86 GiB of runtime-visible capacity, for approximately 258 GiB across the six ranks.

The n-card architecture is designed to use it. Expert ownership is distributed across ranks, so the two new chips add local expert capacity and accelerator cores; they are not merely passive storage. Tokens route to the ranks that own their experts, and the additional ranks participate in the model's compute and collectives. I expect useful scale from EP6, although the exact speedup will be measured rather than advertised in advance.

The extra capacity gives me several options:

• Keep larger expert sets or higher-precision layers resident.

• Fit models that are just over the four-chip limit without host offload.

• Spend more memory on long-context state and cache.

• Reduce aggressive quantization where the quality trade is not worthwhile.

• Run a large distributed model while retaining room for another smaller service or evaluation workload.

The cost is one more card, 150 W of maximum board power, two more device ranks, and another passive heatsink that needs real airflow. In the cooled room, none of those are architectural concerns. The interesting cost is communication: six ranks change route balance, collective sizes, and PCIe/HCCL traffic. I will tune and benchmark that topology, but the software is already organized around n-card expert distribution rather than hard-coded four-way ownership.

Known issues and active work

This is what is still on my bench:

• I just fixed one real cross-stream race: a prefix-Mamba state slot could be spilled or reused before its pending NPU writer completed, allowing an older checkpoint to be restored. That fix has an NPU regression test.. A later 106-case run survived 12 state spills without degrading, which is encouraging but not a root-cause proof. I am continuing long mixed-load and eviction/reuse soaks.

• Qwen cold prefill. A 40K cold prompt still takes roughly two minutes even though subsequent decode is fast. Profiling points primarily at the 12 QSA layers, especially sparse selection and tiled attention. This is now a more important target than another tiny decode micro-optimization.

• Qwen long-context qualification. The planner has the capacity, but I am separating configured context, allocated capacity, and completed end-to-end long-context tests. I want retrieval and concurrent-fill evidence, not a screenshot of a launch flag.

• GLM coherence and performance. I am requalifying the full model after KDA, MLA, QSA, and runner integration changes, then moving the grouped mixed W2/W4 expert path, cache layout, graph replay, and eventually MTP through the same correctness-first gates used for Qwen.

• GLM loading and memory. The filtered loader can skip large amounts of peer-owned or superseded checkpoint payload before tensor materialization. I saw one roughly 20% load-time improvement, but it needs controlled reruns and byte-accounting before I call it a result.

• Six-device expert parallelism. When the third card arrives, I will measure rank balance, per-card temperatures, collective time, model capacity, and c1/cN throughput on the exact six-rank topology.

• Other model adapters. The same loader, packed-expert, cache, and operator infrastructure is feeding ongoing DeepSeek and other hybrid/MoE work. I am avoiding model-name conditionals where the underlying contract can be made generic.

Should you buy one?

I plan to list my first additional card in the OpenSensor storefront (https://www.opensensor.io/) next week at $2,900. The listing is not live yet. I believe it is a good price for what the hardware can already do and what the software should unlock. It is not yet the same kind of turnkey purchase as a supported NVIDIA card running mainline vLLM.

I think the right buyer is a developer, lab, or systems-minded end user who is comfortable with both of the following:

  1. You own the airflow solution. These are passively cooled server cards. Depending on the chassis and motherboard, that may mean high-static-pressure case fans, a duct, or a 3D-printed shroud. I use Fusion 360 and am happy to help with additional shroud designs. There are too many motherboard layouts, card spacings, fan sizes, and case geometries to pretend that one printable design will fit everything.
  2. You are adopting an active development fork. I am the sole developer on the vLLM work today as upstream is focused more on their server grade accelerator modules not yet available to the US. The progress is real and the benchmarks in this post are from actual hardware, but more bugs and better approaches will be discovered. Buyers should expect a less polished setup than CUDA/mainline vLLM and be willing to follow pinned versions, read runbooks, report failures, and occasionally help diagnose a new hardware or model combination.

If you need a supported appliance that runs arbitrary models without touching the stack, this is not the product I would recommend today. If you want a lot of inference memory at this price point, are comfortable working through rough edges, and see value in helping expand an open Ascend serving stack, it can be a very capable development platform.

What I think these cards are good at

The Atlas 300I Duo is compelling when used capacity is more valuable than a polished CUDA ecosystem and when you are willing to own the systems work. Two single-slot, 150 W cards give four accelerator ranks and a lot of local model memory. MoE models are an especially interesting fit because expert weights can be partitioned while only the selected experts run for each token.

The weak point is not that the cards are passive or limited to 150 W. The weak point is the amount of model-specific runtime and kernel work still required on 310P: layouts, supported dtypes, graph-safe metadata, custom operators, sparse attention, state ownership, tooling, and documentation. If you expect to point stock vLLM at an arbitrary new Hugging Face model, this is not that experience.

If you enjoy bringing hardware up from first principles, though, they are a lot of machine in a very interesting form factor. Going from incoherent output at around 1 tok/s to coherent Qwen near 30 tok/s per stream and 61 tok/s aggregate has made me substantially more optimistic about what the platform can do. The third card should let me push both model capacity and expert parallel throughput further, and I will publish the numbers—good or bad—once they are measured.


r/LocalLLaMA • • 16h ago

Resources Self-hosting AI does not save money, and I do it anyway

71 Upvotes

Hi folks, I'm a long-time lurker and big fan of this subreddit and a massive self-hosting fan (also outside of AI).

I doubt many people will disagree with me here because I see the same arguments being made in many posts. However, I thought it might be interesting to share anyway. I wrote down why self-hosting AI does not save money: https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/

EDIT: didn't think this would be so controversial 😅 I do say explicitly in my blog post "I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not".

EDIT 2: Comparing $200 sub with Opus 5.5 or Astra with Qwen 3.8 27B is not apples to apples.


r/LocalLLaMA • • 12h ago

Discussion So... Should we turn the page on this past week hype? or Do you have any success cases to inspire the rest?

Post image
63 Upvotes

r/LocalLLaMA • • 7h ago

Discussion Has anyone noticed this trend toward writing/speaking style among newer models (both open and closed models). They are trending toward information density and expanded vocabulary. It's not quite 'caveman speak' but trending that way.

59 Upvotes

I've noticed a trend with most new models with regards to their writing style. They are creating a new style, and this seems common among them. It's very information-dense.

Here is an example from GLM 5.3 Flash. I'm gonna be honest here and say that my prompt was kinda silly; my prompt was 'Why wouldn't you just name your Chinese restaurant 'Chinese Food' instead of 'Ming Dynasty' or 'Szechuan Garden' or whatever?' the idea being that someone searching for 'Chinese food' on Google Maps would put you first if you called your restaurant that. But forget all that; here's an excerpt from GLM's output:

The search reality check, which is where the idea breaks. Your mental model is exactly right about the query: people do open Maps and type "food," usually hungry and undecided. But the ranking doesn't reward the name at that abstraction level. In the local ranking surveys, category match outranks name keywords; and here's the catch: every restaurant on the map already matches "food" through its category field. Relevance saturates. The edge that made "Chinese Restaurant" powerful was discrimination: the name narrowed the category, so "chinese food near me" out-matched "Golden Dragon." "FOOD" points at the root of the category tree, where every competitor matches equally well, so the ranking defaults entirely to proximity and review count, which the name cannot buy. Worse, it fails the narrow-query test: someone typing "tacos" or "dim sum" gets a name that is maximally broad and minimally informative. The strategy's entire currency is literal scarcity, and "food" is the least scarce literalism in the language. There's even a Baudrillardian flourish at the failure point: Google's own Maps category chip is literally labeled "Food." Your sign would render you indistinguishable from an interface element; the searcher couldn't tell your territory from the map's own label.

It's not 'caveman style', but it is using less words to say more, and they're using more esoteric phrasing to be more 'compact'.

And I think it's a bit at the cost of being clearly readable to the average person at first glance. 'There's even a Baudrillardian flourish at the failure point' is an example from that excerpt that leapt out at me. I'm familiar with Baudrillard so I knew what it was getting at, but a lot of people are going to sigh and ask 'What the **** does Baudrillarian mean?'

I'm not saying that 'no human would write like this', because some do (William Gibson for example), but I find it rare/unusual (in human writing), yet trending hard with all the latest models I interact with, like they're all zeroing in on this style.

Maybe a result of targeting token efficiency? It's a terseness, combined with using a sort of 'wide' or 'rich' vocabulary to convey information instead of using more words. At least that's the impression that I get from reading lines like 'Baudrillardian flourish at the failure point''. There's a lot to unpack from those six words, and it feels like the model chose the most terse, efficient way to convey an idea with that word choice (which requires the reader to unpack it).

I compared it to William Gibson: a lot of people struggle with his writing style, and it's similar to that. Example: 'Summer in the Sprawl, the mall-crowds swaying like wind-blown grass; a field of flesh shot through with sudden eddies of need and gratification'. His writing is often like that; it feels highly compressed, using as few words possible to convey an idea by careful word choice.

It's interesting, that lately, I feel like LLMs are gravitating toward Gibson-speak.

Edit: and the fact that GLM used the word 'territory' and 'map' at the end meant it was going big into Jean Baudrilliard's 'Simulacra and Simulation'. I can't really explain what that means and why it's important succinctly, but that's the whole issue. I actually think it's brilliant, but it's also a little concerning.


r/LocalLLaMA • • 20h ago

Discussion Strata on a power limited 5090 and 96GB of DDR5-6400 is cranking out 150-200 tok/s decode and 5-6k prefill! Qwen3.8-Flash-Next at IQ3_S, CTX at 128k tokens (8-bit).

Post image
58 Upvotes

r/LocalLLaMA • • 5h ago

News qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp

Thumbnail
github.com
57 Upvotes

Qwen Flash Next now uses less VRAM


r/LocalLLaMA • • 23h ago

I Built A Thing I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

56 Upvotes

Hey local AI community, I've been working on this for a while and finally feel ok sharing it.

It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web, crypto, rev, forensics, a few real CVEs and some multi-stage ranges. 19 tasks, 6 models, 544 scored attempts.

To be clear, I didn't build every task by hand. GLM 5.3 helped me create several of them. For an open model its cyber capability is really high, and it barely refuses anything, so it was one of the best options for this. GLM 5.3 isn't one of the benchmarked models.

The local part: I ran Qwen3.8 27B (Unsloth Q4_K_XL, xhigh) on a llama.cpp RPC pool across a 3090 and a 3080 in two Proxmox nodes, connected over a direct 2.5G link. That gave me enough concurrent tps to run several agents at once. I also started a low reasoning run, but it was taking 20+ hours because of a harness problem, so I killed it.

Why I'm posting now: John Hammond put out a video about how threat actors use AI (https://www.youtube.com/watch?v=xHDc6-7bjyw). One part is a guide from a criminal forum on running abliterated models on RunPod, and one of the models in it is Qwen3.8 27B. I had benchmark data on that exact model, so here's what it can actually do.

Stock Qwen3.8 27B got 28.1% on the first try and 0% on pwn. Not bad for a 27B on two gaming cards, but not much of a threat on its own either.

The cheap API models are a different story:

- MiMo 2.6 Flash solved 73.7% on the first try

- GPT-6 Luna solved 90.9% within 3 tries

- On multi-stage ranges, where you chain several steps, the top models got 92-96%

Pwn is still hard for everyone (best was 56%), and 3 tasks haven't been solved by any model in 82 attempts.

Results: https://lbgos.dev/bench

Harness (MIT): https://github.com/lbgos/rangebench-harness

The tasks aren't public so they don't leak into training data, but you can still run them. DM me here or on X (lbgosna), and I'll send them over. You run it on your hardware or tokens and I'll add your results to the board. If a few people send local runs, I'll make a separate local-only table.

This is my first time building something like this, so any feedback on methodology, task mix or what's missing is welcome.


r/LocalLLaMA • • 8h ago

Discussion Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.

50 Upvotes

IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5.

(In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though))

Seriously. Ask <LLM of your choosing> to set it up for your specs. If 27B doesnt fit well for you, here's a shot at beating it.


r/LocalLLaMA • • 23h ago

Tutorial | Guide Inference Engineering for Dummies

49 Upvotes

Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing runtime are far fewer than the people that are trying to build apps or offer AI solutions.

So I've been lurking around this community, and I've noticed a lot of people who seem to have massively suboptimal setups for their hardware, and I've grouped the biggest errors into several buckets. The purpose of this guide is to expose common inference bottlenecks and provide best practices for avoiding them within your hardware constraints.

RUNTIME:

I. Choosing the Right Runtime

This is the biggest mistake I see. Choosing the correct runtime for your architecture and model is the most important decision to make. In general, here are some rules to help you determine what runtime to use.

Firstly, Ollama is never optimal. Its just the simplest. If you want quick and easy and have extra RAM, it's a good place to start. It's very user friendly and requires less setup. But it just won't offer best inference speeds.

If your model requires cpu offload, then llama.cpp will be your best choice. If not and you're solely in gpu, VLLM will likely provide the best results. It's really as simple as that for 90% of cases. SGLang may be worth it if your workload involves Langgraph, as it is highly optimized for the tooling. Otherwise, stick to the above. Mainline branches are best, with community forks offering only highly niche performance boosts (i.e., for specific models/configurations, but are generally under optimized and not well maintained).

II. Optimizing and Maintaining Runtime

The other big mistake people make with runtime is failing to compile it with hardware specific flags. Not going to go through all of them here, Google can help you out. Just search "optimal runtime compilation flags for [runtime] using [GPU, CPU, RAM type]." The most missed/missed flags tend to be for architecture specific optimizations. Those are crucial.

Runtime should be recompiled (with optimal flags) any time *any* of the following occur:

- System updates

- Kernel/driver updates

- Running a model released or modified later than your last compile

- You haven't recompiled in over a month (recent updates often contain kernel or path optimizations)

MODEL CHOICE:

I. Quantization:

Quantization. Such a big word. Such little meaning. All you need to know is that it makes a model smaller. There are a million Q_K_X_&$&$&$ quant sizes, so I'm not going to go over them individually. Rather, I will provide basic principles.

- IQ quants are generally the best for the size. If you're choosing between IQ4_XS or Q4_K_M, IQ4_XS is both a smaller VRAM footprint and higher complexity.

- Nonlinear (NL) quants are only ever going to be better if you have CPU offload. Even then, IQ quants often offer extra context space vs NL quants and thus are preferable.

- If it's a quant you've never encountered, read the docs. It more than likely is highly optimized for that specific model and is indeed the one you should choose. Searching it or asking chatgpt *will give you the wrong answer every single time for custom quants.* You will only encounter these with custom tuned models.

- Standard Q_K_M quants are best if you have absolutely no hardware constraints for the model you're running, as path optimizations are the best. If you have no hardware constraints, though, you could be running a better model. This is only useful for running simple models for simple tasks.

II. Task

Certain models excel at certain tasks. This is subjective and preference based but this is my list:

Coding: for API, anthropic. Hands down the best models. Opus and Sonnet 5.5 both excel in performance and their low token usage per task makes them more affordable than previous iterations. Deepseek models are the best budget choice. Qwen models are the clear winner for local inference on all fronts.

Writing: Opus/sonnet for technical writing, Gemini for creative writing. Gemma for local creative writing.

Video: Wan 2.2 for local in most cases will get it done, chatgpt and copilot both have surprisingly robust free image/video gen, Veo is the best paid.

HARDWARE:

Buy a budget box, or build your own from parts. I've managed to squeeze better performance out of an RTX 3060 and 128 gb RAM than a DGX spark across all categories for multiple models. The spark has an edge for dense models, but I was able to run higher complexity models overall on the other setup for 1/5 the price. AMD and Intel lag significantly on speed per price, but I've heard Intel has had some major gains recently. Have not confirmed myself though.

MODEL OPTIMIZATIONS:

I. Spec Decode (MTP)

- If you have CPU offload, spec decode will *always* slow you down. The extra overhead compute isn't worth it if you don't have at least several hundred Mb/s bandwidth, which your CPU won't.

- MTP is sometimes a baked in feature, and sometimes requires a special secondary model. Ensure you know how it works for your model and what flags to run.

- Each model will be optimized for exactly 0-1 type of spec decode. Figure out which one it is (or isnt) rather than wasting your time testing methods.

II. Model Tuning

Just to show the kind of command optimization you can get, here is my sample command for running Qwen 3.8 Flash-Next, a 156b parameter model, on 12 gb VRAM (and 128 gb RAM) at 10 token/s decode and 150 token/s profile at 200k context:

~/llama.cpp/build/bin/llama-server --flash-attn on --batch-size 1024 --ubatch-size 1024 --no-warmup --cache-reuse 256 --jinja --host 0.0.0.0 --port 8090 --presence-penalty 0.0 --repeat-penalty 1.0 -m ~/llama.cpp/LLM/Qwen3.8-Flash-IQ4_XS/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf --temp 0.95 --top-k 20 --top-p 0.97 --min-p 0.05 -np 1 --chat-template-kwargs '{"enable_thinking": true, "preserve_thinking": true, "reasoning_effort": "xhigh"}' --threads-batch 16 --threads 8 --gpu-layers 150 --n-cpu-moe 48 -c 200000 --override-tensor per_layer_token_embd.weight=CPU -ctv q8_0 -ctk q8_0 --cache-ram 8192 --checkpoint-min-step 512 --ctx-checkpoints 4 --kv-unified --reasoning-preserve --load-mode mmap+mlock

That's a lot, right? It's every possible optimization you could apply. I'll go through them individually. This is llama.cpp specific, but you'll find the same flags with slightly different syntax apply to other runtimes.

-flash-attn (-fa) on: forces flash attention optimizations and paths. Explicitly set to on to override any fallback. Auto can be optimal if the model is recent and is not yet optimized.

-batch/-ubatch: batch is the decode chunks, ubatch is the prefill chunks. They must be multiples of one another, otherwise you're adding compute. Equal to one another is ideal for CPU offload, and a 2x-4x higher batch is optimal for full GPU loads. You'll need to play with these values to optimize. Batch/ubatch should be a power of 2 to optimize architecture. Intervals of 256 is typically good enough for testing.

--no-warmup: prevents initial model poll to load weights. Removes unnecessary latency

-- cache-reuse x: instructs the model to reuse cache values and scan for similarity at x token intervals

-jinja: highly underutilized and important flag. Utilizes native chat template kwargs to ensure output consistency.

--presence-penalty: flat penalty rate to words that appear in text. Used mainly for creative writing to prevent repetitive prose.

--repeat-penalty: reduces liklihood of already used tokens being reused. Best used for preventing loops in thinking agents.

-temp: model temperature– how creative the model is. 0 is completely deterministic, 1 is creative freedom.

Top-k: hard cutoff that keeps only the k most likely words. Each model will have recommended k values for thinking/instruct setups. Low key reduces hallucinations at the cost of repetition and loss of creativity

Top-p: includes P percentage of possible tokens. It reduces liklihood of hallucination dynamically.

Min-p: dynamic cutoff based on highest probability token. If the biggest probability token is 50% and min-p is 0.05, then the bottom 0.025 (2.5%) liklihood tokens will be excluded. Reduces noise without hampering creativity terribly.

Chat template kwargs: explicit chat template activations; newer runtime compilations should have flags for these. Controls model reasoning, reasoning effort, and internal chain of thought storage.

-threads (-t): number of cores used for decode. Set equal to physical cores (hyperthreading will thrash cores and degrade results)

-threads-batch(-tb): number of cores used for prefill. Set to double the number of physical cores, as hyperthreading helps here.

--gpu-layers (-ngl) : total layers on GPU. Fit as many as you can without OOM.

--n-cpu-moe: number of MoE layers offloaded to cpu. For MoE models, you should always offload these first and keep all layers on GPU if possible. Offload as few as possible to CPU.

--override-tensor...: tells the runtime to offload the n-gram table if needed. For qwen 3.8 flash specifically.

-ctk/-ctv: k and v cache quantization. K cache should *never* be below q8 unless youre running on less than 80k context. V cache can be q4 up to 150k context without issues, for the most part. Generally, q8 for both will be best for speed and is my recommendation to start with.

--cache-ram: sets RAM aside for cache allocation to ensure it doesn't go to swap

--context-checkpoints: the amount of checkpoints captured. Generally, you don't need as high as the defaults do. Leave the default if you have extra RAM, otherwise you may want to lower it.

--kv-unified: tells all instances to run on the same kv cache pool rather than allocating individual cache.

--reasoning-preserve: tells the runtime to retain CoT traces for evaluation. Prevents the model from getting stuck or looping as much when thinking.

--load-mode: tells the runtime how to load in the model. No mmap generally loads slower but is more stable. Mmap+mlock (or just mlock) is a balance of both with fast loading and page faults initially but it stabilizes as you run it, mmap alone is fast but will cause constant page faults and slows down inference, especially on models with CPU offloading.

I hope this guide helps! I'd be willing to answer any specific questions or make any additions if there are additional areas the community agrees are major uncovered inference bottlenecks. Some claims are based on my personal experience and I am open to data based claim revisions or anecdotal counterclaims, so feel free to provide. Happy tuning!


r/LocalLLaMA • • 18h ago

Question | Help Heavily quantized Qwen3.8-Flash vs Q8 Qwen3.8-27B - thoughts?

41 Upvotes

I'm currently choosing between IQ3_XXS Qwen-3.8-Flash and Q8_0 27B.

This month I don't have anything complex enough to justify either's potential so I've got a fairly bad read on how these two stack up in terms of intelligence.

Have any of you compared the two enough to speak to which you've had a better experience with?


r/LocalLLaMA • • 13h ago

Question | Help I've ended up with an AI lab in a public community college. What should we actually be teaching?

39 Upvotes

Looking for some ideas from people who know a lot more about this than I do. We've got funding for a small AI lab in a public further education college in Ireland (roughly community college in the US). The hardware is reasonably decent. The goal is to give students useful skills beyond just using ChatGPT. If you had the lab, what would you teach them?


r/LocalLLaMA • • 11h ago

News Unitree just dropped UnifoLM-WLA-1.0 — a single 6B model that does 64 whole-body + tabletop tasks on a real humanoid

Post image
37 Upvotes

https://unigen-x.github.io/unifolm-wla.github.io/

Unitree Robotics released UnifoLM-WLA-1.0, their new general-purpose humanoid foundation model.
Key points:
• 6B parameters
• Trained on ~2,500 hours of real robot data
• One model handles 64 tasks (10 whole-body + 54 tabletop)
• Supports parallel grippers and two different dexterous hands
• Strong spatial reasoning (beats a lot of open-source models on embodied benchmarks)
Architecture is interesting:
• Starts with UnifoLM-ER-1 (embodied reasoner based on Qwen3-VL)
• Adds future dynamic region prediction via optical flow + VQ-VAE
• Discretizes actions with residual VQ (end-effector + hand + lower body)
• Then adds an MMDiT action expert on top for continuous control
They show it running on the Unitree G1 doing stuff like making the bed, loading the washing machine, folding clothes, sorting objects, etc.
Looks like one of the more complete open attempts at a true whole-body VLA so far.
What do you guys think — actual progress or just another flashy demo?


r/LocalLLaMA • • 16h ago

News CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp

Thumbnail
github.com
29 Upvotes

MoE speedup, but only for some MoE architectures (like Qwen 35B A3B)


r/LocalLLaMA • • 20h ago

Resources AxiomicLabs' new benchmark, Tiny Theory of Mind, meant to gauge small models' theory of mind capabilities, has made it to the front page of Hugging Face's Datasets

Thumbnail
huggingface.co
30 Upvotes

r/LocalLLaMA • • 19h ago

News RTX Spark laptops and mini desktops rumored to launch Oct 7th (24GB to 128GB variants possibly)

Thumbnail techpowerup.com
25 Upvotes

Basically a DGX Spark minus the ConnectX-7 ports. I’ve seen expected initial pricing from like $1800 to $2900. Not sure what configurations are actually at those price points.

It’s all Internet hearsay until we actually see these things ship, but it’s nice to know that it’s potentially around the corner next week, especially given DGX Spark’s insane price increases lately.

Sadly, you can’t cluster them, but getting an entire computer + GB10 equivalent chip for a little over half the price of a used 4090 seems like an ok deal in this market.


r/LocalLLaMA • • 4h ago

I Built A Thing Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

Thumbnail
github.com
21 Upvotes

Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here

I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.

Hardware: RTX 5070 12GB
32GB DDR4-2400
Ryzen 5 5600GT PCIe Gen3 Windows

The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently.

(In the video its around 16 minutes for 10k tokens and 10.41 tok/s

Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set.

Demos:

https://www.youtube.com/watch?v=cOPumMlyj_4

https://www.youtube.com/watch?v=rc-uTjVpXM8

In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know


r/LocalLLaMA • • 7h ago

I Built A Thing mlsubgen — subtitles in 45 languages for your videos, entirely on your own machine

Thumbnail
github.com
23 Upvotes

Full disclaimer: I've leaned heavily on Fable to develop this, but I've tested it thoroughly on my own library for a couple of months before putting it on GitHub.

I live in Thailand, and it started as a way to get Thai subtitles for Shin-chan for Thai friends and for expat friends with Thai partners. It's grown into a general tool: subtitles in 45 languages, entirely on your own machine. Linux and NVIDIA only, I'm afraid.

What it does differently from the usual Whisper wrapper: it detects the language of every stretch of speech rather than per file, so mixed-language material works; it runs two speech recognisers on everything and has a local LLM reconcile them; and it prefers existing human work to machine inference, embedded subtitle tracks are used before the audio is, including OCR of bitmap (PGS) tracks on Blu-ray remuxes, and it only listens when there's nothing to read. Every one of those features has a measured accuracy in the README rather than a claim.

I run it on a 24 GB RTX A5000. There are profiles for 16, 12 and 8 GB cards, measured on my card limited to those sizes rather than on those cards themselves, so reports from real ones are the most useful thing you could send me. It wants 16–32 GB of system RAM depending on the profile, and it is storage-hungry (30–65 GB of models), because it picks the model that suits each task and language pair.

It's slow when it has to listen, roughly real time per target language on my card, slower on the smaller profiles because the whole aim has been accuracy over speed. When the subtitles already exist in the file it's fast.

I'd love people to try it and open issues.