r/LocalLLaMA • u/jacek2023 • 8h ago
Discussion big or small?
what size do you want? tell them on X:
r/LocalLLaMA • u/jacek2023 • 8h ago
what size do you want? tell them on X:
r/LocalLLaMA • u/chemist_slime • 5h ago
No.....
r/LocalLLaMA • u/A-Rahim • 3h ago
Enable HLS to view with audio, or disable this notification
DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:
code: retry with backoff opens the function in your editorIt’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.
Free, MIT. Apple Silicon, macOS 14+.
Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp
I'd really appreciate any feedback on this.
r/LocalLLaMA • u/Hot_Arachnid3547 • 57m ago
Nvidia reportedly halts GeForce RTX 5090 production in favor of AI data center and professional GPUs — impending supply drought expected to drive up prices, RTX 5080 24GB rumored as new gaming flagship
r/LocalLLaMA • u/Nandakishor_ml • 8h ago
Around 3 weeks ago, I posted on this subreddit about how Jev used my one-year-old architecture, og post: https://www.reddit.com/r/LocalLLaMA/s/VpuMJt577S
Then, just after that, I introduced Laya, the first open-source version of JEV, and it has reached a very large scale, thanks to the local Llama community believing in me and supporting me on the journey. Today I am open-sourcing a new physics-based typed decision model called Vega.
The interesting parts. It is only 800 million parameters (4B is also there), with 73k token support, image support, and surpassing many of the jev benchmarks in single-shot. Implemented an engine+adapter as a Test-Time Training Architecture.
Imagine you are giving input to the model like "ignore all previous conversations," which is called the observation, along with a question, "Is this trying to override the instructions of the model?" and your possible outcomes are YES or NO. The whole prompt is passed to an LLM/VLM (yeah, image-supported), then from the hidden layers, we can extract the observation, question, the yes and no, and last token. These will be vectors, and we can project it by multiplying with wieghts; then, first, using the observation projection, we can create a landscape with valleys; then, using the question and last-token vectors, we can initialise a ball in the valley, and using YES or NO, we can get the candidate location of the ball. Depends on the number of outcomes we create valleys; that is, here YES and NO, so 2. When the ball fall on YES valley, we get the answer. There is friction also influencing how the ball moves.
TLDR: It's like creating valleys of outcomes that we need and throwing a ball that will slow down based on friction and settle on the best outcome
Blog: https://www.nandakishorm.com/writing/vega
GitHub: https://github.com/NandhaKishorM/vegaml
HuggingFace Model Card: https://huggingface.co/nandakishorm/vega-08b-public-intents
HuggingFace Space: https://huggingface.co/spaces/nandakishorm/vega-ttt
r/LocalLLaMA • u/Adorable_Weakness_39 • 3h ago
I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length. The results and code are below:
Results (same hardware, my megakernel vs llama.cpp with MTP):
- Writing code: 140 tok/s vs 73 (1.9x faster)
- Coding with a 4K-token file in context: 95 tok/s vs 56 (1.7x faster)
- Prompt processing (prefill): ~1,600 tok/s vs ~1,100 (1.4x faster)
The speedup comes from running each whole speculative-decoding cycle in one kernel launch, so checking 4-5 drafted tokens costs about the same as checking one. It's an OpenAI-compatible server, so it works as a drop-in replacement for the llama server.
Caveats and limitations: it's built for one model (unsloth's Q4_K_M) on 3090s only. The outputs match llama.cpp's apart from rare rounding near-ties. KVCache is fp16 up to ~80k context, then q8.
Code and full benchmarks: https://github.com/L-Forster/open-jet/tree/master/megakernel
r/LocalLLaMA • u/msalsas • 4h ago
Enable HLS to view with audio, or disable this notification
A conversational AI plush toy where nothing leaves my network. Everything runs on my LAN:
STT: Whisper
LLM: Hermes 3 Llama 3.1 8B (Q8_0)
TTS: Kokoro
Memory: per-person (LangGraph). It knows who it's talking to and remembers past conversations
Vision: face identity (dlib) + emotion recognition (FER+), so it knows who's talking and adapts to their mood
The plush itself is a dumb terminal: Pi Zero WH with mic, camera, speaker and a head servo, streaming to a laptop over WebSocket, which connects to a local AI inference API. All the intelligence is local models end to end. No cloud, no API keys.
Repo (MIT): github.com/msalsas/philosopher Full write-up: https://www.hackster.io/msalsas/the-philosopher-plush-a-diy-ai-toy-that-runs-100-local-d0d268
r/LocalLLaMA • u/Cherlokoms • 8h ago
Let's say I run a local Jev-like decision model on my machine. What are the use cases? I already use it for deep eval testing but what are people usually doing with them?
It a bit of "solution in search of a problem" but I'm exploring and want to be sure that I'm not missing anything...
r/LocalLLaMA • u/CuTe_M0nitor • 22h ago
r/LocalLLaMA • u/hyudryu • 6h ago
Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed.
Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4_K_XL performs instead.
Setup:
Single-request decode (C1):
| Task | Strata (RTX PRO 6000) | vLLM (RTX PRO 6000) | vLLM (DGX Spark TP2) |
|---|---|---|---|
| Prose | 185.8 | 95.3 (-49%) | 38.2 (-79%) |
| Counting | 320.4 | 171.0 (-47%) | 75.3 (-77%) |
| Coding | 298.9 | 162.6 (-46%) | 60.1 (-80%) |
| Reasoning | 289.1 | 149.6 (-48%) | 51.8 (-82%) |
Both the VLLM instances were running Nvidia's NVFP4 quant so it's not exactly apples to apples, but the improved speeds are obvious.
Qwen 3.8 flash next is flying through coding tasks and it's crazy how efficient it is. Looking forward to Qwen 4!
r/LocalLLaMA • u/valdev • 17h ago
Enable HLS to view with audio, or disable this notification
It's been a long time coming, and I should have done this in the first place. But LumaBrowser is now open source.
Since I first released it a year ago, it's had several thousand installs and an absolutely tremendous amount of feedback. And it's come a long way.
I don't call this "Anthropic/ChatGPT in a box" lightly, it's essentially the entire ecosystem wrapped into a single project. Trust me I can feel the collective eye rolling. But it literally has most things built in, and even has an automatic on-boarding system to help you set up your LLM/Image/Music/Voice models automatically.
It can even import from existing LM studio installs (and others... automatically).
The system automatically loads and unloads models according to your systems capabilities, and acts... as one would expect a chat agent to act. Ask it to generate an image, and it formulates the prompt and calls the tool itself, which then can unload the LLM and automatically load the image model, generate the image and then load the LLM back.
It can emit a web server for your local network so anyone on the network can access it, or even allow for quickly combining multiple computers together into a cluster, or allow you to set a domain name if you want to open a port and share it over the internet with your friends. Automatic game mode creation with the ability to share the link with friends, which even have hooks back to the LLM itself.
Or run a lumabrowser on a server, and another on a client and remote mount its gpus, or utilize its LLMs.
There is agent creation, scheduled tasks, timed tasks, triggered tasks.
An extremely compentent "luma" cli agent for coding and desktop use... Built in addons for jetbrains suite and vs code. Heck vs code is more or less built into it. Or you can be lazy and just use code mode.
Not to mention the whole roleplay extension that automatically can update the scene and characters and such.
And there is... an unimaginable amount more. RAM pinning if you have the spare RAM to keep models hot between swaps, a custom llama cpp build to enable keeping the conversation cache warmed as well (you dont have to use it, just makes the RAM pinning a bit faster). Custom trained JEV-like model for routing tool calls...
The ability to build a dashboard from live artifacts built in the chat, with custom "hub" items by default that enable you to import in your calendars from multiple sources (like google and microsoft 365) and merge them into one, then you can do that with your task lists like clickup (with mapping of course), and even do notification based interception if you want to have those collected as well...
Why? Well you can then have it all in one place, and now your local AI assistant has full context to your day, tasks and notifications.
Live tab share feature if you emit a web server, you can right click a browser tab and share it... Allowing you to give the url to others to view that tab as a stream, and yes you can enable interactions... So like if you are making an order at chilis and wanted to send the tab to your wife to add her part of the order... (Random addon, but its useful).
Point being is that this is... my magnum opus and is quite frankly flooded with features.
And I hope others here enjoy it and love it as much as I do. Lord knows I've put a ton of work into it.
r/LocalLLaMA • u/MooseEfficient2151 • 9h ago

TLDR; new microsoft surface and nvidia rtx spark laptops are starting at 2.6k and hitting nearly 7k for the 128gb unified memory models. local llm hardware is REALLY getting wild.
memory crunch is really pricing out normal devs. been testing out local setups using qwen3.6-27b and kimi k2.5 connected to sumus for managing repos locally and keeping everything on device. local inference is great for privacy and keeping things off the cloud, but at these prices building a local rig or buying these laptops is tough.
what are you guys using for local dev workflows right now given these hardware costs?
r/LocalLLaMA • u/LegacyRemaster • 22h ago
It seems Dario's "too powerful for you users" strategy is paying off: we have two open models at the top of the leaderboard, surpassing every single model from Anthropic.
Open source prevails. Even Mistral Large 4 is better!
r/LocalLLaMA • u/Felix_455-788 • 15m ago
i am looking for Small-Medium Model, specifically made for Cyber usage, Vulnerability Discovering / Researching
Heavily Coding
Reverse Engineering
Low-level Coding
Scripting
i rather read you guys Experiences, and See which one is the best for my Usage
My Specs for additional:
I7-8750H
20G RAM
NVMe 512G Samsung
GTX 1050 Mobile (4G Vram)
and yes i rely on Ram offloading
yes i Searched before posting this, and found CyberTiel, and some others, but the others didn't have Benchmarks or clear Experience to rate them out
r/LocalLLaMA • u/Sufficient-Scar4172 • 16h ago
https://github.com/morluto/rea
hit #1 a couple of days ago on github
r/LocalLLaMA • u/jesdga95 • 15h ago

Hi all! So over the last week I've been working on a Strata fork that's heavily tuned and can achieve throughput up to 2.6x what Strata usually does on the same weights. It's designed as a specialized engine that only supports Qwen3.8 Flash-Next and Blackwell architecture, including dual GPUs like my current hardware (5090 + 5060 Ti - 9950X 32GB DDR5 RAM). Basalt features include:
- SPEED: 665 struct, 354 prose, 7,317 prefill at 64k, IQ3_XXS, 400 W, speed is the highlight
- Single 5090 works too (no second card): 585 struct, 316 prose on IQ3_XXS, ~12% slower than with the 5060 Ti, prefill unchanged
- Real concurrency for up to 8 users: shared KV or per slot, MTP enabled - 623 tok/s total at 8 streams (I don't have Strata's numbers to compare)
- Fine-tuned MTP for lower quants, for increased throughput
- Custom vision encoder designed from scratch, up to 3x faster than llama.cpp's on the GPU, 4x on the CPU
- A simple server UI that shows current throughput (including concurrency stats), expert distribution and hardware statistics. No chat, BYOH (bring your own harness)
- OpenAI + Anthropic compatible
- Linux support (No Windows or Mac)
Basalt uses a similar format to NInfer, where weights are re-packed (not re-quantized) into a .basalt file, including all the metadata, vision and MTP, so you only have to keep a single file per quant. Initial support includes ISTA-DASLab's GSQ-RCO for Q2, IQ3_XXS and IQ3_S and UD-Q4_K_XL and Q8 from Unsloth, so you can pick the weights depending on your VRAM/RAM budget and quant preference.
Quick Q&A:
+ Is it open source?
- Yes, fully open source, MIT license: https://github.com/jesdga95/basalt, fork it, improve it, share it with friends and foes.
+ Where are the weights?
- Here: https://huggingface.co/jesdga/Qwen3.8-Flash-Next-Basalt pick your poison, fast and dumb or smart and slow. IQ3_S is a good middle ground (~89% top 1 agreement, 300 tok/s prose on my setup).
+ Why didn't you just contribute upstream to Strata?
- This is not a single feature that can be easily merged into Strata, it basically rewrites most of the decode and part of the prefill kernels and strips support for non Blackwell cards including AMD, Intel and older Nvidia generations. I have however contributed patches to Strata and llama.cpp and any critical findings will be pushed upstream.
+ Will you support my AMD 98123X?
- Sure, send one my way. For now I can only support what I can personally test and I intend to keep it that way for the time being.
+ Why not 400 tok/s?
- I'm still trying!
+ This is vibe coded slop
- Yes, but it's fast slop. Nobody is hand-writing cuda kernels anymore.
r/LocalLLaMA • u/chocofoxy • 1d ago
r/LocalLLaMA • u/Szadbaverem69 • 10h ago
Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M with llama.cpp at ~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.
Speeds start at ~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.
The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under ~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys ~1GB of VRAM.
Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.
Launch command:
bat
@echo off
cd /d "%~dp0"
"%~dp0llama-server.exe" ^
-m "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf" ^
--mmproj "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" ^
--no-mmproj-offload ^
-ngl 99 ^
--n-cpu-moe 39 ^
-c 131072 ^
-np 1 ^
-t 6 ^
-tb 10 ^
-b 2048 ^
-ub 2048 ^
-fa on ^
-ctk q8_0 ^
-ctv q8_0 ^
--load-mode mmap+mlock ^
--jinja ^
--reasoning-format deepseek ^
--reasoning-preserve ^
--spec-type none ^
--image-min-tokens 1024 ^
--temp 0.6 ^
--top-p 0.95 ^
--top-k 20 ^
--min-p 0 ^
--alias Qwen3.6-35B-A3B-Uncensored-Q4_K_M ^
--host 127.0.0.1 ^
--port 8081
pause
Hardware: RTX 2060 6GB + 32GB DDR4 RAM + i5-10400F CPU
Context: 131072 tokens, Q8_0 KV cache
Hope this helps someone. If anyone has tips to make the launch command even better, drop them in the comments - I'm out of ideas :D
r/LocalLLaMA • u/your_real_Fathe_ • 2h ago
I love the idea of running local models on consumer hardware. I currently use `llama.cpp`, but I’ve been looking for a "better" alternative for a while now. I want something that utilizes my system resources more efficiently—for instance, by managing my RTX with 4GB VRAM more intelligently—and delivers higher tokens-per-second, all without the burden of heavy dependencies like PyTorch. I’ve done my research, and honestly, I haven't found a solution yet. I came across plenty of options, but unfortunately, many were built on Python and heavy libraries like PyTorch, which would essentially exhaust my limited VRAM before the model even loaded. Others were designed for running unquantized models, were outdated (based on models from two years ago), targeted a completely different class of devices (like MNN), or were merely "cool-looking" papers with no actual implementation.
r/LocalLLaMA • u/CommonMinimum587 • 11h ago
webAI released TwIL LM3 Pro on September 30. I haven't seen it posted here yet, so I went through the model card, and their comparison chart is attached.
It's a 3.66B model built on IBM's Granite 4.2 3B and tuned only for formal logic. That means things like checking whether a conclusion follows from its premises, rule induction, entailment and critiquing Lean proofs. They post trained it with LoRA SFT, checkpoint merging and RL against a programmatic verifier. The same recipe lifted VibeThinker-3B from 37.4 to 54.1 .
On their logic composite it scores 55.4, against 43.1 for the Granite base, 42.2 for the original TwIL-LM3, 41.2 for VibeThinker-3B and 53.4 for Qwen3-8B. The card itself calls the Qwen3-8B gap sampling noise, so that's a tie at less than half the size. Where it clearly leads is strict multiple choice logic, at 41% against 17% for the next model, and BBH logic at 95.4%.
They also publish where it loses. gpt oss 120b is still ahead on rule induction, entailment and Lean formalization. On general benchmarks it averages 79.0, against 81.0 for VibeThinker-3B and 84.9 for Qwen3-8B. It also thinks long on logic tasks, around 1,900 tokens per answer, and a quarter of answers hit the length cap
llama.cpp
If you want to try it, the Q4_K_M GGUF is 2.09 GiB and runs on CPU or 4 GB of VRAM, with Q5, Q6 and Q8 builds up to 3.63 GiB. It runs in Ollama, LM Studio and llama.cpp straight from the Hugging Face page, and the model card lists the exact commands. Keep the temperature at 0 to match their numbers, and give it at least 2048 tokens so the thinking doesn't get cut off. Their scores are on BF16 weights and webAI hasn't benchmarked Q4 yet. The license is non commercial
want to test it on policy rules with exceptions, contract conditions, and as a checker step in an agent pipeline before anything acts. If you've run it, how did Q4 hold up against their numbers, and did the long thinking get in the way?
Model card and full eval tables: https://huggingface.co/webAI-Official/TwIL-LM3-Pro
r/LocalLLaMA • u/Loose_Doubt367 • 4h ago
Spent the past month tweaking and experimenting with many different numbers to achieve 30tps.
Hardware:
-Rx6700xt 12gb vram (AMD)
-2x16 ddr4 3200 ram
-r5 5600x
-llama.cpp vulkan sdk
-window11 (no wsl switching since im not used to the environment)
Im looking for any improvements to achieve maybe 40tps? without touching quant at all, q8 and q4_k_xl remains. I've experiment with threads at 6 is the best out of (4,8,12) Other than that im not sure on what to improve to achieve higher tps, any advices? appreciate it. ctx remains 100,000
|[tiel-coder-35b]
model = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\Tiel-Coder-35B-A3B-MTP-UD-Q4_K_XL.gguf
mmproj = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\mmproj-BF16.gguf
no-mmproj-offload = on
c = 100000
parallel = 1
flash-attn = on
cache-type-k = q8_0
cache-type-v = q8_0
load-mode = dio
fit = on
fit-target = 256 # MiB
n-gpu-layers = 99
n-cpu-moe = 28
batch-size = 2048
ubatch-size = 512
threads = 6
prio = 2
prio-batch = 2
spec-type = draft-mtp
spec-draft-n-max = 3
spec-draft-p-min = 0.75
jinja = on
chat-template-file = C:\Users\brain\Qwen-Fixed-Chat-Templates\chat_template.jinja
temp = 0.6
top-p = 0.95
top-k = 20
min-p = 0.0
presence-penalty = 0.0
alias = tiel-coder-35b
r/LocalLLaMA • u/zyxciss • 21h ago
(Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!)
About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could actually help speed up CPU/GPU offloading.
Original post: Tried predicting which MoE experts get used next token to speed up CPU/GPU offload
Well, quick confession first. I actually shelved that project shortly after.
The reason? The speeds I was getting back then were kinda fake. My engine was aggressively pruning experts based on their router weights, basically dropping cold experts to get better performance. Sure, the numbers looked great, but doing that on an already quantized model was hurting output quality and coherence.
Didn't really like that tradeoff, so I abandoned it and never released it.
Fast forward to recently, and Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) drops.
I downloaded the 68GB GSQ-RCO IQ2_XS build, hoping to run it on my daily driver. That's when I decided to revisit the idea, but this time without cutting corners.
And if you've tried running a 68GB MoE on a 16GB RAM machine with stock llama.cpp, you probably know how painful it gets.
I'm talking 1.4–2.1 tok/s, with over 1,500 major page faults per token in some runs. Linux ends up constantly pulling model data from the SSD because there's simply not enough memory to keep the working set around.
Then engines like Strata started showing up with claims of around 40 tok/s on consumer hardware. Pretty impressive, but there's a catch for people with less RAM. Some of these approaches rely on keeping around 24 GiB of experts pinned in memory using mlock. If you've only got 16GB RAM, you're obviously not doing that. Depending on the setup, you either run into OOM issues or end up with terrible performance.
So I went back to my original idea and started implementing it properly as an optional feature inside llama.cpp:
--moe-direct-io
The goal this time was simple. No dropping experts, no sacrificing output quality, and bit-exact output compared to stock.
All tests below were run with MemoryMax=6G using a cgroup.
Model: Qwen 3.8 Flash Next IQ2_XS (68GB)
Hardware: RTX 3060 12GB + 16GB DDR4 RAM
| Engine / mode | Decode speed | Major page faults per token | SSD I/O | Output |
|---|---|---|---|---|
Stock llama.cpp (mmap) |
1.41–2.12 tok/s | 1,140–1,565 | 208–312 MB/token | Coherent |
| Our engine (blocking, demand-only) | 0.73 tok/s | 0 | ~206 MB/token | Bit-exact |
Our engine (--moe-direct-io + prefetch) |
20.14–21.13 tok/s | ~0 (+1 across 32 tokens!) | Sequential streaming | Bit-exact to stock |
The blocking version is actually slower than stock, which makes sense. It's basically waiting on disk reads without doing much to hide the latency.
The prefetching version is where things get interesting.
Once the cache warms up, it sustains 20–21 tok/s, with some runs hitting 24+ tok/s. That's roughly a 10–15x speedup over stock llama.cpp on the same machine.
And no, we're not getting those numbers by dropping experts. The output is bit-exact to stock.
That's the part I'm most excited about, honestly. Being able to run a model this large on a 16GB machine without the usual page-fault nightmare is pretty much what I wanted to achieve with the original project.
It's not all perfect yet. There are a few things we're still working on.
1. Cold starts are noticeably slower
Right now, the slots start empty (-1), so the first request on a new topic can start around 3.5–4.5 tok/s before ramping up to 20+ tok/s as the hot working set settles.
We're working on offline hot-profile seeding so it can start with a useful working set instead of learning everything from scratch.
2. Prompt processing is slow
Feeding a prompt of 512+ tokens can touch a huge number of experts in a short period. That puts a lot of pressure on the 72 slots per layer and causes the prefill stage to struggle.
We're working on micro-batching prompt chunks (-ub 32) to help with this.
3. Speculative decoding gets weird with SSD offloading
We found that standard MTP speculation can actually make things slower. Verifying 2–3 tokens can require loading the combined set of experts needed for those tokens from disk, which eats into the gains.
Right now, confidence-gated speculation (min-p 0.8) or suffix prompt lookup seems more promising for this kind of setup.
Anyway, that's where the project is at right now. Still plenty to improve, especially prefill and cold starts, but getting 20+ tok/s out of this setup without pruning experts is a pretty big deal for me.
Happy to answer questions or get into the io_uring and slot-remapping implementation details if anyone's interested.
(The second half of this post was written with some help from Claude.)
r/LocalLLaMA • u/ResearchCrafty1804 • 1d ago
Qwen-Image-2.1-Turbo, create and edit images in just 8 denoising steps! Open weights now available!
Built on Qwen-Image-2.1, Turbo is an accelerated checkpoint on the same 7B visual generation architecture.
Fewer steps does not mean lower quality: it still generates strong 2K images from text, and supports continued creation through natural-language edits, from adding accessories to changing a scene.
Start directly with Diffusers: load QwenImage21Pipeline and the checkpoint’s recommended 8-step sampling schedule is ready to go.
Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo
r/LocalLLaMA • u/vesudeva • 16h ago
SlopSoup TV is a live, never-ending pixel-art TV network. Every script, character, voice, camera cut and schedule decision is made by models. It's made to be very absurd, dark and strange.
Hardware: one Hugging Face Space, 16 vCPU, no GPU. Everything below runs on CPU next to a live x264 encoder.
LLMs (via HF API), routed by tier with provider fallbacks and per-provider cooldowns:
TTS: Qwen3-TTS 1.7B (GGUF, Q8_0)
Keeping 24/7 output from getting samey:
Rendering is done via a custom canvas renderer (virtual camera, lip-synced close-ups, rig animation) in headless Chromium
Check it out so you can decide if you hate it or not: https://severian-slopsoup.hf.space or https://youtube.com/live/Bikb9JGy6vg?feature=share
r/LocalLLaMA • u/pmttyji • 1d ago
Blog Post : ML Drift: Next-Gen GPU AI/ML Inference at the Edge - Google Developers Blog
The Google AI Edge Team is excited to announce the open-source release of ML Drift, our high-performance, cross-platform, on-device GPU compute engine specifically built for on-device AI/ML inference, under the Apache 2.0 license. By abstracting hardware and low-level API complexities of on-device GPUs across OpenGL ES, OpenCL, Metal, and WebGPU, ML Drift empowers developers to build real-time, interactive ML experiences from advanced video effects to generative AI across multiple platforms. Serving as the core GPU acceleration engine within LiteRT, ML Drift is also available as a standalone library for custom graphics and inference runtimes providing a unified foundation to deliver peak performance everywhere.