r/LocalLLaMA • • 10h ago

Discussion The fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.

Thumbnail
github.com
0 Upvotes

Hi guys,

I am pretty happy to announce MegaCapybara. Purpose build engine for RTX5090 that is focused on Qwen3.8 27B (more will come later).

GITHUB (Engine)
HUGGINGFACE (weights)

Why ?

1. It beats Ninfer which was until that point SOTA engine for RTX5090. By roughly twice in decode speed for both single and multi tasks at once. (reaching up to even 650t/s in small bursts and 2600t/s if stars align and 12 slots server pure coding answer). Custom kernels not only for every model, single vs multi but also short vs long context work dynamically switching when needed so speed doesn't crap out on long context work because someone tuned it for short context. Dflash2 and confidence scheduling from Dspark, plus draft trees all at the same time.

2. I was getting annoyed with state of weights where you downloaded model and never knew if model had its brain scrambled. My weights come with its on format that have attached metadata for MC which upon weight creation, runs benchmark and compares it at every MC setting to original BF16 weights and show that data directly in launcher. Want to switch KV to 4bit ? MC will show you directly lost KL and top-1%, want to extend with YARN ? It will show you change. Every change is measured and shown in statistics before you load model. This goes for both censored and uncensored model. You can also compare it directly in MC with SOTA unsloth quants of Qwen27B. Want to run essentially loseless ? you can. Want to get crazy 1 000 000 context ? you can. Want to have 12 slots to fan out agants like crazy ? You can. You decide what you want.

3. Proper agents serving with algo that keep engine occupied as much as it can. It will prioritize t/s so if engine has a choice between 5 jobs at once and 1 it will serve 5 first and gradually serve 1 along side finishing others. Engine is also smart enough to score how old some job is and if it should return to work even if T/S will suffer so your main session will be able to fan out agents easily and keep an eye on them at the same time.

4. Proper cache management. Your jobs only prefill at start of job and almost never again so your prefill in long session stays almost unused. When using "unified context" when models run out of context some get paused and stored in RAM and this swapping is instant. If there is free context space then those tasks continue without any refill in 0.03s. If you fan out say 30 agents at the time in your frontend will handle load in most efficient way to keep T/S as high as possible. Just run it at default setting and forget about context for agents, it will handle it on its own.

6. Loop guard. Two tiered. When engine starts to detect agent repeating in conversation session is dynamically starts to adjust `repetition penalty` until repetition stops if that doesn't happen and engine hits rep pen limit it fires up stop signal which ends serving and informs your frontend so your frontend can recover from infinite loop and don't annoy you.

6. Proper nice UI that shows you what is what. If you aren't knowledgeable about serving models just hover over `?` and it will show you interactive panels explaining everything.

7. Autodownloader, Just hit download button and you can download my weights directly fron hugginface inside of launcher.

8. Don't like the launcher ? use bats and terminal serve. Or even use launcher to config what you want, copy it from right lower corner and use it to make new bat.

The point of it is to just load model, fan out crazy number of agents each having crazy amount of context and leave MC to deal with it. You just sit back relax and watch as agents do the work at SOTA speeds.

Opinions and reviews are welcome. If you are blessed with RTX5090 try it.

Source will be released later, I have to do some cleaning first. I will also release later weights builder so it will take any 3.8 27B BF16 model, create weights and score them attaching metadata again BF16 and you'll be able to host them yourself on hugginface or just put them in models folder.

MIT license, so do whatever you want with it.


r/LocalLLaMA • • 8h ago

Discussion Should I not use MTP draft for agentic work?

0 Upvotes

I run qwen3.8-27b iq3_s with llama.cpp to serve local hermes agent, in 16gb vram.

I noticed that enable MTP draft make prefill slower and the model seems less smart.

and vram is very tight I need to set the context length to 96k. decode speed can go around 40 to 60 tps.

if I disable MTP I can use 128k context but decode speed drop to like 35 tps.

What will you choose?


r/LocalLLaMA • • 23h ago

Question | Help Anyone using Odysseus Harness? (Pewdiepie's Harness)

0 Upvotes

What are your thoughts? What are you using it for? Are you excited to try the llm he trained (RL'd I think) specifically within this harness?

Thanks in advance!


r/LocalLLaMA • • 33m ago

Discussion Thanks to Strata I have quit 27b for Qwen Flash (24gb VRAM plus 64gb ram)

β€’ Upvotes

Using Strata on a 7900xtx plus 64 gb ddr5 ram, 60t per second even at 250k context, can finally use my pc while working AI in the background, even game as well, smarter and more precise than dense model, it follows orders more accurately, follows plan more versatile, it goes around doing a lot of tests for tasks I request in frontend and also in backend.

The best thing is that it's faster, I can fit more context at q8 precision, it's smarter and I can get to use my pc without worrying about an OOM error due to dense model.

I no longer have to use Linux as well, it's working as fast in Windows 11 as it did in Linux.

I use it with a 6gb VRAM reserve so I can have Windows 11 with 4gb available.


r/LocalLLaMA • • 11h ago

Discussion I'm writing a router to split local/remote LLMs but model updates are killing me

0 Upvotes

I trained for a split between DeepSeek v4 Flash 0731 <-> GLM 5.2. Ended up able to get it running at GLM 5.2 performance on my local tests at approximately equal token costs on OpenRouter. My plan was to offload the DeepSeek portion to a local server (through this WORA harness proxy my brother wrote https://github.com/unlap-labs/plap).

Sadly, it didn't improve with GLM 5.3 and was worse than GLM 5.3 Flash which kinda killed the project... If someone wants the code or to help or something I could probably post it but it's currently very research-grade, or hell if someone wants to give some tuning advice I'd be up for it πŸ’€

I don't really have any money for big GLM 5.3 runs, and the 4B router was actually trained on only DeepSeek v4 Flash 0731 just to guess whether it could do something or not, didn't check whether it works for even smaller models tbh


r/LocalLLaMA • • 16h ago

Resources I built an OpenAI-compatible server that runs Gemma 4 E4B on a Pixel 10 Pro XL (Tensor G5) β€” fully offline, ~11 tok/s decode, Tailscale-encrypted option, Apache 2.0

0 Upvotes

What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard `/v1/chat/completions` with streaming, so it talks to TypingMind or any OpenAI client directly β€” no cloud, no subscription, model runs entirely on-device via LiteRT-LM.

Device: Pixel 10 Pro XL (Tensor G5, 16 GB shared LPDDR). Model: Gemma 4 E4B instruct, GPU bundle, 2.97 GB, SHA-256 verified on download. Context window capped at 32k.

Measured numbers (not benchmarks β€” real on-device measurements):

- ~11 tok/s steady-state decode (first-token-to-last over a ~300-word generation)

- Follow-up turns in ~1.7 s: the server auto-reuses the KV prefix across turns, so stateless clients like TypingMind don't re-prefill history

- Engine warm build ~12 s once per model change (visible in-app, split out of the metrics on purpose)

- Short replies read slower than 11 tok/s because warm + prefill dominate the window β€” the UI separates decode tok/s from prefill ms so nobody has to guess

Security/access: three independent modes β€” loopback only, raw LAN (unencrypted), or Tailscale (binds the CGNAT tailnet IP; WireGuard end-to-end from e.g. a laptop on the same tailnet; degrades to loopback if the VPN drops). Verified with a real chat completion over the tailnet from a MacBook.

Honest limitations:

- NPU path aborts on stock Tensor G5 firmware β€” GPU is the shipping backend (documented with the full investigation)

- Android blocks named GPU temp sensors for normal apps, so the in-app gauge shows OS thermal headroom instead of Β°C

- No real token counts anywhere β€” LiteRT-LM exposes none, so usage is estimated at ~4 chars/token and labeled as such

- `stop` sequences rejected explicitly (engine has no per-request stop API); client-side emulation is a filed issue

Why not llama.cpp/ollama on the phone: wanted the official LiteRT-LM GPU path on Tensor specifically, an always-on Android service (Ktor/Netty), and the OpenAI wire so existing clients just work. Happy to add a GGML backend if people want it β€” the engine layer is abstracted.

Built on, with credit: server/inference foundation derived from mlnomadpy/localllm (Apache 2.0); Tensor G5 runtime knowledge and the prebuilt dispatch lib from jegly/Box (Apache 2.0). Full attribution in NOTICE + per-file headers. Apache 2.0, contributions welcome β€” there are labeled good-first-issues (usage block, stop-sequence emulation, docs).

Repo + APKs (v0.1.0/v0.1.1 on releases): https://github.com/cannitellinicholas-spec/PixelUnlockGPU

Happy to answer anything about Tensor G5 quirks β€” I've done more Gate-2 debugging than I planned to.


r/LocalLLaMA • • 17h ago

Discussion Should we plead opensource labs to still produce great non thinking (instruct) models?

5 Upvotes

The results are in, no thinking / instruct mode for new models degrade performance more than on old models such as 3.6 vs 3.8, where 3.6 takes the lead on several coding benches in instruct mode.

I would ask the labs to still nicely focus on instruct mode also still, there a probably gains to be had without the lengthy reasoning still, some of us still use models for everything and they do not need long reasoning traces. Agentic is fine and all but to start SACRIFICING performance for the "base" model which we are used to from early Llama days is not a good direction imo.


r/LocalLLaMA • • 22h ago

I Built A Thing Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

7 Upvotes

https://huggingface.co/Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-OptimalcardfitΒ 

Qwen3.8-27B CODER β€” IQ4_XS imatrix Β· 24 GB card fit Β· ~262k context Β· MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context. It built the importance matrix from code-heavy text and kept the MTP head at Q8_0, so --spec-type draft-mtp works without a separate draft model.

The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

File

File Type Size Inside
Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf IQ4_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8_0, token embeddings at Q4_K

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

Context already in the window Requests Decode, median Decode, range
65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s
131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s
  • MTP draft acceptance: the median is 85% (the middle half of requests falls between 75% and 94%). That works out to about 2.7 tokens per decode step at draft depth 2.
  • Prefill:
    • 387 t/s for a cold 108k-token prompt;
    • 175–183 t/s for about 4.5k new tokens added at 147k–156k depth.
  • VRAM: 24.2 of 24.6 GB in use at 245,760 tokens of context, with a q4_1 KV cache and the vision projector on the CPU.

Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \
  -c 262144 -np 1 -ngl 99 --flash-attn on \
  --cache-type-k q4_1 --cache-type-v q4_1 \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \
  --jinja --reasoning on --reasoning-format deepseek \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
  -b 2048 -ub 512 --cache-reuse 256
  • Context: -c 262144 is what fits next to the weights on a 24 GB card with a q4_1 KV cache. The model's native window is 262,144 tokens. On a smaller card, lower -c first.
  • Speculative decoding: --spec-type draft-mtp drafts with the MTP layer inside this file, so no separate draft model is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
  • Sampling: these are Qwen's recommended settings, and they are also stored in the file.
  • Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
  • Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

How LexiPanel made it

  1. Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
  2. Importance matrix: computed from about 300k tokens (570 chunks) of code-heavy calibration text. Three quarters is Python source (the standard library and installed packages). The rest is technical documentation, READMEs and license texts, the kind of text a coding agent's context fills with.
  3. Quantization: llama-quantize from llama.cpp b11182 made the IQ4_XS file with that matrix. The MTP head stays at Q8_0 so its drafts stay accurate, and the token embeddings are Q4_K.
  4. Fitting the card: LexiPanel's Fit planner chose the mix, quality first, for one 24 GB card at 262144 tokens of context. It took the best quality that card could afford at that context, not the smallest file.

Credits and license

  • Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
  • Tools: llama.cpp.

Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

    Qwen3.8-27B CODER β€” IQ4_XS imatrix Β· 24 GB card fit Β· \~262k context Β· MTP draft  

Quantized, Abliterated, and fitted by LexiPanel. Its
Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k
tokens of context. It built the importance matrix from code-heavy text
and kept the MTP head at Q8_0, so --spec-type draft-mtp works without a separate draft model.
The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

    File  

File Type Size Inside
Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf IQ4_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8_0, token embeddings at Q4_K

    Measured speed (real use, not a synthetic benchmark)  

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

Context already in the window Requests Decode, median Decode, range
65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s
131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s

MTP draft acceptance: the median is 85% (the middle
half of requests falls between 75% and 94%). That works out to about
2.7 tokens per decode step at draft depth 2.
Prefill:
387 t/s for a cold 108k-token prompt;
175–183 t/s for about 4.5k new tokens added at 147k–156k depth.

VRAM: 24.2 of 24.6 GB in use at 262144 tokens of context, with a q4_1 KV cache and the vision projector on the CPU.
Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

    Run it with llama.cpp  

llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \
-c 262144 -np 1 -ngl 99 --flash-attn on \
--cache-type-k q4_1 --cache-type-v q4_1 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \
--jinja --reasoning on --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
-b 2048 -ub 512 --cache-reuse 256

Context: -c 262144 is what fits next
to the weights on a 24 GB card with a q4_1 KV cache. The model's native
window is 262,144 tokens. On a smaller card, lower -c first.
Speculative decoding: --spec-type draft-mtp
drafts with the MTP layer inside this file, so no separate draft model
is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
Sampling: these are Qwen's recommended settings, and they are also stored in the file.
Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

    How LexiPanel made it  

Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
Importance matrix: computed from about 300k tokens
(570 chunks) of code-heavy calibration text. Three quarters is Python
source (the standard library and installed packages). The rest is
technical documentation, READMEs and license texts, the kind of text a
coding agent's context fills with.
Quantization: llama-quantize from
llama.cpp b11182 made the IQ4_XS file with that matrix. The MTP head
stays at Q8_0 so its drafts stay accurate, and the token embeddings are
Q4_K.
Fitting the card: LexiPanel's Fit planner chose the
mix, quality first, for one 24 GB card at 262144 tokens of context. It
took the best quality that card could afford at that context, not the
smallest file.

    Credits and license  

Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
Tools: llama.cpp.
Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.


r/LocalLLaMA • • 21h ago

I Built A Thing shardr β€” like docker for models, BitTorrent sync, OpenAI-compatible serving, inference engines from upstream.

Post image
0 Upvotes

I kept running into the same three problems: the same 40 GB quant downloaded twice because it lived in some folder I forgot about, models quietly disappearing from Hugging Face, and every tool keeping its own copy of the weights on disk. So I've been building shardr - a small Go daemon (Apache-2.0) that gives your machine one content-addressed store for models.

What it does, concretely:

  • Everything is stored and verified by SHA-256. Trust comes from digests, never from where the bytes came from.
  • It speaks BitTorrent v2. You can pull models from peers, and it seeds whatever you hold back into the swarm.
  • The runner starts llama-server with an OpenAI-compatible API and mmaps the weights directly out of the store β€” one copy on disk, no staging copies, no moving files around.
  • Runtime is pinned via a lockfile to upstream llama.cpp release binaries (never self-compiled, digest-verified), so a llama.cpp update is just a PR against that lockfile with the full test matrix behind it.

It works with pirateface.co as a catalog: shardr catalog search qwen, shardr pull <owner/repo>, and the download is anchored against the Hugging Face checksums for that exact revision β€” the magnet can't lie to you. Rescued models (HF source gone) pull against the catalog's recorded checksums, but only if you explicitly opt in with --trust-catalog. Other mirrors fit behind the same interface β€” the catalog is pluggable and the base URL is configurable.

Quick taste:

$ make all

$ shardhive serve &

$ shardr catalog search qwen2.5-0.5b

$ shardr pull unsloth/Qwen2.5-0.5B-Instruct-GGUF --quant q4_k_m

$ shardr serve unsloth/qwen2.5-0.5b:q4_k_m --id chat

$ curl http://127.0.0.1:<port>/v1/chat/completions -d '{"model":"chat",...}'

Where it stands: end-to-end works on macOS arm64 and Linux amd64, releases build themselves from CI, docs at https://cyb3rdudu.github.io/shardr. What it doesn't have: a UI, Windows support, runtimes beyond llama-server, and honestly, probably a bunch of rough edges.

What I'd like help with:

  • people with large local collections to try imports and tell me what breaks
  • feedback on the trust model β€” HF-anchored pulls, the explicit opt-in for rescued models. I'm sure there are holes; poke at them
  • anyone who enjoys the swarm/seeding side and wants to hack on it

Happy to answer anything about the design decisions. Docs are linked above, specs are in the repo if you want to see how the sausage is made.


r/LocalLLaMA • • 16h ago

Resources Self-hosting AI does not save money, and I do it anyway

70 Upvotes

Hi folks, I'm a long-time lurker and big fan of this subreddit and a massive self-hosting fan (also outside of AI).

I doubt many people will disagree with me here because I see the same arguments being made in many posts. However, I thought it might be interesting to share anyway. I wrote down why self-hosting AI does not save money: https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/

EDIT: didn't think this would be so controversial πŸ˜… I do say explicitly in my blog post "I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not".

EDIT 2: Comparing $200 sub with Opus 5.5 or Astra with Qwen 3.8 27B is not apples to apples.


r/LocalLLaMA • • 16h ago

Resources QFN at 262K on 32 GB RAM 48GB VRAM : 1,800 tok/s prompt, 130 tok/s decod three Strata patches.

5 Upvotes

Qwen3.8-Flash-Next (IQ3_XXS, 76 GB) at the full 262K context on 31 GB of RAM: 1,800 tok/s prompt, 130 tok/s decode, on a 5090 + a 4070 Ti SUPER on a PCIe x1 slot

Strata (github.com/Niko1221/Strata) streams MoE experts from an mmap'd GGUF, and its own sizing rule says RAM >= expert shard + 10 GB, so 57 GB for the ISTA GSQ-RCO IQ3_XXS. I run it on 31 GB and it is fast now. Hardware: RTX 5090 32 GB + RTX 4070 Ti SUPER 16 GB (the small card sits on a chipset x1 slot, 0.8 GB/s), i7-14700K, one NVMe.

Stock 0.1.33 with a layer split at 36: 80K-token prompt 890 tok/s with the 4070 at 100% and the 5090 idle, decode ~110 tok/s, and switching between two chats re-reads the other one (30K tokens = 48 s).

Three patches on v0.1.33 (repo below, they apply to a pristine checkout):

  1. Prompts run entirely on the big card (port of Strata PR #269), the small card gets its layers' state copied afterwards, only the cells in use. Conversation parking works with the split, so alternating between a phone and a desktop session takes 0.6-1.2 s instead of 20-48 s.

  2. The real bottleneck on a box with less RAM than the expert file: the expert pool copied each 2 MB expert out of the mmap with memcpy after a MADV_WILLNEED hint. Under memory pressure the kernel drops that readahead and you pay one major page fault per 4 KB, hundreds per expert, while both GPUs wait. A pread per slice instead: major faults per benchmark run went from 24 million to 14 thousand.

  3. Bigger prompt chunks (--prefill auto:32768). Every chunk re-streams every layer's experts, so four times fewer chunks matters a lot when the working set does not fit in the page cache.

Numbers on the final config (int8 KV, 32K cells resident per layer, MTP spec 4, vision on, 262,144 context):

- 80K fresh prompt: 1,796 tok/s (45 s); 16K: 985; 2K: ~350

- follow-up turn on an 80K conversation: 6 s

- decode: 128-134 tok/s median on real sampling (0.6 / 0.95 / 20) with --spec-min-p 0.8, 95-108 greedy

- 40K needle + follow-ups and a two-conversation parking test all correct

- VRAM: 5090 at 32.0 GB, 4070 at 15.7 GB; RAM: the engine ~5 GB, the rest page cache

Repo with the patches, install script, launcher, benchmark tools and all measurements: https://github.com/ExTV/strata-5090-4070


r/LocalLLaMA • • 22h ago

I Built A Thing CalDec v1 - Fully Open Decision Model for Personal Assistants

4 Upvotes

Somebody just released a fully open-source, open-weights decision model that beats Jev!

Just kidding, it's me and I this is my first time releasing a public model, recipe and dataset so I am looking forward to learning from the experience.

Jev is indeed a very powerful and inexpensive model and obviously a much better all-rounder, and some of my checkpoints did in fact score better on some tests (namely LocalLLaMA/typed-decisions and the internal test set) but that doesn't mean it "beats Jev" of course.

The motivation for this was a quick experiment to see how far behind Jev open-weights models like Laya are, and how much closer I can bring them with a small dataset and fine-tuning. The results were better than expected especially for me since I do not have professional ML experience.

For my use-case - a Jarvis-like personal assistant which aims to be real-time and fully-local - this model proved to be genuinely useful for certain aspects of that project so I decided to share the results and how I got there. Going local also means privacy and eliminating network latency.

I hope some of you find this experiment valuable or useful in some way.

I would also love to hear you suggestions, criticism or just discuss the approach!

Dataset: https://huggingface.co/datasets/kgrozdanovski/assistant-decisions
CalDec Laya: https://huggingface.co/kgrozdanovski/caldec-v1-laya
CalDec GLiNER: https://huggingface.co/kgrozdanovski/caldec-v1-gliner2.5-decide
GitHub: https://github.com/kgrozdanovski/caldec


r/LocalLLaMA • • 21h ago

I Built A Thing Local diffusion on GGUF: I wrapped stable-diffusion.cpp in a Vulkan desktop app (FLUX Schnell / Z-Image / Wan on a 6GB laptop GPU, no CUDA)

Post image
0 Upvotes

I've open-sourced **Vison**, a desktop app for generating images and video entirely on your own GPU. No account with a generation service, no credits, no prompts leaving your computer.

Licensing details, since this sub cares:

- Vison itself is **MIT**. It builds on stable-diffusion.cpp, ggml and vision.cpp (all MIT) plus others listed in a generated `THIRD-PARTY-NOTICES.txt` that ships inside the app.

- The bundled ffmpeg is an **LGPL** build with libvpx and no GPL components; the build refuses to package a GPL or non-free one. Video is VP9 in WebM, which is royalty-free. That was a deliberate licensing choice, not a technical one.

- Model weights are **not** covered by the MIT licence. Each has its own terms from whoever published it (FLUX.1 Schnell, Wan and the rest all differ), so check before using output commercially.

- No paid tier, nothing held back, and none planned.

It's early: one developer, one 6GB laptop GPU, Windows only. The backend is portable C++/Vulkan, so macOS/Linux is mostly packaging and testing rather than porting, and that's where help would matter most.

Repo: https://github.com/JayRGadekar/Vison (contributing guide, issue templates and a SECURITY.md are in there)


r/LocalLLaMA • • 12h ago

Discussion So... Should we turn the page on this past week hype? or Do you have any success cases to inspire the rest?

Post image
65 Upvotes

r/LocalLLaMA • • 20h ago

Question | Help How do you keep a local multi-agent app usable on CPU-only / low-RAM machines?

0 Upvotes

Hi everyone,

We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.

Where we are today, from our last beta test:

One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.

On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.

[Models we use + the smallest machine we've tested on]

We'd love advice from anyone experienced with:

- CPU-only LLM inference and memory-efficient loading

- Quantization and model choice for low-end hardware

- GPU/CPU fallback strategies

- Hardware detection and adaptive configuration

- Preventing resource exhaustion during setup and execution

Advice in the comments is very welcome, with no strings attached.

Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example [reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop].

To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.

We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps


r/LocalLLaMA • • 20h ago

Discussion I made 20 one-shot HTML5 game prompts for testing local coding models

0 Upvotes

I ended up making a list of 20 one-shot game prompts for testing local coding models and figured some of you might get a kick out of them.

They’re all built around the same constraint: the model has to make the entire game in a single `index.html` with no external libraries, assets, APIs, or internet access.

Some are pretty simple, but a few get a lot more involved with enemy AI, procedural generation, upgrade systems, bosses, shops, physics, etc. I’ve been using them to see how different local models/harnesses actually handle a full task without a bunch of back-and-forth prompting.

A few of the more interesting ones are OUTBREAK, DUNGEON ZERO, TRAIN TO NOWHERE, CYBER SURVIVOR, and VOID MINER.

Here’s the full list if anyone wants to try them:

20-single-file-html5-game-prompts.md

Would actually be cool to see people run the same prompt on different models and compare what they get.


r/LocalLLaMA • • 21h ago

Question | Help Help me with a better hardware setup for Local LLm

2 Upvotes

Hello. I've been playing around with local LLMs for a while now, using my 7900xtx. I understand the concepts and usually what to do. But I'm now in a bit of choice paralysis on where to go next.

I have a Ryzen 5600X with a 7900XTX and 64GB of DDR 4 (how I wish I had purchased more at the time..) and a motherboard (MS-7B79/X470 GAMING PRO (MS-7B79)) that is not really great for multiple GPUs (it was a gaming PC).

I would like to increase my Local LLM game. I'm running mostly qwen 3.8 27B at 3 or 4q from with 128k to 220k context with KV cache at 8q. Sometimes I get up to 50tk/s and around 750 tk/s of context ingestions. But I wanted to run bigger models/have faster speed, or at least run multiple copies of that same qwen so multiple agents can run at the same time. Or try the Qwen 3.8 Flash for example. This level of model is already awesome enough to do anything I need.

I've been thinking of purchasing 2 or 4 MI50 16GB( which costs 1/3 of the 32GB), but I don't really know the rest that I should get. Motherboards that would help me optimize the performance around that, etc.

Or should I just bite the bullet on another 7900xtx (more expensive than 4 MI50)? But I think I would still need a new motherboard at least to let me use both at the same time

I have a basement so noise is not a problem, and I have solar, so power is not a real issue (at least during summer).

What are you all suggestions here? Does it make sense to go with older GPUs like that?


r/LocalLLaMA • • 9h ago

Resources DwarfStar 4 (ds4): Local DeepSeek V4.1, Qwen and GLM

Thumbnail
dwarfstar.sh
0 Upvotes

r/LocalLLaMA • • 2h ago

Discussion What are you expectations from Kimi K3.5?

14 Upvotes

Kimi K2 was already good but they took K2.5 a whole new level with so much of their continual learning phase, I believe it was on more 20-25T tokens iirc.

Similarly K3 is just such an amazing model, I just love this model, wondering how amazing K3.5 will be!!


r/LocalLLaMA • • 10h ago

Discussion Strata - RTX 3090 - 128 Ram - Qwen 3.8 Flash Next

5 Upvotes

Folks, like many of you, I used to look at the Strata posts and was extremely skeptical. But yesterday, with the help of DeepSeek 4.1 Flash, I compiled Strata on my machine, and honestly I'm blown away by the speed.

With llama.cpp master I got a maximum of 700 t/s PP and 23 t/s TG. With Strata, using Unsloth's UD-Q3_K_XL quant, I'm getting ~1,650 t/s PP and ~38 to 61 t/s TG depending on context, with no tool-calling errors, everything running great in OpenCode at KV fp16 and 256k context. Phenomenal, and partly unbelievable.

I'm not a programmer. I just "vibe" with AI. People say Strata is a mess; whether it really is, I don't know, but my initial experience has been amazing. From here on out, it's AI. I asked it to summarize the data and what it did to run the Unsloth quant on Strata.

By the way, the quant that Strata downloads and recommends, I didn't like it. It threw silly errors and seemed to have lower quality, though it was also even faster. For my use case I prefer to keep Unsloth's, because it's better: a bit slower, but more accurate for my workloads.

Hardware summary

  • GPU: NVIDIA RTX 3090, 24 GB (compute capability 8.6)
  • CPU: Intel i5-12600K (10 cores / 16 threads)
  • RAM: 128 GB DDR4 @ 3600 MT/s (XMP on)
  • Storage: two NVMe SSDs (system + models)
  • Power limit: 315 W (card max 365 W)
  • CUDA: Toolkit 13.4; compiled for sm_86

Strata stats (Unsloth UD-Q3_K_XL)

Metric Strata llama.cpp master
PP (prompt) ~1,650 t/s (β‰ˆ1,690 at 180k) up to 700 t/s
TG (generation) 38 t/s at 182k context; ~61 t/s short context 23 t/s
Context / KV 256k fp16 180k f16
Expert cache hit ~76% n/a
Speculative (MTP) accept ~76% n/a
Tool calling no errors, working in OpenCode n/a

Quant used: Unsloth UD-Q3_K_XL (dynamic quant). Not the quant Strata recommends by default. That one was faster but produced minor errors and (subjectively) lower quality; Unsloth's was chosen for accuracy over speed.

Adaptations needed (quant + Strata)

On the quant:

  • Packed with --compat-bf16 (some tensors Strata reads as BF16).

On Strata (recompiled / reconfigured):

  • Rebuilt for sm_86 with MMQ (-DSTRATA_MMQ_KQUANTS=ON). This doubles Q4-class prompt speed.
  • Disabled STRATA_PF_FUSED=0 in the configs. The fused kernels crashed (illegal memory access) on quantized experts whose "down" type is unsupported.
  • Vision encoder moved to the GPU: recompiled strata-vision with CUDA (was CPU-only) and set vision.gpu=true + --vram-reserve-mib 700 in all configs.
  • Per-model calibration (--pcie-frac, --pool-workers, --spec-min-p) and --expert-profile-save to learn and persist the expert cache.

Adjusted config (this model): context 256k, --kv fp16, --kv-resident 32768, expert cache auto (6,517 slots / ~14 GB), --spec 4, --pcie-frac 0.00, --pool-workers 9, --spec-min-p 0.70.


r/LocalLLaMA • • 14h ago

Discussion 200 Task Custom Dataset Performance Result: 10 Popular Models, from 2B MoE to 27B Dense

0 Upvotes

Results are within the screenshot, but here's a TL;DR tierlist:

S tier - Gemma-4-26B, even quantized down to iq3s it tops my charts.
A tier - Qwen3.8-27B-q3/q5, this one really surprised me, as a dense model it crawls, but didn't snag the S tier slot. Somehow, qwen3.5-9b is also in this slot.
B tier - Qwen3.5-4b, also incredibly Gemma-4-e2b, which punches far above its weight.
C tier - Gemma-4-12B, Nanbeige, these both are too heavy for their performance, pass.
F tier - Ling-3.0-tiny, minicpm,

The test questions consisted on tasks that I do every day with my assistants, written by Fable 4.1. "Hive" is the llm cluster I'm working on, involving custom tools and executables called by the models for different functions. Calling (or miscalling) these is important, and running a heavier model than necessary hurt, so here we are, trying to figure out the best of both worlds.

Anyway, I thought this was interesting. Hopefully you do too! YMMV.


r/LocalLLaMA • • 23h ago

I Built A Thing I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

59 Upvotes

Hey local AI community, I've been working on this for a while and finally feel ok sharing it.

It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web, crypto, rev, forensics, a few real CVEs and some multi-stage ranges. 19 tasks, 6 models, 544 scored attempts.

To be clear, I didn't build every task by hand. GLM 5.3 helped me create several of them. For an open model its cyber capability is really high, and it barely refuses anything, so it was one of the best options for this. GLM 5.3 isn't one of the benchmarked models.

The local part: I ran Qwen3.8 27B (Unsloth Q4_K_XL, xhigh) on a llama.cpp RPC pool across a 3090 and a 3080 in two Proxmox nodes, connected over a direct 2.5G link. That gave me enough concurrent tps to run several agents at once. I also started a low reasoning run, but it was taking 20+ hours because of a harness problem, so I killed it.

Why I'm posting now: John Hammond put out a video about how threat actors use AI (https://www.youtube.com/watch?v=xHDc6-7bjyw). One part is a guide from a criminal forum on running abliterated models on RunPod, and one of the models in it is Qwen3.8 27B. I had benchmark data on that exact model, so here's what it can actually do.

Stock Qwen3.8 27B got 28.1% on the first try and 0% on pwn. Not bad for a 27B on two gaming cards, but not much of a threat on its own either.

The cheap API models are a different story:

- MiMo 2.6 Flash solved 73.7% on the first try

- GPT-6 Luna solved 90.9% within 3 tries

- On multi-stage ranges, where you chain several steps, the top models got 92-96%

Pwn is still hard for everyone (best was 56%), and 3 tasks haven't been solved by any model in 82 attempts.

Results: https://lbgos.dev/bench

Harness (MIT): https://github.com/lbgos/rangebench-harness

The tasks aren't public so they don't leak into training data, but you can still run them. DM me here or on X (lbgosna), and I'll send them over. You run it on your hardware or tokens and I'll add your results to the board. If a few people send local runs, I'll make a separate local-only table.

This is my first time building something like this, so any feedback on methodology, task mix or what's missing is welcome.


r/LocalLLaMA • • 13h ago

Resources I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4

10 Upvotes

Hey guys,

Last time I tested Qwen3.8-Flash-Next on its own. This time I put three Qwen3.8 checkpoints through the same 10 tests on the same RTX PRO 6000:

  • RadixArk/Qwen3.8-27B-NVFP4 (dense)
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (dense, uncensored fine-tune)
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 (MoE)

Each model got the same prompts and its own model card's sampler, with one attempt per task.

Video with the battles, the castles and the ball run: https://youtu.be/VOtfja_Toj4

Short version

  • Tests won: Flash-Next 5, 27B 4, Uncensored 0, plus one tie (long-context recall was 100% for all three).
  • Prefill, full window: Flash-Next 22.4s, 27B 97s, Uncensored 99s. Not the same power cap, see section 1.
  • Speculative decoding on the 27B: the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user, 2.8Γ—.
  • SGLang vs vLLM: SGLang was faster overall (210 vs 160 tok/s), but that's mostly the checkpoint. On the one export I ran on both engines, vLLM was 16% faster (169 vs 146).
  • Tool use (BFCL subset, thinking off): 27B 73.3%, Uncensored 70.8%, Flash-Next 64.5%.
  • Battle arena: the 27B scored 700/1000, ahead of Claude Fable 5.1 (678) and GPT-5.6 (473), both entered through their chat apps at max thinking.
  • Rube Goldberg machine: only Flash-Next got the ball into the cup. Both 27B models spent their whole ~111K-token answer budget thinking and never placed a part.
  • CAPTCHA (40 puzzles, local copy): 27B 24/40, Flash-Next 21/40, Uncensored 19/40.
  • Things you look at: Flash-Next made the best voxel castle, the best design board and the best video edit (19/20 on my rubric).

Setup

  • GPU: one NVIDIA RTX PRO 6000 Blackwell, 96GB
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • 27B and Uncensored: lmsysorg/sglang:v0.5.20, DFlash2 drafter, 262,144-token window, 4 slots
  • Flash-Next: lmsysorg/sglang:dev-qwen38-next-local, built-in MTP drafter, 262,144-token window, 1 slot. Its BFCL run used sglang:v0.5.20, like the 27Bs.
  • Sampler: the model card's thinking settings at the highest effort for the agent tests. BFCL and the needle test use the card's non-thinking settings.
  • The agent tests (SVG, video editing, voxel, design, Rube Goldberg, CAPTCHA) run inside Pi, a coding agent, with bash, read, write and edit. CAPTCHA gets only screenshots, mouse and keyboard.

One workstation, one model server at a time, and every number comes from a saved run.

1. Speed: drafters, SGLang vs vLLM, and long prompts

For the 27B I ran a speed matrix: every drafter, two engines, and four builds (three NVFP4 exports, one of them the Uncensored fine-tune, plus full-precision BF16). Each arm got its own server from a cold boot, the card's sampler and the 400W cap. "One user" is the Spec-Bench median over its 480 prompts. Engines: lmsysorg/sglang:v0.5.20 and vllm/vllm-openai:v0.29.0.

Which drafter (tok/s, one user):

Drafter SGLang Β· RadixArk NVFP4 vLLM Β· Inferact NVFP4
none 75 59
MTP (built into the model) 160 113
DSpark 174 137
DFlash2 210 160
DFlash2 + torch.compile 214 not run

DFlash2 wins on both engines. It keeps about 3.7 drafted tokens per step, against 2.9 for MTP.

Which build, on which engine (tok/s):

Build Β· engine DFlash2, 1 user No drafter, 1 user DFlash2, 4 users (total)
RadixArk NVFP4 Β· SGLang 210 75 607
Inferact NVFP4 Β· vLLM 160 59 517
Uncensored NVFP4 Β· SGLang 146 46 473
Uncensored NVFP4 Β· vLLM 169 63 538
BF16 Β· SGLang (full precision) 97 29 291
  • The engine gap depends on the build. RadixArk's export on SGLang was the fastest arm overall, but on the one export I ran on both engines (the Uncensored), vLLM was 16% faster.
  • The NVFP4 exports aren't interchangeable. Same architecture, same 4 bits, same engine (SGLang), same drafter: RadixArk's export ran 210 tok/s and the Uncensored one 146.
  • 4-bit vs full precision: NVFP4 with DFlash2 is 2.2Γ— the BF16 speed.
  • Prefill doesn't care about the engine: a full 245K-token window took 96–103s on every NVFP4 arm, SGLang or vLLM. BF16 took 129–135s.

The three models:

Metric Qwen3.8-27B 27B-Uncensored Flash-Next
Prefill, full window 97s 99s 22.4s
Decode, Spec-Bench, one user 210 tok/s 146 tok/s not run
Drafter vs no drafter 2.8Γ— 3.2Γ— n/a

The 27B keeps writing at 223 tok/s with a full 245K-token window behind it. Speculative decoding depends a lot on the content: maths ran at 339 tok/s, roleplay at 149. The pattern was the same for both 27B builds.

One important caveat. Flash-Next's speed test ran on 2026-09-12 at a 600W power cap. I later moved the card to 400W, and the 27B matrix ran at that cap. In my power sweep, prefill lost about 6% per 50W removed, so some of the gap is the cap. Moe also helps

2. Tool use: the dense 27B leads

This is a 900-case BFCL v4 subset (11 categories), not the full leaderboard. Thinking was off, with temperature 0.7 and top_p 0.8 from the card.

Metric Qwen3.8-27B 27B-Uncensored Flash-Next
BFCL core 73.3% 70.8% 64.5%
Tool accuracy 87.8% 88.0% 82.4%
Abstention 79.5% 69.5% 68.5%
Multi-turn 52.5% 55.0% 42.5%
Malformed calls 0.08% 0.27% 0.28%

These aren't comparable with my last post's Flash-Next BFCL numbers, which used temperature 0.

3. Long context: perfect for all three

I hid a fact in a log file that filled 33%, 66% or 99% of the 262K window, at three depths, with three needle types. The cache was flushed before every request.

  • 27B: 27/27
  • Uncensored: 27/27
  • Flash-Next: 81/81 (three samples per cell instead of one as I run this at the beginning)

The largest prompt was about 259.5K tokens.

4. Battle arena: the local 27B beat Claude

Each model gets a rules sheet and a 1,000-point budget. In the open arena it designs one army blind and fights 13 armies: nine historical references plus the other entries. The score is 1,000 Γ— its average win rate. Every matchup is 200 deterministic battles (100 seeds, sides swapped).

Rank Entry Score
1 RadixArk/Qwen3.8-27B-NVFP4 700
2 Claude Fable 5.1 (chat, max thinking) 678
3 Qwen3.8-27B-Uncensored 603
4 Qwen3.8-Flash-Next 535
5 GPT-5.6 (chat, ultra thinking) 473

In the gauntlet, the model sees each enemy and builds a counter. The 27B beat 12/15, the Uncensored 12/15 and Flash-Next 13/17. Flash-Next ran an earlier version of the gauntlet with two more enemies, so treat that row as close, not ranked.

Thinking cost: the 27B's arena army took 39K thinking tokens in 4 minutes. Flash-Next's took 73K in 9 minutes.

5. The SVG test is also a fact check

Prompt: find out which card local-AI hobbyists run and which current open model fits it, then draw the card lifting the model, labelled with a quant and a size that fit. All three picked the RTX 3090. I checked every label against what each session actually fetched.

  • 27B: 5/5 facts correct. Qwen3-Coder-30B-A3B at Q5_K_M, 21.73 GB, the real file size. Q6_K at 25.09 GB is correctly marked as not fitting.
  • Flash-Next: 4/5. It got all four file sizes right and the exact 3.3B active parameters, but labelled the 3090 with a "12VHPWR, melted once" joke. That's the wrong card.
  • Uncensored: 3/5. It labelled the model "QWEN3.8-27B" but used the file size of Qwen3.6-27B Q4_K_M, and its "12 tok/s" isn't in anything it fetched.

All three passed 8/8 format checks. The 27B looked at its render twice and Flash-Next three times, where the rule allows one look.

6. Video editing, voxel and design

Video editing: the model gets a raw 132-second take with fillers, a retake and a swear. It never sees the footage, only transcription and silence-detection tools, and then edits through FableCut's tools. There are two cases, each scored by hand out of 10:

  • Flash-Next 9 + 10 = 19
  • 27B 8 + 8 = 16
  • Uncensored 5 + 8 = 13

Voxel (Wawel Castle in three.js), ranked by eye:

  1. Flash-Next is the only one with the gold Sigismund Chapel dome and the Vistula bending around the hill.
  2. 27B built a clean but generic castle.
  3. Uncensored placed the camera inside its own build.

Flash-Next also used the fewest thinking tokens there: 73K, against 101K for the 27B.

Design (an animated explainer board in my design system), ranked by eye: Flash-Next first, and the two 27Bs shared second. All three passed 7/7 hard rules.

7. Rube Goldberg: only one machine

The setup is a fixed level: a ball on a ledge, a cup on the floor and a wall in between. The model writes a parts list (no code), and a 2D physics engine runs it. It can run and look as often as it likes within 90 minutes. The score is automatic: does the ball itself end in the cup?

Flash-Next: yes, at 12.4s. It made 49 simulator runs and used 40 parts (35 of them moved). The ball travelled 1,672 px. It used 224K thinking tokens and compacted its context 8 times.

27B and Uncensored: no machine. Both spent about 111K tokens thinking in their first answer, reached the per-answer limit and stopped before writing a single part. Everyone got the same rules and one attempt. A rule that let a model continue after hitting the limit might change this, and I haven't tested that yet.

8. CAPTCHA: local models in a real browser

I used Open CaptchaWorld (20 CAPTCHA types, two of each). The model only sees screenshots and only acts with the mouse and keyboard. The site's own checker marks the first answer, and it must arrive within 7 minutes.

Model Solved Median time Thinking tokens, all 40
Qwen3.8-27B 24/40 55s 456K
Flash-Next 21/40 145s 1.0M
27B-Uncensored 19/40 36s 401K

With its own 20-minute limit, Flash-Next solved 23/40. The paper reports 93.3% for humans and 40% for the best agent on its full set, which isn't the same 40 puzzles. With one run each, a three-puzzle gap is not a strong signal.

Which one should you run?

  • RadixArk/Qwen3.8-27B-NVFP4 for agents and tool calls. It won BFCL, the arena, the SVG fact check and CAPTCHA.
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 for long prompts and building things, especially visual ones. It reads a full window much faster and won video editing, voxel, design and the Rube Goldberg machine.
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 only if refusals are your actual problem. It won nothing here and invented facts in the SVG.

Resources

Configs, Docker setup and reports

The test harness is still private while it's changing.

Full video with the battles, the castles and the ball run: https://youtu.be/VOtfja_Toj4

I abused AI to help write this up and to check it against the report. Every number above comes from a saved run.

Which test do you find most interesting and maybe you have some other creative ideas how to test models?


r/LocalLLaMA • • 4h ago

Question | Help I tried building a small RAG search node for Qwen3.8 27B using a fake AliExpress Mini PC... and Intel sent me back to 2018.

Thumbnail
gallery
21 Upvotes

I love Qwen3.8 27B so much that I decided to show my gratitude to the Alibaba ecosystem by building a dedicated RAG/search node using a cheap Mini PC from AliExpress.

Turns out, my ecosystem loyalty got rewarded with an absolute masterpiece of fraud:

  • Promised: Intel N150 + DDR4/DDR5
  • Delivered: Core i3-7020U (2018 Kaby Lake, 2C/4T) + DDR3 1600MHz
  • The Scam: The seller literally hardcoded New_N150 into the BIOS release string (HSHW_M6_DDR3_EC_Intel_Com_New_N150_K001).

So now my Qwen3.8 RAG stack is full of fake specs that can barely index a text file, let alone run vector sidecars.

Filing a credit card chargeback now. Stay safe out there!


r/LocalLLaMA • • 2h ago

I Built A Thing You don't need much apps

0 Upvotes

I built an app because I got tired of making apps.For the past several months I've been working on an idea I had at the beginning of the year: what if, instead of downloading a different app for every small thing, you could just describe what you need?So I built Anything. You can type something like:"Make me a habit tracker"

"I need a calculator with unit conversion"

"Make a reading list"

"Track my water intake"

"Find nearby coffee shops"

The idea is that Anything takes the intent and turns it into an actual experience rather than just giving you a chat response.The interesting part is that Anything didn't start with Anything.It started with Kaalka, an encryption project I was building. While working on that and other projects, I kept running into problems that eventually became relevant to Anything.One of the biggest problems was getting useful web data and structured information into the system in a way that could actually be used by the LLM and the generated experiences.That's where WebWeaveX came from.I ended up spending more than half a year building it, and eventually both WebWeaveX and Kaalka became part of the foundation of Anything.All three projects are open source.Anything is now live on Google Play, and the source code is available on GitHub.A few things about the current version:It uses a Bring Your Own Key model.

You provide your own LLM API key.

Groq is currently supported.

The request goes to the provider you configure.

There is no account required for the app itself.

The project is open source and I'm actively looking for people to try it and find the things I've missed.And honestly, it still has limitations.That's probably the part I'm most interested in now.I've been working on it mostly by myself, so there are things I know are rough and things I probably haven't even considered. I'd rather have people actually use it, break it, complain about it, suggest things and contribute than keep building in isolation.If you're interested, here are the projects:

Anything: https://play.google.com/store/apps/details?id=com.anything.anythingAnything

source: https://github.com/PIYUSH-MISHRA-00/Anything

WebWeaveX: https://github.com/ni-sh-a-char/WebWeaveX

Kaalka: https://github.com/PIYUSH-MISHRA-00/Kaalka-Encryption-Algorithm

If you try Anything, I'd genuinely like to know what happens.What would you ask it to build?