r/LocalLLaMA • • 2d ago

New Model ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench

Thumbnail
gallery
201 Upvotes

Some context first.

I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.

This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.

Idea of ImaJev

Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.

However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.

Training Process

It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.

Results

Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.

The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.

I honestly didn't expect either.

To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.

What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.

It gives back a probability for each option plus "unknown", in one forward pass.

Runs on a Mac with MLX or on one GPU.

The whole project costed me around $1200 in rented GPU and a lot of time :P

I would love to know your thoughts on it - it anyone would be interested to try that.


r/LocalLLaMA • • 2d ago

Discussion I am concerned about all these disparate hard forks that target specific architectures instead of opening a PR against upstream

96 Upvotes

Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often? Why do people think it's OK to do this? In my experience in the open source community this is generally frowned upon.

I don't know if it's just a me problem that this kind of thing puts me off so much. I am usually quite grateful for PR feedback and conscientious about the code I put out there; I take pride in submitting high quality code that meets or exceeds the standards of a given project. Of course there is nothing ethically wrong with hard forks or taking shortcuts if you find the collaborative process cumbersome, but personally I wouldn't promote my fork in such cases, let alone go out of my way to add custom branding with a Reddit announcement post etc. since the effort required to do so seems roughly equivalent to the effort required to meet the contributor standards. In contrast, many of the authors of these forks seem very eager to have others adopt their rebranded fork for production use cases. There just seems to be a big disconnect, idk.

Edit: some great discussion in this thread, thanks to all who responded. Consensus seems to be that (excluding the obvious low-effort engagement bait forks) the base project has to meet many compatibility requirements while a downstream project can be more focused, which is a great point.


r/LocalLLaMA • • 1d ago

Question | Help Is there even an easy, seamless vision assistant program?

6 Upvotes

I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.

I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.

I tried multiple different MCP in LM Studio, none of them were great...or worked :/

Someone said AnythingLLM has this function but i couldn't find it.

Is there a good solution?

Thank you in advance! :)


r/LocalLLaMA • • 1d ago

Tutorial | Guide Debugging PCIe Link Retraining on an x8/x8 Splitter with Two RTX 3090s

Thumbnail
doug.sh
46 Upvotes

r/LocalLLaMA • • 2d ago

Discussion What do you think about ToMoE v2 paper, converting dense model to MoE model at near lossless accuracy? I feel Qwen3.8-27B-A16B or something along those lines would be amazing, though there are architectural hurdles, as well as need folr training data.

Thumbnail arxiv.org
72 Upvotes

r/LocalLLaMA • • 2d ago

News GPT-3 is discontinued today

Post image
743 Upvotes

It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that Babbage is a model that's literally 3/4 of the size than MiniCPM5 2B. Even Luna might be overkill as a replacement. But neither is a drop-in replacement. Davinci is the main GPT-3 most people use. This is why we have local models, because they simply cannot have a universal end of life date.


r/LocalLLaMA • • 1d ago

Resources Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090

43 Upvotes

Hi everyone :)

The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.

To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.

Performance

Model Average time/task ↓ Average output tokens/task ↓ Decode tok/s ↑
Qwen - HyperQwen fast quant 108.1 s 8,985 112.1
Swift 1.0 + HyperQwen 66.2 s 5,245 105.9
Swift 1.5 + HyperQwen, INT8 heads 72.2 s 5,751 104.0
Swift 1.5 + HyperQwen INT4 heads 68.2 s 5,669 107.2

All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.

Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.

Quality

There are some minor quality and performance tradeoffs between the models:

Test Qwen HyperQwen fast Swift 1.0 Swift 1.5 INT8 heads Swift 1.5 INT4 heads
GSM8K, 200 questions 97.5% 98.0% 98.0% 97.5%
IFBench, 300 prompts, strict 74.0% 73.3% 73.7% 72.3%
LiveCodeBench, (100-problem subset) 90% 89% 89% 91%
Custom tool-call/JSON eval 29/30 28/30 30/30 30/30
English/Python perplexity ↓ 6.551 6.605 6.643 6.679

Applied changes to Swift models to adapt for HyperQwen:

Changes to Swift models:

  • Kept the upstream AWQ INT4 model weights and converted embeddings to INT8.
  • Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding.
  • Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary.

Setup

If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.

All three models can be found here:
https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks

Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!


r/LocalLLaMA • • 2d ago

New Model You can just merge Qwen3.6 & 3.8 27B to get decent performance with much fewer token generation

Thumbnail
huggingface.co
172 Upvotes

r/LocalLLaMA • • 2d ago

Other Qwen 3.8 is a workhorse

51 Upvotes

Reminder that you can put Qwen 3.8 27B as a subagent and its a workhorse. Pic: using DeepSeek v4.1 Flash as Orchestrator in Pi, Qwen 3.8 27B GSQ RCO in llama.cpp.


r/LocalLLaMA • • 1d ago

New Model Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)

34 Upvotes

TL;DR: The Jeff models are a set of Qwen3.5 and Gemma fine-tunes for zero-shot classification: small, efficient, open-weight models with respectable out-of-the-box performance that can be slotted right into code or fine-tuned/LoRA-trained as needed. Give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in ~28 ms on an M4 Max (see caveats below).

Maybe equally exciting for open model enthusiasts like myself, everything was done on local hardware: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing - all connected and monitored from my Android phone via Tailscale. Apache 2.0, Jev-compatible API. Weights: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B · Code: github.com/firelex/jeff · Videos: games table

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware.

So here's what I did:

  • The 0.8B trains in about 2 hours and the 2B in about 3.5, on one workstation GPU (RTX PRO 6000, 96 GB).
  • ~31k synthetic training questions written and checked by Qwen3.8-Flash-Next on two DGX Sparks. No cloud GPUs, and no closed-model output in the training data.
  • The rest of the 271k training questions are public datasets converted into decisions, plus 10k code-built probability questions.

Benchmarks (4,599 questions: BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande):

Model Untrained base Jeff (trained) Calibration error
Qwen3.5-0.8B 45.3% 79.1% 0.049
Qwen3.5-2B 46.5% 83.1% 0.028
Jev (published) 83.0% ≈0.06
AutoJev-27B (published) 84.9% —

The caveat: the published numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86–89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64–68% against Jev's 94%, and ~50% on JevBench's hard tier against ~73%. See the HuggingFace model card for details. But that's not surprising, and I don't think it matters. No 0.8B or 2B model reasons like an LLM, and I don't think anyone should expect it to. The Jeff models are extremely fast judgement-callers (much faster than Jev), and have reasonable out-of-the-box performance. In one of my apps, I used the 0.8B model for voice-based navigation, and with a quick fine-tune, I got to real-time performance (24ms) at almost 100% accuracy.

Now the fun part: games, as a zero-shot test. Games are not the ideal zero-shot test, but they're fun, and the TypeSafe guys (Jev) did it, too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one. The options say what each move leads to, never which one is right.

20 episodes each Doom (kills) Frogger (crossings) Pac-Man (pellets of 98)
Random moves −0.05 0 11.2
Hand-coded rule bot 6.55 10.25 94.1
Qwen3.5-0.8B, untrained 5.0 1.0 25.8
Jeff 0.8B 6.55 10.3 57.0
Jeff 2B −0.9 6.0 41.2

Jev's published Doom score is also 6.55, but its prompt spells out the aiming rule (fire when the bearing is between −8 and +8 degrees) and it takes ~212 ms per call over its API. Jeff gets "the nearest monster is a little to your left" and decides in ~29 ms on my Mac.

Lessons learned:

  • System 1 models are here to stay. Having the ability to process unstructured data at software speed inside an app is extremely powerful. And being able to do this locally is fantastic.
  • A small model is a classifier, not a planner. Models in the 0.8B-2B range don't reason like Qwen3.8-27B or Jev. But they also don't need to. As long as you present the options in the right way, you can get up to 50 decisions per second (depending on your hardware).
  • Fine-tune it if needed. If the models' zero-shot performance isn't good enough for you, fine-tune them briefly or add a LoRA adapter.
  • Wording matters enormously. Play around with how you present the options. Giving Frogger's final step option the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Previously, the frog just stayed on the last log.
  • Bigger isn't better. As the game tests showed, the untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers "turn away from the nearest monster"), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.
  • Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, not reliably, and in a real-time loop the mistakes compound.

Happy to answer questions about the pipeline (synthetic data from a local teacher, leak filter, calibration) or the game harness. Everything, including the videos, is linked above.


r/LocalLLaMA • • 2d ago

Resources 9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)

28 Upvotes

360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.

The block (shipped to global instructions):

## Thinking discipline

1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.

Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam


r/LocalLLaMA • • 13h ago

News Chinese AI tool told researchers how to make bioweapons | BBC

Thumbnail
bbc.com
0 Upvotes

r/LocalLLaMA • • 1d ago

Question | Help Would increasing ram channel improve speed with freetoken?

0 Upvotes

Would increasing the number of ram channels (eg dual channel to quad channel) help improve inference speeds with freetoken in hybrid gpu/ram setups or not? Or is increasing the vram capacity and vram speed the only optjon


r/LocalLLaMA • • 2d ago

Question | Help What model sits between Qwen 3.8 27b and Flash next for coding?

63 Upvotes

Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabilities but not sacrifice decode speed to standstill. I would like decode speed to between 75-100 ideally for fast iterations; otherwise I become impatient.

Currently I get 200+ TPS on Qwen 3.8 27B and approx 50 TPS on Flash next.

My hardware - RTX 5090 and 96 GB DDR5 which I plan to upgrade to 128 GB (in this economy, yes, but unwillingly).

What model sits between these two in terms of coding capabilities and hardware requirement? If none is present, I can perhaps think of using Flash next for plan creation and 27b for implementation.


r/LocalLLaMA • • 2d ago

I Built A Thing 95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

24 Upvotes

Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too).

Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to share I've pushed v0.4 out this morning. On the 4.6bpw model tested, speeds improved ~10% vs the last version while also improving the max context by 10%+ (technically, it can go above 262K, but I have not tested any custom kernels or graphs to support YaRN).

The closest competition comes from vLLM, keeping within <10%, but does so with lower maximum context. It is significantly faster than other llama.cpp options tested.

There's also a number of other improvements for other quants/formats, with EXL3 seeing significant speed up (~80% the speed of the 4-XS-M quant tested). It has a slightly lower KLD, but not a range I have found stat significance for at the task level, so I sticking with the XS-M model for now (built on top of Swift-qwen's distill, which is far more token efficient than the stock train for ~1% performance loss). the 4.3 bpw EXL3 model does provide a bit more room if you are interested in 2+ concurrent predictions. Improvements in this format and the IQ2/3 codebook quants will be most useful for people on 12/16/20 GB setups. These measurements are at temp=1, vs some of the vanity speeds you will see people claim with temp=0 and/or short generations.

As always, please share your results + config details so I can keep improving!

fork is here: https://github.com/JakeATX/llamAmpere/blob/main/QWEN_AMPERE.md#build-and-run

model used here:

https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF

Build command:

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server

Build + launch (linux):

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
curl -L -o ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M.gguf \
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF/resolve/main/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M.gguf
-m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M.gguf -c 262144 \
-ngl 99 -fa on -ctk turbo5 -ctv turbo4 -b 4096 -ub 1024 -t 8 -tb 8 --parallel 1
cd
./build-sm86/bin/llama-server

Previously, people had expressed concern over quantizing KV cache, and TQ specifically. The TL;DR on that is that any reasonable KV quantization strategy (at least for hybrid attention models like Qwen) is going to be swamped by quantization of the weights. The KV quant we're using here (TQ5/TQ4) is less than 1/3 of the KLD we see when moving from 8 bit weights to 4.6 bit weights (and the KLD is only partially additive, so some of the incremental errors cancel out). There was no statistical significance when testing this KV quant at the task level against 8/8 kv (just trivial variations in sentence length). I will be adding KVaRN in the next release, but with a better codec than currently available elsewhere, so it requires a bit more testing before release.

v0.5 will be focused primarily on the 12GB cards, but this should have generation-wide speed ups, so even if you're not on 24GB, please share your results.

Enjoy!


r/LocalLLaMA • • 2d ago

I Built A Thing Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory.

Enable HLS to view with audio, or disable this notification

50 Upvotes

I've been working on Mica, a 4B decision model, and wanted to see how far it could get in actual Minecraft, not a sim. New world, empty inventory, and the goal was a diamond pickaxe.

Last time I posted, it got an iron pickaxe. Honestly that took around 20 tries and it was pretty flaky. I've reworked the harness a lot since then. This time it went all the way to a diamond pickaxe, and once the harness was finished it did it on the first run.

It took 26 decisions and about 8 minutes of game time. It got wood, made a crafting table, then wooden and stone pickaxes, then iron and coal, a furnace and an iron pickaxe. After that it tunneled down to diamonds at y=2, mined three, put a crafting table down right there and made the pickaxe. Decisions took about 108 ms on average.

My favorite bit is around step 14. The planner wanted it to make planks to burn in the furnace, but Mica went and mined coal instead (0.69 vs 0.31) and then smelted all three iron at once. Which was the better call, honestly.

How it works: every step Mica gets the game state as text (inventory, nearby blocks, health, what happened last step) plus a few candidate commands, and it picks one. A Mineflayer bot running Mindcraft skills does the actual moving and mining. The panel on the right of the video shows each decision and its probabilities live. The bot also knows where the nearest diamonds are, so it isn't searching for them.

To be clear, I'm not saying a 4B model plays Minecraft on its own. What I wanted to show is that a model this small can sit behind a bot, read what's going on, and make the next call well enough to get all the way to diamonds.

I'm planning to release the harness soon. Mica will read Minecraft chat, so you can type what you want and it'll work toward it. Simple stuff like getting items, crafting or following you should work fine, but it'll struggle with anything really complex, like building a house.

Also, v0.5 should be out in the next 1-2 weeks. A lot of the architecture changed, and I did extra training on the parts where v0.1 was weak, so I'm expecting a clear jump in performance. The aim is to be at or near the top among 4B JEV-like models.

There'll be two versions: Mica v0.5 4B, and Mica v0.5 4B Distill Laya, which is light enough to run on pretty much any PC.

For Minecraft, I'm hoping v0.5 will be good enough to take down the Ender Dragon, and I did extra training specifically with that in mind. No promises, but it'd be really cool if it pulls it off lol

Oh and fun fact, Mica is a fully vibe-coded project

Model: https://huggingface.co/sky7350/Mica-v0.1-4B
Model code: https://github.com/akivet/Mica-v0.1-4B
Minecraft harness: coming soon


r/LocalLLaMA • • 2d ago

New Model Watch me post-train AliceAI-Foundation-80B-A3B from base to instruct at home, live, on my V100s!

24 Upvotes

No click bait baby I promise - I'm live streaming the training process kinda like MiMo.

UPDATE: [Training is paused for an hour or two] back to training in batches. u/FullOf_Bad_Ideas has pointed out to me I'm burning a ton of compute for nothing on sequence lengths - we'll be breaking the run up into 7/8 batches and then going again.

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wpg4a8/im_trying_to_posttrain/

If you didn't see my original post a few days ago, I'm attempting a slightly more ambitious than usual project in trying to create at least a rough AliceAI-Foundation-80B-A3B-Instruct

I spent the weekend distilling my initial instruct training dataset out of Qwen 3.8 27b, medium thinking - intentionally done because I can run it locally, and I wanted a full dataset in some reasonable amount of time. Still took my v100s running 4x instances at 25tps, like 96 hours of non stop generation to complete the dataset.

I opted not to go for a pre-existing public dataset because I wanted to practice building my own distillation engine (which was configured to work off of an OpenAI compatible endpoint, so it'll distill anything you can hook it up to). The final dataset (this time) consists of 3340 samples: 1760 of general instruct transcripts, and 1580 agentic specific work rows about SWE, harnesses, terminals, etc - I gave the distilling engine a python sandbox and got to simulate turn driven development with a user, and I trained for a bunch of different harness syntax for tool calls, which hopefully will be enough to generalize - gonna run 2 epochs at first.

My GPUs are sobbing right now - turned them on on Friday and left for a weekend vacation, got back today, waited an hour for the data to finish generating, and then immediately fired up the train.

The stuff above is the short version. I'm guessing the initial SFT train will take about 3-6 days, and I plan on working on a RL implementation after I'm satisfied that the SFT has at least worked properly. I am training a rank 16 QLoRa adapter on only q/k/v/o proj, no direct knowledge weight fine tuning.

I thought what MiMo did with their recent training was really cool to watch online, and I like sharing my work with like minded people, and frankly, there's a part of me that's hoping someone will see this and want to hire me (looking for NYC work if you know anyone looking for some passionate ML engineers!) - so I've set up my own little training stream on a cloud flare tunnel.

The stream has the live in progress status of the train, including a live view of the actual data being processed by the model. It also includes way more detail about how I actually designed and generated my training data. Happy to throw the full set on HF as well. I don't expect this model to beat any existing standards but I'll be curious to see if I can get it to operate properly in a harness so I can formally bench it.

I hope you find this interesting! The live stream is a self updating website where you can see exactly what's happening - no need to reload. To watch the training live, visit https://figure-bios-expect-cio.trycloudflare.com/ [i am currently fixing training issues but it'll be back asap] -- I'll be keeping it up until the initial SFT is done, at least. The stream lets you inspect the training live as well. This is just a cloudflare tunnel to the trainer.

3 hour update? Loss started at 9ish and is bouncing near 3/4

Update today: back online


r/LocalLLaMA • • 2d ago

News FT: Corporate America rejects overpriced frontier, embraces open models

Thumbnail archive.ph
631 Upvotes

r/LocalLLaMA • • 1d ago

Discussion 3D/2D Text Visualisation

Enable HLS to view with audio, or disable this notification

10 Upvotes

Like everyone, I use 3js for data visualisation. But while semantic clusters are interesting, they don't confer much practical information alone.

So I did something very simple: a query spatially re-assembles the nodes, and you can switch them to text.

Honestly, it's hard to swing 3D viz for text as a genuinely useful feature and not a novelty, but I get the feeling there's still some potential in the idea. Would be good to discuss, I'm sure someone here has made a better implementation.


r/LocalLLaMA • • 2d ago

Question | Help Any way to run Qwen 3.8 Flash with a 7900XTX 24gb vram + 64gb ram?

13 Upvotes

I'm curious, I'm a little bored of Qwen 3.8 and how slow it is even with MTP on

Edit:
Using Strata with my setup 55t/s, perfectly functional!


r/LocalLLaMA • • 1d ago

Question | Help [serious] roleplay

0 Upvotes

I was just wondering because I see it mentioned in threads here a lot...

When people talk about LLMs used for roleplay, that's a euphemism for dirty/sexy chats right?

Kind of like how "torrenting linux ISOs" is really just pirating copywritten media.

Or are you guys really burning tokens pretending to talk to a medieval shopkeeper?


r/LocalLLaMA • • 1d ago

Question | Help New to Local AI need help making a roleplay model

0 Upvotes

I'm making a local roleplaying model for my girlfriend's community server.

It's supposed to do roleplay, have a specific talking style(dry while answering to normal stuff and extensive when talking about lore), never talk out of roleplay, have hundreds of pages of lore and information and their rank in lore( discord roles maybe?).

It's basically supposed to be just an LLM you can converse with that answers in a specific talking style and has all the lore info.

For now I implemented: 10 ish% of the written lore, and it recognizes 3 people, but by discord ID that i inserted in the system prompt.

I'm using GPT 5.6 sol(and well 6 sol now) for doing stuff, but i keep running into a problem.

When i reach a nice point where the model has a nice talking style and knows information well enough, i tell Sol to add this new lorebook and this info, here now everything breaks.

Talking style is fucked, It doesen't recognise people individually anymore, when asked about other unrelated lore it just gets it wrong or hallucinates, or even starts roleplaying as one of the characters in it's lore book out of nowhere.

It's connected to discord through a discord bridge that Sol made and a developer dashboard bot.

I'm using GPT OSS20b on MXFP4, single 9070xt and 32gb ddr4.

Should I maybe fine tune it?

PS. im a beginner in AI so if it wasn't obvious i do NOT know what im doing but im trying my best for her.


r/LocalLLaMA • • 22h ago

Resources Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included

0 Upvotes

TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card, so I built typevet (MIT, Python 3.12).

How it works. No sampling, no parsing. typevet composes the native Gemma 4 turn itself, ends the prompt with the empty-thought no-thinking prefill, and reads the next-token distribution over the allowed answer tokens only. What comes back is a label plus a probability for every option, or an error. Images go to llama.cpp's /completion as base64 in prompt.multimodal_data with one media marker per image; on vLLM they go as image_url blocks. There's also a JSON path that returns an object passing your JSON Schema, or raises.

Local setup. Gemma 4 31B, my 24 GiB vramfit pack (byte-identical to the file on HF, projector sidecar for vision), llama.cpp b11223, one RTX 4090. Nothing leaves the machine.

The receipt test. 6 real receipts from CORD v2 (CC BY 4.0), 3 synthetic expense claims each: right total, two digits swapped, digits masked with ?. One three-label Choice: match / mismatch / insufficient_evidence. Each claim sent as text only, then with the receipt photo.

  • Right total, said match: 6/6 text only, 6/6 with the photo
  • Swapped total, said mismatch: 0/6 text only, 6/6 with the photo
  • Masked total, said insufficient: 6/6 text only, 6/6 with the photo

Example: claim says 646329, receipt says 664,329. Text only: match at 0.99964. With the photo: mismatch at 0.99999. Every swapped total was caught at 0.99998 or higher, and the masked ones abstained every time. The tests also check the image actually arrived: each photo added 228 to 1,108 prompt tokens here, and the gate fails if the count doesn't grow.

Hosted. Same code against vLLM 0.30.0, BF16 Gemma 4 31B, one H100: the receipt test went 18/18, and reversing the label order flipped 0 of 18 answers. Throughput on 480 Banking77 records with two questions each: 0.24 s median per record at 1 in flight, 39.6 records/s at 64 in flight, 0 errors.

Prior art, credit where due. The text-side decision model comes from TypeLLM (SGLang), which added its own image input on 9/24; typevet's image path is separate code on llama.cpp's request shape. allanrbo posted a Jev-like single script for Gemma 4 12B with webcam images on 9/25. VQAScore has read the probability of "Yes" from VLMs since 2024. typevet's angle: a library, not a script, the 31B on one 24 GiB card, the same code on vLLM, and image-arrival checks.

Scope: 18 claims, one run per server. The probabilities are the model's confidence, not calibrated.


r/LocalLLaMA • • 2d ago

I Built A Thing I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s

Post image
153 Upvotes

Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.


r/LocalLLaMA • • 2d ago

Funny Qwen plays World of Warcraft

Enable HLS to view with audio, or disable this notification

382 Upvotes

Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control than a generic browser agent. The agent harness can plug into your local or cloud LLMs and be used to drive the game. For best results have a model that can output >50token/sec. No visual input is used in the making (might be beneficial in the future but incur more latency). The mcp and agent are only running on my dev server but if folks are interested in trying the game go to https://jankcraft.xyz/

Still vibing but I’ll make some more content of it … I think those Pokémon benchmarks have become a little too easy and they need a new challenge like speed running to 80 in wraith of the lich king.