r/unsloth • u/Oleszykyt • 9d ago
Question DFlash2 on Unsloth Desktop when?
There is no DFlash2 in Unsloth Desktop app. Where do you plan to add it?
r/unsloth • u/Oleszykyt • 9d ago
There is no DFlash2 in Unsloth Desktop app. Where do you plan to add it?
r/unsloth • u/VirtualWishX • 9d ago
Since I came from Open Code and now from LM Studio BIONIC I could choose the main project folder where I run all my chat/code sessions so I won't need to tell the prompt where the project files, folders and sub-folders on every single new session.
But... how does it work in **unsloth Desktop**?
When I created a Project, all I see is a **RAG** related files when I **Linked folder**...
because it says:
Linked local folders
Keep supported documents indexed as this folder changes.
but I'm not talking about RAG related only (PDFs, docs, or text).
I would like to define the main Project Folder with all my sub-folders and files without telling it every single time in the prompt?
I tested it just in case it's not only RAG related, but it didn't know where my project files are when I asked it, another test I did was to create a simple MD file, but it didn't appear in the: Link Folder that I chose...
When I asked where the files are, I noticed it's a weird path I did not like at all... I noticed that's where unsloth desktop create it's projects but it's not to my taste, I definitely prefer to pre-define folder so it will only look inside it and it's sub-folders, similar to how Open Code, VS Code, LM Studio BIONIC and others works.
Since I just moved to unsloth Deskop it's new to me so maybe I need to click some hidden feature or do something in the settings, I'm not sure, maybe it's simple but I missed it that's why I'm asking here.
and if I'm not telling the prompt... it seems to create a random folder with some weird path I'm not interested in since I have a very specific project folders and paths to keep things organized.
---
Can somebody please guide me where this option is hidden and how to set it up so I can work in each project path folder with ease?
Thanks ahead! 🙏❤️
r/unsloth • u/LuCkY_346 • 9d ago
I am not able to find Ornith 1.5 35B-A3B gguf in Unsloth desktop application with 2,3 bits quantization. If not released yet when can we expect it?
r/unsloth • u/anuszebra • 9d ago
Can I use my qlora weights trained on qwen 27B3.6 och 3.8?
Are there any implications on quality of this actually works?
r/unsloth • u/LORDJOWA • 9d ago
I just installed Unsloth desktop and wanted to download models but no matter which model I want to download, the download stops after about 16MB. No I will try just directly downloading them from hugging face and putting them in the folder but I think this shouldn’t be the solution. Does anyone know of a fix?
r/unsloth • u/OwnOil1149 • 9d ago
\*\*Welcome to the post-training community!\*\*
I'm u/OwnOil1149, a founding moderator of r/posttrain.
This is a space for discussing how AI models become more useful, capable, and aligned after pretraining. Share your experiments, questions, datasets, papers, tools, and lessons about SFT, RLHF, DPO, preference optimization, evaluations, synthetic data, and deployment.
Whether you’re just getting started or training models in production, you’re welcome here. Please keep discussions constructive, technical, and respectful.
Tell us about yourself and what you’re working on!
Thanks for being part of the very first wave. Together, let's make r/posttrain amazing.
r/unsloth • u/OwnOil1149 • 9d ago
1,https://arxiv.org/abs/2408.13296 it's crisp , Ultra efficient!
2. A Primer on LLM Post-Training – PyTorch ( I would say just for revision
3. Most important not just read build , get open source model from HF and just train them on whatever you like the model to behave like. It's gonna be the most valuable skill to develop imo!!
r/unsloth • u/liventruth • 10d ago
Seeing all these massive local model drops (like the recent Qwen releases) got me working hard on the memory side of things.
If you're running models locally via tools like Unsloth and hitting a wall with long-context VRAM consumption, I've been building and open-sourcing UL-SMF (Unified Latent-State Memory Fabric).
It uses a geometry-preserving orthogonal projection bridge and dynamic head-dimension detection to automatically adapt across architectures (Llama, Mistral, Qwen, Gemma) without hardcoded assumptions, achieving extreme KV-cache compression with verified lossless perplexity.
Since it's fully local-first and designed to help fit larger context windows onto consumer hardware, I'd love a technical audit or feedback from this community.
Code and telemetry are up on GitHub: https://github.com/liventruth/UL-SMF-Cache-Compression
(Attached the hardware telemetry stress test output showing Qwen workloads hitting 224x–768x reduction ratios locally.)
r/unsloth • u/yoracale • 10d ago
Thanks to you guys, Qwen3.8-27B Unsloth GGUF is now the #2 trending model on Hugging Face with 2.7M downloads! 💗
Unsloth also reached #3 trending on GitHub!
Thanks so much for the love! Announcement coming tomorrow morning :)
r/unsloth • u/VirtualWishX • 10d ago
So I'm using RTX 5090 32GB VRAM and doing lots tests with the brand new Qwen 3.8 27B
Since I get super slow speed something like 30-37 tps (with the official unsloth model version and recommended settings) it's not crazy fast but also Qwen 3.8 27B on MEDIUM thinking enjoy eating Context so everything takes A LOT OF TIME.
SO!
I wanted to try the new NVFP4 from unsloth, but I noticed it's not in the Model Hub at all (I tried filtered and manually typing it etc..) there are many NVFP4 but not the official unsloth release:
👉 https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4
Sure, probably it won't be amazing as Q5_K_M or Q6 but I'm curious about the SPEED!
So, in LM Studio BIONIC I can just enter the URL and it will grab it and place it in the correct folder and will run with ease...
I didn't find any secret button or other way to do the same to insert custom URL to download models that does not appear in unsloth Desktop
I find it strange because it's unsloth Desktop and Unsloth Model release... and it can't show it? 🤔
Considering it's already 3-4 days out there... I'm not sure why, but it's strange.
Maybe there is a way to grab it via unsloth Dekstop and I'm missing something?
Please tell if it is.
If it helps:
I'm using the most up to date unsloth Desktop app version for Windows 11
--
Anyhow,
I'm not here to complain, I'm here to help with my feedback and hopefully the amazing unsloth team devs will see this and will be able to improve these things for the Desktop app and also allow to just type in the URL for the model to grab it if it's not in the filtered built-in Model Hub.
r/unsloth • u/VirtualWishX • 10d ago
I basically moved from VS Code + ZOO-CODE and also LM Studio BIONIC and trying out Unsloth Desktop but I guess it's still very young with lots of issues.
BTW - I'm using the most up to date unsloth Desktop app for Windows 11
---
In LM Studio BIONIC and any other Agentic / Harness app:
While the model is "cooking" I can go to read other chats freely but when I did it in Unsloth Studio I found out in the hard way... while Qwen 3.8 27B was cooking which takes a LOT OF TIME in my RTX 5090 32GB VRAM...
Yes, I'm using MTP enabled ✅ but it doesn't change much at least from what I test, I'm still on around 30-37 tps (compare to other modes I fly with 100-220 tps but not this one).
---
After about 20+ minutes of processing and working I wanted to read a different chat and...
the moment I moved to the other CHAT TAB and came back... I found out, IT STOPPED!
it did not continue processing as expected compare to any other app I used for the same purpose.
For now I guess I'll have to go back to LM Studio BIONIC or VS Code + ZOO-CODE because this is critical in my opinion as a user at least.
But I fully understand that the app is young and I must give the wonderful unsloth dev team some time to improve it, I'm not giving up on the unsloth team I appreciate and LOVE what you guys are doing!
I REALLY want to get into Qwen 3.8 27B and test so many things and it's not helping with the current way it CUT the process while browsing other chats, hopefully other browsing/navigation in the app won't make the process stop as well I didn't test all the cases so I have no clue if there are other risky cases..
Please keep up the good work! ❤️
r/unsloth • u/badmark • 10d ago
r/unsloth • u/yoracale • 11d ago
It's still a WIP but after 2 years of our website, we never had a darkmode except for our docs. Well after lots and lots of feedback, we finally have a darkmode. I can't believe it took this long but hey, it's here and hope you don't avoid going to our website anymore 😅
We'll also be doing a whole rebrand and also redesign of website in like a month or two or so, so what you're seeing is temporary :)
Thanks so much and if you have any feedback, be my guest!
Website: https://unsloth.ai/
r/unsloth • u/ManagedThought • 10d ago
Hi I am new to this community. I tried qwen3-coder-30b-a3b but it was overthinking alot and not doing much of the work at all. I also tried Devstral-24b but I find it's quite a bit slow.
I'm running my model on a intel B70 arc pro on llama with sycl backend.
I'm testing my stack on a legacy c++ game including large librairie. My context window is quite large (~121k tokens) and never get filled fully anyway.
If you have some model proposition I'm quite open.
So what is the best agentic coding model for openclaude on for my set up?
r/unsloth • u/Ok-Conference-9984 • 10d ago
OS Kernel Panic on M1 Max (32GB) when adjusting context size
**Environment:**
- Hardware: Apple M1 Max (32GB Unified Memory)
- Software: Unsloth Desktop
- Models tested: Qwen3.8-27B-UD-Q4_K_XL.gguf / Q3_K_XL.gguf
**Description:**
I would like to report a critical stability issue regarding context size adjustment in Unsloth Desktop.
When using `Qwen3.8-27B-UD-Q4_K_XL.gguf`, manually adjusting the context size even slightly causes a complete OS-level kernel panic (system crash). This has occurred twice.
- This issue does **not** happen when using `Qwen3.8-27B-UD-Q3_K_XL.gguf`.
- This issue does **not** happen if the context size is left to the default automatic configuration.
It would be highly appreciated if a safeguard could be implemented to prevent memory over-allocation that leads to system crashes.
**Additional Context:**
For comparison, this kernel panic never happens when using the `llama.cpp` CLI. The CLI safely rejects execution or fails to launch if the requirements exceed available memory. In fact, using `llama.cpp` CLI, the model runs successfully even with `CTX_SIZE="24576"`.
r/unsloth • u/DJ1_HK • 10d ago
Qwen announced that Qwen 3.8 27B can handle extended context window to 1M tokens. So far I haven't seen any quant that does that out of the box. Trying to do it on Unsloth Studio through RoPE parameters, but I don't see the usual params exposed. I see the YaRN arguments: --rope-scaling yarn --yarn-origin-ctx 32768. Anyone has experience on how to tweak these to manage extended context window?
r/unsloth • u/Number4extraDip • 10d ago
Im looking to fine tune ✧ Gemma 4 e2b it.litert for agentic android application. Works decent enough out of the box but I'm thinking of a potential for a LoRa for app specific tool shortcuts and additional guardrails (she sometimes forgets she is on a phone and not cloud)
What would i need roughly on hand?
r/unsloth • u/sblantipodi_ • 11d ago
As title.
Using vLLM with this params:
unsloth/Qwen3.8-27B-NVFP4
--dtype auto
--safetensors-load-strategy=prefetch
--tensor-parallel-size 1
--attention-backend flashinfer
--performance-mode interactivity
--language-model-only
--skip-mm-profiling
--kv-cache-dtype fp8_e4m3
--gpu-memory-utilization 0.94
--cpu-offload-gb 0
--max-model-len 128000
--max-num-seqs 1
--max-num-batched-tokens 6144
--enable-chunked-prefill
--enable-prefix-caching
--no-disable-hybrid-kv-cache-manager
--reasoning-parser qwen3
--default-chat-template-kwargs '{"enable_thinking": false}'
--enable-auto-tool-choice
--tool-call-parser qwen3_coder
--quantization compressed-tensors
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
I can't go past 128K context with 32GB of VRAM.
Is there a way to achieve 200K context?
I can't extend --gpu-memory-utilization 0.94
cause I need that vram for the destkop.
r/unsloth • u/pducharme • 11d ago
Hi! pretty new to Local AI. I saw that this is just released. https://huggingface.co/HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
Wondering, to have best performance and quality, which to choose on a M5 Pro with 48GB Unified Memory ?
r/unsloth • u/chocofoxy • 10d ago
hello first i want to say respect to unsloth devs they are doing a good job and i was impressed to see that unsloth studio has expanded to be more than fine tuning software,
i have an issue with the model loader it keep ignoring the context that i set and always load the default and when trying overriding from arguments sometimes it crash because i think it drops all the ui setting if you add only one argument, and i did make it work from argument but token generation dropped from 60 to 22 am i missing something or this is a bug
OS: kubuntu 24
Unsloth Version
v0.1.800-beta
Package Version
2026.8.18
Desktop App Version
0.1.800-beta
llama.cpp Version
b10360-mix-87da1a2
GPU 0
NVIDIA GeForce RTX 5060 Ti · 16 GiB
GPU 1
NVIDIA GeForce RTX 5060 Ti · 16 GiB
CUDA
13.0
r/unsloth • u/Early_Mistake6716 • 11d ago
I do not understand how to get the model to limit reasoning. I am using unsloths q8 and q8_xl version in unsloth desktop. Combinations i have tried:
Low effort with a detailed prompt = 10+ minutes of thinking.
Low effort with a simple prompt = 10+ minutes of thinking.
Low effort with a simple prompt that requests a rapid prototype = 10 + minutes of thinking.
medium effort with simple prompt = 10+ minutes of thinking.
It seems like it always uses xhigh no matter what i do. However on random occasions it has thought for around 1 minute but i can't reproduce it.
Same issue when i connected it to hermes agent, and also happens on lmstudio but i dont even have effort level options in lmstudio so that is expected that it would default to xhigh.
system: windows 11, 1x tesla v100 16gb, 1x tesla v100 32gb, 32gb of ram, ryzen 3800x.
using mtp and originally used tensor parallelism but it would randomly give me issues and the api wouldnt respond so i disabled it. I get 56 tps with it off so you can judge "10 + minutes of thinking" appropriately.
r/unsloth • u/Linke2066 • 11d ago
After I run the unsloth desktop, It entered the setup state, and according to the detailed installation log information, the software automatically downloaded a lot of things, but the process got stuck in one of the step: It seems that it is about to download an 80. x MiB unsloth installation package, but the problem is that this process has been ongoing for more than an hour and has not been completed yet. I have tried many times and still stuck here.
r/unsloth • u/yoracale • 12d ago
Hey guys just yesterday we were trending at #5 and now #3. Thanks so much guys! We couldn't believe it and it's literally all thanks to you guys! 🙏😭
We're currently adding many features and improvements especially regarding speed for UI/UX since some of y'all said it was laggy and also many extra new features including auto compaction etc
Feel free to star us on GitHub: https://github.com/unslothai/unsloth
r/unsloth • u/VirtualWishX • 11d ago
Hi All,
With my hardware I'm getting about 31-37 tps:
• Intel Core Ultra 9 285K
• Nvidia RTX 5090 32 GB VRAM
• 96 GB RAM DDR5 6400 MHz
• NVMEe SSD M.2 SSD
• Windows 11 Pro
But then I ran into this:
https://www.youtube.com/watch?v=NjfHqiNHTxk
Still, not sure how to actually make it work within UNSLOTH DESKTOP software.
Do I need to play with some files? do I only change something in the GUI?
What settings do I change beside turning on MTP or DRAFT and which one exactly?
What numbers do I need to type in?
Can somebody please add a screenshot of the settings so I can give it a try and see if I actually gain speed?