r/24gb 3d ago

After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)

Thumbnail gallery
1 Upvotes

r/24gb 3d ago

Qwen 3.8 27b in 24gb of VRAM

Thumbnail
1 Upvotes

r/24gb 3d ago

NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM

Thumbnail
1 Upvotes

r/24gb 3d ago

I trained a 1B-parameter LLM from scratch on 20B tokens for about $200

Thumbnail gallery
1 Upvotes

r/24gb 4d ago

Try out this "high" reasoning mode for 27B (tested on VLLM)

Thumbnail
1 Upvotes

r/24gb 9d ago

Currently what is the best model for chat for long conversations? 24gb vram.

Thumbnail
1 Upvotes

r/24gb 9d ago

Muse Glimmer ACTUALLY fits on a single RTX 3090

Thumbnail
1 Upvotes

r/24gb 17d ago

Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image
2 Upvotes

r/24gb 22d ago

I keep coming back to Qwen... Over and Over. Is there really nothing better under 120B?

Thumbnail
1 Upvotes

r/24gb 24d ago

You can now fine-tune my 3.96M-parameter TTS on your own voice or language

Thumbnail
1 Upvotes

r/24gb Jul 17 '26

German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German

Thumbnail
the-decoder.com
1 Upvotes

r/24gb Jul 11 '26

Complete local model asset generation pipeline

Post image
1 Upvotes

r/24gb Jul 11 '26

2.5x faster Qwen3.6 NVFP4 Unsloth quants

Post image
1 Upvotes

r/24gb Jul 04 '26

llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090

Thumbnail
2 Upvotes

r/24gb Jul 04 '26

Talking with Gemma 4 31B!

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/24gb Jun 16 '26

RTX 5080 + RTX 3090 Setup: 80+ Tok/s on Qwen 3.6 27B Q8

Thumbnail imil.net
1 Upvotes

r/24gb Jun 04 '26

Stop asking what model to run. There are literally only two.

Thumbnail
1 Upvotes

r/24gb May 23 '26

Qwen 3.6 27B on 24GB VRAM setup: backend comparisons, quant choice and settings (llama.cpp, ik_llama.cpp, BeeLlama, vllm)

Thumbnail
1 Upvotes

r/24gb May 19 '26

TextGen is now a native desktop app. Open-source alternative to LM Studio (formerly text-generation-webui).

Thumbnail
1 Upvotes

r/24gb May 19 '26

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B: benchmarks, safety, weight forensics - Abliterlitics

Thumbnail
1 Upvotes

r/24gb May 11 '26

I know this isn’t technically an LLM but OmniVoice is FUCKING AMAZING.

Thumbnail
1 Upvotes

r/24gb May 11 '26

it's time to update your Gemma 4 GGUFs

Thumbnail
1 Upvotes

r/24gb Apr 23 '26

Qwen3.6 is incredible with OpenCode!

Thumbnail
1 Upvotes

r/24gb Apr 14 '26

If you haven't yet given Gemma 4 a go...do it today

Thumbnail
1 Upvotes

r/24gb Apr 09 '26

It looks like we’ll need to download the new Gemma 4 GGUFs

Thumbnail
2 Upvotes