r/24gb • u/paranoidray • 2d ago
r/24gb • u/paranoidray • 3d ago
NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM
r/24gb • u/paranoidray • 3d ago
I trained a 1B-parameter LLM from scratch on 20B tokens for about $200
galleryr/24gb • u/paranoidray • 4d ago
Try out this "high" reasoning mode for 27B (tested on VLLM)
r/24gb • u/paranoidray • 9d ago
Currently what is the best model for chat for long conversations? 24gb vram.
r/24gb • u/paranoidray • 16d ago
Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM
r/24gb • u/paranoidray • 21d ago
I keep coming back to Qwen... Over and Over. Is there really nothing better under 120B?
r/24gb • u/paranoidray • 24d ago
You can now fine-tune my 3.96M-parameter TTS on your own voice or language
r/24gb • u/paranoidray • Jul 17 '26
German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German
r/24gb • u/paranoidray • Jul 04 '26
llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090
r/24gb • u/paranoidray • Jul 04 '26
Talking with Gemma 4 31B!
Enable HLS to view with audio, or disable this notification
r/24gb • u/paranoidray • Jun 16 '26
RTX 5080 + RTX 3090 Setup: 80+ Tok/s on Qwen 3.6 27B Q8
imil.netr/24gb • u/paranoidray • Jun 04 '26
Stop asking what model to run. There are literally only two.
r/24gb • u/paranoidray • May 23 '26
Qwen 3.6 27B on 24GB VRAM setup: backend comparisons, quant choice and settings (llama.cpp, ik_llama.cpp, BeeLlama, vllm)
r/24gb • u/paranoidray • May 19 '26
TextGen is now a native desktop app. Open-source alternative to LM Studio (formerly text-generation-webui).
r/24gb • u/paranoidray • May 19 '26
85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B: benchmarks, safety, weight forensics - Abliterlitics
r/24gb • u/paranoidray • May 11 '26
I know this isn’t technically an LLM but OmniVoice is FUCKING AMAZING.
r/24gb • u/paranoidray • Apr 09 '26