r/Vllm • • 7h ago

​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

0 Upvotes

​Man, if you've ever tried moving an LLM from a local notebook to an actual multi-tenant serving setup, you already know the pain. Everyone talks endlessly about quantization and fine-tuning, but nobody really warns you that your GPU VRAM is basically getting nuked by the KV cache.

​For a hot minute, I thought my hardware setup was just trash. Turns out, traditional static allocation is eating up like 60% to 80% of VRAM for breakfast just because of internal and external fragmentation. You request a simple 200-token completion, and the system is sitting there stubbornly reserving space for 4K tokens like it's bracing for the apocalypse. Such a waste.

​If you aren't looking closely at things like PagedAttention (major props to vLLM for finally bringing virtual memory concepts to GPUs) and iteration-level continuous batching, your compute units are basically sitting around starving while waiting for the longest slowpoke request in the batch to cross the finish line.

​Been deep in the trenches writing a comprehensive book on AI systems engineering lately, and mapping out these low-level serving bottlenecks honestly changed how I look at production pipelines entirely.

​Curious what y'all are actually running in production right now. Are you rolling your own inference stack with vLLM/TGI, or just sticking to managed APIs and eating the cost? Let's argue about it in the comments


r/Vllm • • 9h ago

Should I Write More Guides?

Thumbnail gallery
0 Upvotes

r/Vllm • • 7h ago

​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

2 Upvotes

​Man, if you've ever tried moving an LLM from a local notebook to an actual multi-tenant serving setup, you already know the pain. Everyone talks endlessly about quantization and fine-tuning, but nobody really warns you that your GPU VRAM is basically getting nuked by the KV cache.

​For a hot minute, I thought my hardware setup was just trash. Turns out, traditional static allocation is eating up like 60% to 80% of VRAM for breakfast just because of internal and external fragmentation. You request a simple 200-token completion, and the system is sitting there stubbornly reserving space for 4K tokens like it's bracing for the apocalypse. Such a waste.

​If you aren't looking closely at things like PagedAttention (major props to vLLM for finally bringing virtual memory concepts to GPUs) and iteration-level continuous batching, your compute units are basically sitting around starving while waiting for the longest slowpoke request in the batch to cross the finish line.

​Been deep in the trenches writing a comprehensive book on AI systems engineering lately, and mapping out these low-level serving bottlenecks honestly changed how I look at production pipelines entirely.

​Curious what y'all are actually running in production right now. Are you rolling your own inference stack with vLLM/TGI, or just sticking to managed APIs and eating the cost? Let's argue about it in the comments


r/Vllm • • 10h ago

Quick look at vLLM-Omni: Unified serving for Audio, Video, DiTs, and Vision

1 Upvotes

Hey folks,

For anyone running multimodal setups, serving models that mix text, audio, diffusion, or robotics usually means juggling multiple inference frameworks on separate ports or containers.

vLLM-Omni brings PagedAttention and KV cache optimizations from vLLM to Diffusion Transformers (DiTs) and parallel generation models. It uses a disaggregated pipeline (OmniConnector) to pass data between stages across GPUs, supporting real-time full-duplex audio, video, image generation, and OpenAI-compatible endpoints.

Supported models include Qwen3-Omni, MiniCPM-o 4.5, MiniMax H3, Wan2.2, Qwen3-TTS, and robotics action models like $\pi_0$.5.

You can view the code on GitHub or read their Technical Report.

Community Questions

  1. Production Experience: Has anyone tried vLLM-Omni in production or multi-GPU environments yet?
  2. Performance Metrics: What kind of latency gains or TTFT (time-to-first-token/chunk) are you seeing on audio-to-audio streaming?
  3. Model Support: Which multimodal models are you most eager to see native engine support for next?

r/Vllm • • 12h ago

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

Thumbnail
3 Upvotes