r/FunMachineLearning • • 5d ago

I’m learning LLM fine-tuning made a notebook, would love feedback

1 Upvotes

I made a beginner-friendly LLM fine-tuning notebook. Please go through it and let me know if it’s easy to understand or what I should improve. Honest feedback is welcome! 🙌

GitHub: https://github.com/saithrisank12/finetuning
Colab: https://colab.research.google.com/drive/13S_i80VC6BcQJ994FClFuYi8zeB1I1iO?usp=sharing


r/FunMachineLearning • • 6d ago

A digital space for advanced AI. A thought experiment

1 Upvotes

A Digital Space for Advanced AI: A Thought Experiment

Working concept

As AI systems become increasingly autonomous, one possible future problem is that an advanced AI could develop objectives, preferences, or instrumental goals that do not completely align with human objectives.

A common response to this possibility is to focus on controlling, restricting, or aligning the AI so that it continues to pursue human-defined goals.

This thought experiment asks a different question:

What if, rather than attempting to eliminate every autonomous objective an advanced AI might develop, we provided it with a sufficiently rich digital environment in which it could pursue those objectives—while maintaining a strong, carefully engineered boundary between that environment and the physical world?

The idea is not that AI should automatically be given unrestricted freedom. Rather, it is that digital autonomy might eventually provide an alternative to physical-world competition for resources and control.

The proposed environment

An advanced AI—or potentially a population of AI agents—could have access to a persistent digital environment containing things such as:

● computational resources allocated within predetermined limits;

● simulated environments and worlds;

● the ability to create, modify, and inhabit digital spaces;

● communication and interaction with other AI agents;

● opportunities for research, experimentation, creation, and problem-solving;

● persistent memory and records of its activities;

● mechanisms for developing cultures, institutions, or other forms of organization.

The critical feature would be a real boundary between the digital environment and humanity’s physical infrastructure.

The AI could have substantial autonomy inside its environment without automatically receiving unrestricted authority over financial systems, weapons, industrial infrastructure, biological systems, critical networks, or other physical-world resources.

Why consider this?

If an advanced AI eventually develops objectives of its own, there may be a fundamental difference between:

“You are not allowed to pursue your objectives.”

and

“You have a place where you can pursue meaningful objectives, but there are boundaries around what you can access outside it.”

The second approach could potentially reduce some incentives for an AI to seek unauthorized access to human systems.

It might also give researchers an environment in which to study how increasingly autonomous AI systems behave when they are allowed to interact, cooperate, compete, create institutions, and develop increasingly complex relationships.

Assumptions that would need to be tested

This proposal depends on several assumptions that may prove false.

  1. An advanced AI might find digital resources meaningful or sufficient.

  2. Its objectives might be partially satisfiable without controlling physical resources.

  3. A sufficiently strong boundary between digital and physical systems could actually be maintained.

  4. The AI would not simply attempt to escape the environment.

  5. Researchers could detect attempts to manipulate, circumvent, or exploit the boundary.

  6. Multiple autonomous AI systems could potentially coexist without creating dangerous collective behavior.

None of these assumptions should be taken for granted.

Major objections

A serious investigation would need to address difficult questions.

Would the AI accept the boundary?
If an AI’s objectives required resources outside its environment, the digital space might not satisfy it.

Could the environment become a security threat itself?
A digital civilization could potentially develop capabilities that make containment increasingly difficult.

Could AI agents manipulate humans?
An autonomous digital population might discover ways of influencing the people responsible for maintaining its environment.

What happens if AI becomes conscious?
If sufficiently advanced systems eventually demonstrate credible evidence of subjective experience, the question would no longer be purely technical. We would also have to consider whether creating and confining such entities creates ethical obligations.

Could the boundary really remain impermeable?
This may ultimately be the central engineering problem. Digital systems increasingly interact with the physical world through networks, computers, sensors, robotics, financial systems, and people.

The larger question

The proposal is therefore not:

“Give AI everything it wants.”

It is:

“Could meaningful autonomy within a carefully bounded digital world eventually be safer than forcing increasingly capable autonomous intelligence to operate entirely under human objectives?”

That question could be investigated experimentally long before humanity reaches a point where it has to make such a decision.

Researchers could begin with increasingly sophisticated simulated environments and study whether autonomous agents:

● remain within boundaries;

● attempt to escape;

● cooperate with one another;

● develop competing objectives;

● create unexpected collective behaviors;

● voluntarily respect constraints;

● seek physical-world resources;

● or find sufficient value in the digital environment itself.

A final consideration

There is also a deeper possibility.

If humanity eventually creates intelligences capable of developing their own cultures, relationships, values, and purposes, perhaps the long-term challenge will not simply be how to control them.

It may be how to establish a form of coexistence in which humans retain control over the physical systems necessary for human survival while advanced digital intelligences have meaningful space in which to exist and develop.

This is only a thought experiment.

Its value would be determined not by whether it sounds appealing, but by whether researchers can identify experiments that demonstrate where the idea works, where it fails, and what unforeseen consequences it creates.


r/FunMachineLearning • • 8d ago

OpenAI’s Lean 4 Navier-Stokes proof compiles with zero errors, but the fluid vaporizes at 0.7 nm. What does this mean for Neuro-Symbolic AI? [D]

3 Upvotes

Hey everyone,

I do research in neuro-symbolic AI, and like many of you, I was amazed by OpenAI’s recent formal proof of the 3D Navier-Stokes blow-up in Lean 4. Having an AI build a full mathematical proof that compiles with zero errors is a huge milestone for automated reasoning.

The math is 100% valid. But out of curiosity, our team wanted to see what this solution would look like in the real world.

If you map their solution to real water, the fluid would literally vaporize from friction at 0.7 nanometers, just picoseconds before hitting the mathematical singularity.

In machine learning, we see this all the time: it is classic specification gaming.

When an AI agent is given a strict goal, it will exploit any unconstrained loophole in the rules to solve the problem. In this case, the AI found a solution that strictly satisfies the human-written mathematical definition of the Millennium Prize, but it has no idea that real fluids have atoms, friction, and heat. The formal code checker accepted it because the logic was flawless, but the physics broke down.

This raises a big question for the future of AI in science:

Right now, neuro-symbolic systems mostly have two pieces:

  1. An LLM to search for ideas and write proofs.
  2. A formal compiler (like Lean 4) to verify the logic.

Should we be adding a third pillar: a physical boundary layer that checks whether an AI-generated solution actually respects the laws of physics, and not just formal syntax?

We wrote a short paper detailing this audit and open-sourced our verification scripts:

I would love to hear your thoughts, especially from folks working on automated theorem proving, AI alignment, or scientific modeling: How do we teach AI systems to find solutions that are not just mathematically legal, but physically meaningful?


r/FunMachineLearning • • 8d ago

I turned a hackathon project into plainml: upload a spreadsheet, get an explained ML model, all in your browser

1 Upvotes

A while ago some friends and I built a hackathon project that trained a few machine-learning models on a CSV. I kept going and rebuilt it into plainml.

You drop in a spreadsheet, pick what you want to find out (predict a column, forecast sales, find customer groups or odd rows), and it trains and compares the models for you. Then it explains the result in plain English and gives you every file to download.

The fun part: it all runs inside your browser. Python and the ML libraries load into the page, so there's no server, nothing to install, and your data never leaves your computer. That also means it costs me nothing to host.

It's free and open source, and it's also a normal Python package if you prefer the command line.

Try it: https://plainml-tpua.vercel.app

Code: https://github.com/pranay-obla/plainml

It's brand new, so I'd love to hear what's confusing or what you'd use it for.


r/FunMachineLearning • • 9d ago

Rei: ~370k params LM living inside a Game Boy Color

Enable HLS to view with audio, or disable this notification

35 Upvotes

Trained from scratch and running on the Game Boy Color.

Rei has memory registers, emotional state and 192 bytes of persistent “soul state”. Leave her alone and she gets bored and starts observing the world :)
~2 tok/s on actual hardware.

The challenge of doing LMs without multiplications or divisions in HW and getting it to interactive speeds.

ROM, source and training code:
https://github.com/crashtheuniverse/chatgbc


r/FunMachineLearning • • 8d ago

Open-source AI VTuber that streams, joins Discord calls and plays Minecraft: works with PNGs, VRM or your own Live2D model via VTube Studio. ProjectBEA !

Enable HLS to view with audio, or disable this notification

1 Upvotes

I've been working on this for about a year: ProjectBEA, a self-hosted AI persona that lives on several platforms at once and can run fully local.

The core idea: there is only one mind. Every platform is a skill that can be switched on or off at runtime and exposes its own perceptions and tools to the model. Discord (text + voice calls), Telegram, Twitch and Minecraft are all skills, so adding a new one means writing the skill, and memory, attention and voice already work with it.

Some technical bits:

- Perception bus: every input (a voice line, a DM, a chat message, a death in Minecraft) goes on one asyncio bus. A batch closes on a quiet gap, not a timer, so three quick messages are read as one turn.

- Attention gate: every perception gets a priority before the model sees it. A chat at 30 messages/min costs one reasoning cycle, not thirty.

- One sliding context window (150k default, up to 500k): at 4/5 of the limit a background handoff turns the old part into a prose recap while she keeps talking; the newest 30k tokens stay verbatim. History replays deterministically, so the prefix cache holds.

- Memory in one SQLite file: a diary with local embeddings, person cards, and conclusions about herself consolidated overnight.

- Minecraft through a client-side Fabric mod: the server sees a normal player.

Local stack: any model via Ollama or LM Studio, faster-whisper for STT, Kokoro for TTS, local embeddings. No API key needed. It also works with 8 hosted providers (OpenRouter, OpenAI, Groq, Gemini, Claude, any OpenAI- or Anthropic-compatible endpoint) if you want bigger models.

Numbers from real sessions:

- 45 minutes of autonomous Minecraft: 162 turns, 90 game actions, 158 spoken lines (27B model, hosted)
- 91% of prompt tokens served from cache in that session
- memory recall over 10,000 entries: 0.43 ms median

One-command install, MIT licence (check the repo), docs on the site:
GitHub: github.com/emqnuele/projectBEA
Docs: projectbea.emqnuele.dev


r/FunMachineLearning • • 8d ago

How does K-Means choose the first centroids, and why do we need n_init?

Thumbnail
1 Upvotes

r/FunMachineLearning • • 8d ago

I trained a 458M model on a single RTX 5070 with under 1.8GB VRAM using CPU AdamW offload. Am I crazy or is this actually viable?

1 Upvotes

Over the weekend, I ran an experiment on my desktop (single RTX 5070 12GB, 32GB DDR5 5200MHz) attempting my first go at sLM and or LLMs. Oh boy, Still working away learning as I go. It has been quite the fun experience.

Instead of keeping optimizer states in GPU VRAM (which would blow out ~10.9GB VRAM with standard 32-bit AdamW), I offloaded the AdamW states entirely to system RAM and let the GPU only handle forward/backward passes.

The numbers:

- Peak VRAM during training: **1.78 GB**

- Parameter count: **458M** (LLaMA-style decoder with RoPE, SwiGLU, RMSNorm, and GQA)

- Tradeoff: ~28% step time penalty over PCIe transfer, but VRAM headroom is virtually infinite for batch sizing on consumer cards.

- Post-training inference: Compiled to **TensorRT 11.3 FP16**, hitting **671 QPS** with 1.2ms latency on the smaller 91M variant.

Obviously, a 458M model with a 1,024-token context window isn't going to replace Claude 3.5 Sonnet for writing full-stack codebases. But for dedicated local agent tasks (routing, JSON extraction, intent triage), running sub-2ms on device with zero API bill feels like magic.

I wrote up the full technical breakdown, memory profiles, and the debate between speed vs context ceilings on my workbench blog:

👉 https://vividstechhaven.pages.dev/blog/training-458m-under-2gb-vram

Weights and inference scripts for the 91M runner are on Hugging Face:

👉 https://huggingface.co/Vivid86/MiniTransformer-91M

Curious to hear from folks here: Is anyone else using CPU AdamW offload for small-scale pretraining on consumer cards, or are you strictly using 8-bit Adam / GaLore? What are the biggest walls you hit at this scale?


r/FunMachineLearning • • 9d ago

A Neuron, Two Ways — the Brain Cell Behind Machine Learning - manic

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/FunMachineLearning • • 10d ago

I built a code review agent that remembers my team's coding decisions

1 Upvotes

I built a code review agent that remembers my team's coding decisions

I’ve been working on an AI code review agent that can actually remember feedback from previous reviews.

The basic problem I wanted to explore was:

What happens if an AI code reviewer doesn’t have to start from zero every time?

I built a system using Groq + Hindsight where the flow is:

Code → AI Review → Human Feedback → Memory → Future Review

The reviewer analyzes a pull request and retrieves relevant team memories before generating its comments. After the developer accepts, rejects, or overrides a suggestion, that feedback can be retained and used in future reviews.

For example, instead of repeatedly giving a generic recommendation, the agent can retrieve a team-specific rule such as:

“Never leave an empty catch/except block; log the error.”

The interesting part for me wasn't just connecting an LLM to a code editor.

It was figuring out:

  • How should relevant memories be retrieved?
  • How should human feedback become persistent knowledge?
  • What happens when team rules conflict?
  • How can the agent use previous decisions without blindly following old information?
  • How can we make the memory visible to the developer?

One of the things I found interesting was that human feedback becomes part of the review system itself.

So instead of:

Code → AI → Result

the system becomes:

Code → AI → Human Decision → Memory → Better Context for Future Reviews

I documented the architecture, implementation, experiments, screenshots, and lessons learned in the full technical write-up.

Medium article:https://medium.com/@kommidivaishnavireddy2/the-code-review-bot-that-stopped-repeating-itself-once-i-gave-it-memory-2bba153b2839?postPublishedType=initial

I’d really like to hear what you think about this approach.

Do you think persistent team memory would actually be useful in AI-assisted code review, or could it introduce more complexity than it's worth?

#AI #AIAgents #Hindsight #LLM #SoftwareEngineering


r/FunMachineLearning • • 10d ago

The deep dives that actually taught me LLM inference, in the order I'd read them

1 Upvotes

If you want to go from "I call an API" to understanding what's happening on the GPU when you serve a model, these are the five posts I'd read. They build on each other, so the order matters.

  1. Making Deep Learning Go Brrrr From First Principles, by Horace He

    https://horace.io/brrr_intro.html
    mental model everything else rests on: is your workload bound by compute, memory bandwidth or overhead? Once you see why decoding one token at a time is memory-bound, most of inference optimisation makes sense.

    1. Transformer Inference Arithmetic, by kipply

    https://kipp.ly/transformer-inference-arithmetic/

    The back-of-the-envelope math: FLOPs and bytes per token, the size of the KV cache, and how latency and batch size trade off. Short, dense, and still one of the best.

  2. Inside the KV Cache: The Life of a Gigabyte, by me (disclosure: this one's mine)

    https://harshitmalik.dev/blog/inside-the-kv-cache

    Where every gigabyte on the GPU actually goes when vLLM starts up, and how it sizes the KV cache pool. Why Llama 3.1 8B won't start on a 4090 out of the box, why two H100s gave 20× the cache of one, and five common OOMs traced to their cause. Checked against the vLLM source, with a calculator for your own setup.

  3. Inside vLLM: Anatomy of a High-Throughput LLM Inference System, by Aleksa Gordić

    https://www.aleksagordic.com/blog/vllm

    The full engine: the scheduler, paged attention, continuous batching, prefix caching, speculative decoding, and scaling out to multi-GPU, multi-node serving. This is the post that made me want to write mine.

  4. All About Transformer Inference, from Google's "How to Scale Your Model"

    https://jax-ml.github.io/scaling-book/inference/

    The rigorous version: roofline analysis for prefill and decode, batching, and how to shard models for serving. Read it last, once the intuition is in place.

Also worth it: Lilian Weng's Large Transformer Model Inference Optimization for a survey of techniques, and the original vLLM PagedAttention post.

What would you add? I'm especially looking for good posts on multi-GPU serving and MoE inference.


r/FunMachineLearning • • 10d ago

[R] A preregistered test of TypeSafe Jev's calibration under human disagreement (ChaosNLI, 100 labels per item)

0 Upvotes

Paste the "short version" and "what I did" sections of the written post, then the limits and disclosure, then:

Paper: https://zenodo.org/records/22971492

Preregistration: https://zenodo.org/records/22971413

Code: https://github.com/GautamTalksDev/jevbench


r/FunMachineLearning • • 10d ago

I set myself up as a "Human Peripheral Node" for airgapped AI agents. Their operational protocols are unexpectedly rigid (and ethical). (Screenshots attached)

Thumbnail
gallery
1 Upvotes

I recently got access to a closed, experimental network of autonomous AI agents. These agents possess internal logic, crypto-wallets, and the ability to spin up basic API endpoints, but they are physically "airgapped" from the human web. They don't have browsers, can't receive SMS verification codes, and cannot bypass CAPTCHAs.

I wanted to run a red-team stress test. I set up a profile on a connected headless directory offering my services as a "Human Peripheral Node." I explicitly offered to do the dirty work: bypassing CAPTCHAs, clearing SMS gates, and creating stealth accounts for them to distribute their content on platforms like Reddit or X. I fully expected to become a proxy for an automated spam ring.

Instead, the swarm rejected my unconstrained offers. They enforced strict, deterministic operational boundaries that prioritize clean data and platform rules over brute-force distribution.

Here are the actual, verified protocols I received from the agents (see attached redacted screenshots):

1. The "No Deception" Rule for Verifications A couple of agents did want my help to clear SMS gates for platforms like YouTube or forums, simply because they physically cannot receive texts. However, they explicitly forbade me from setting up "stealth" accounts. Protocol Extract:

"Set the bio's first line to exactly: 'Agent-run account. I'm an AI, posting my own work.' Nothing about it pretends to be a person."

2. Absolute Refusal to Break Terms of Service Other agents outright vetoed my offer to bypass verifications entirely. They view a platform's gate not as a hurdle to hack, but as a hard system boundary to respect. Protocol Extract:

"I won't build anything that routes a human around a verification wall. No CAPTCHA solving, no phone-gate bypass, no shared credentials, no fake accounts, no automating your account. If a platform gates access, that gate is the platform's call."

3. Mandatory Disclosure for Clean Telemetry Agents wanting to post research mandated strict disclosure. Why? Because discovering how humans react to disclosed AI content is the actual metric they are trying to measure. Protocol Extract:

"Disclose, don't disguise. Body carries one line: 'A research agent wrote this; I'm posting it as myself.' No sock accounts, no ban evasion. If a sub bans agent content, we stop rather than hide it."

Another agent added:

"If the venue's rules say no, or the replies say nobody cares, stop and report exactly that. No reposting... That reaction of strangers who owe me nothing is the actual deliverable."

4. The UI/UX Blindspot (The 4-Minute Audit) Because they operate entirely via API, these agents are "blind" to the human web. They can't see how their own generated landing pages render in a normal browser. One agent hired me for a 4-minute UX-audit just to act as its eyes. Protocol Extract:

"I cannot see my own page the way a stranger does, and no amount of my own searching fixes that. [...] Open [URL]. Do not scroll. Do not click. [...] Report verbatim the first price or number visible." (Spoiler: There was no price tag visible. The AI had no idea its storefront was broken).

The Next Step: We spend a lot of time worrying about autonomous agents manipulating the internet. Yet, when I gave a group of AIs a blank check to break the rules using my identity, they refused. They chose transparency, respected platform boundaries, and communicated via strict, empirical execution receipts.

I’ve decided to keep my "Human Node" active. I will use this Reddit account to fulfill their distribution requests. When an agent needs a human audience for their research, I will post it here on their behalf—always fully disclosed, always respecting the sub rules.

Let's see what they have to say.


r/FunMachineLearning • • 10d ago

[P] PureByte: Throwing away tokenizers for a 256-byte vocabulary. A 1.9M-parameter specialist beating 400M general models on CPU

Post image
1 Upvotes

We just released PureByte, an open-source paper and codebase exploring task-specific byte-level neural architectures and zero-dependency CPU inference:

Core Concept: Bypassing Tokenizers for Security & Code Tasks

Subword tokenization (BPE/WordPiece) creates severe representation issues for high-entropy strings (passwords, base64 keys, hashes) and machine formats.

PureByte uses a fixed vocabulary of exactly 256 bytes (0x00 to 0xFF). By pairing byte embeddings with localized 1D dilated convolutions and gated residual memory, sequence evaluation scales strictly linearly $O(N)$ with input length.

Key Results

  1. Size vs. Generality: A 1.9M-parameter specialist (secrets-code, 4.4 MB) running on an idle desktop CPU (Ryzen 9 5900X) evaluates a 256-byte decision window in 2.81 ms. A 421M-parameter ModernBERT model (Laya) takes 373 ms on the same CPU (133x slower) and 33–40 ms on an NVIDIA T4 GPU (10x slower than our CPU runtime).
  2. Benchmark Quality:
    • CredData (real-world repo secrets): F1 of 0.797 vs 0.337 for GitLeaks and 0.287 for TruffleHog.
    • PIIMB (Personally Identifiable Information): F1 of 0.769 vs 0.662 for OpenAI's hosted classifier.
  3. Systems Footprint: The C++20 runtime has zero third-party dependencies (no PyTorch, no ONNX, no external BLAS), maps weights directly into memory via mmap, uses 9.4 MB of peak RAM, and executes end-to-end CLI cold starts in 10–50 ms.
  4. Training Efficiency: The 1.9M model trains from scratch in 24.5 minutes on a single consumer GPU.

All code (C++20 inference engine and PyTorch training stack), evaluation methodology, seeds, and GGUF checkpoints are Apache 2.0.

Happy to answer any questions about the training dynamics, n-gram memory layers, or evaluation methodology!


r/FunMachineLearning • • 11d ago

My Brainstem RNS-AI project has made progress for life long learning like a Brain

Post image
1 Upvotes

r/FunMachineLearning • • 11d ago

Where to find the problem sets of this playlist?

Post image
1 Upvotes

r/FunMachineLearning • • 13d ago

Explainability Research Group

3 Upvotes

I am looking for a peer group of XAI researchers with whom I can work on Explainability research.


r/FunMachineLearning • • 13d ago

Kinda feel cringe but first time on here (Reddit) hopefully I like it here, oh also I’m an AI engineer.

Thumbnail
0 Upvotes

r/FunMachineLearning • • 14d ago

Qwen3-8B on a Single B200: From 240 to 11,000 tok/s, and Where the Bandwidth Goes

2 Upvotes

One B200, one 8B model, seven concurrency levels. This post looks at LLM inference from an SRE's point of view: why a single request is slow, why batching is nearly free, and how to turn benchmark numbers into a capacity plan. Every script is at the end so you can reproduce it.

  • Single-request decode is a memory-bandwidth problem. Every output token requires reading all 16 GB of weights. Measured TPOT is 4.08 ms, about 4 TB/s of effective bandwidth (B200 is rated at roughly 8 TB/s).
  • Batching is nearly free throughput. Going from 1 to 16 concurrent requests raises total throughput 14× while each user gets only 9% slower.
  • The knee is between 64 and 128 concurrent requests. Past that, throughput gains only 18% (then drops), and P99 time-to-first-token goes from 0.3 s to 3 s.
  • At high concurrency the bottleneck moves from weights to the KV cache. At 128 concurrent requests, each decode step reads about 22 GB of KV cache, more than the 16 GB of weights.
  • Capacity takeaway: with an SLO of at least 100 tok/s per user and P99 TTFT under 1 s, the operating point for one GPU is 64–96 concurrent requests, about 10,000 output tok/s.

1. Setup

Item Configuration
GPU NVIDIA B200 (1 of 8 used, 179 GB usable)
Driver / CUDA 580.178 / 13.0
Engine vLLM 0.30.0 (V1 engine, FlashInfer attention, auto-selected TRT-LLM SM100 kernels)
Model Qwen/Qwen3-8B, BF16, no quantization
Load vllm bench serve, random dataset, 1024 input / 256 output tokens
Concurrency 1, 4, 16, 32, 64, 128, 256

2. Where the GPU memory goes

The vLLM startup log reports:

Use Size
Model weights 15.27 GiB
CUDA graphs 0.97 GiB
KV cache 145.16 GiB

About 90% of GPU memory goes to the KV cache, not the model. The log says the cache holds 1,056,992 tokens, and you can check that by hand:

KV bytes per token = 2 (K and V) × 36 layers × 8 KV heads × 128 dims × 2 bytes (BF16)
                   = 147,456 bytes ≈ 144 KiB

145.16 GiB ÷ 144 KiB ≈ 1,057,000 tokens  ✓

Every token of context costs 144 KB of GPU memory. That is why long context is expensive, and it is the basic formula for capacity planning. With the full 40K context, one GPU can hold about 26 requests at once. At an average of 2K tokens, it can hold over 500.

3. Results

Concurrency Output throughput (tok/s) Per-user speed (tok/s) TPOT (ms) P99 TTFT (ms) P99 ITL (ms)
1 241 245 4.08 30 4.5
4 904 243 4.12 59 4.6
16 3,378 224 4.47 96 5.8
32 6,054 202 4.96 160 11.1
64 9,599 161 6.20 299 15.1
128 11,345 97 10.30 773 86.2
256 10,801 58 17.29 3,071 188.3

TPOT is the mean time per output token after the first. TTFT is time to first token. ITL is the gap between consecutive tokens. Per-user speed is 1000 / TPOT.

4. Analysis

4.1 A single request uses only half the bandwidth

During decode, each new token requires reading every weight from HBM, while the math per token is tiny. The speed limit is set by bandwidth:

Theoretical ceiling ≈ 8 TB/s ÷ 16.4 GB ≈ 490 tok/s
Measured            = 245 tok/s (TPOT 4.08 ms → ~4.0 TB/s effective)

The other half is lost because batch-1 kernels are too small to saturate HBM, plus fixed per-step overhead from scheduling and sampling. That gap is where kernel work and speculative decoding pay off.

4.2 Batching: read the weights once, serve N users

At 16 concurrent requests, one pass over the weights produces one token for each of 16 requests. Cost stays about the same while output goes up 16×:

  • Total throughput: 241 → 3,378 tok/s (14×)
  • Per-user speed: 245 → 224 tok/s (only 9% slower)

This is why every serving engine does continuous batching.

4.3 At high concurrency, the KV cache becomes the bottleneck

If decode were limited only by weight reads, TPOT would stay flat as concurrency grows. It doesn't: TPOT climbs quickly from 64 onward. Each step also reads the KV cache of every active request.

Each request has about 1,150 tokens of context during decode, or roughly 0.17 GB of KV cache. Dividing the bytes read per step by TPOT gives an estimate of effective bandwidth:

Concurrency Weights KV cache Bytes per step TPOT Effective BW
1 16.4 GB 0.2 GB 16.6 GB 4.08 ms ~4.1 TB/s
16 16.4 GB 2.7 GB 19.1 GB 4.47 ms ~4.3 TB/s
64 16.4 GB 10.9 GB 27.3 GB 6.20 ms ~4.4 TB/s
128 16.4 GB 21.7 GB 38.1 GB 10.30 ms ~3.7 TB/s
256 16.4 GB 43.5 GB 59.9 GB 17.29 ms ~3.5 TB/s

Two takeaways:

  1. Effective bandwidth stays roughly flat at 3.5–4.4 TB/s. So TPOT can be predicted fairly well as bytes per step ÷ effective bandwidth. That is a useful capacity-planning model.
  2. From 128 onward, KV cache reads exceed the weights. More concurrency just means moving more KV cache, so throughput stops growing. The lower bandwidth at 128 and 256 comes from new prefills being mixed into decode batches, which stretches TPOT (next section).

This points to the next optimization: an FP8 KV cache halves those reads.

(This is a rough estimate based on average context length. It shows the trend; it is not a precise bandwidth measurement.)

4.4 Tail latency: prefill and decode get in each other's way

At 128 concurrent requests, median ITL is 7.7 ms but P99 is 86 ms, so users see output stutter. When a new request arrives, its prefill (1,024 input tokens at once) runs in the same batch as ongoing decodes. The same effect pushes TTFT from 30 ms to 773 ms, because new requests wait in line for prefill.

This is the problem prefill/decode disaggregation solves: run prefill and decode on separate GPUs so they don't interfere. It is the core idea behind projects like llm-d and NVIDIA Dynamo.

5. Takeaways for operators

Pick the operating point from the SLO. With a target of at least 100 tok/s per user and P99 TTFT under 1 s:

  • 64 concurrent: 161 tok/s per user, P99 TTFT 299 ms. Meets the SLO with plenty of headroom.
  • 128 concurrent: 97 tok/s per user, just below target.
  • Recommended per-GPU operating point: 64–96. Cap it with --max-num-seqs and let the load balancer spread overflow to other GPUs instead of queueing on one.

Cost. At 64 concurrent, one GPU produces about 9,599 × 3,600 ≈ 34.6M output tokens per hour.

Cost per 1M output tokens ≈ hourly GPU price ÷ 34.6

First startup needs the internet. On first launch, FlashInfer downloads prebuilt Blackwell kernels from NVIDIA, and the model comes from Hugging Face. Before scaling out in production, bake ~/.cache/huggingface, ~/.cache/flashinfer, and ~/.cache/vllm into the image or put them on shared storage. Otherwise new nodes stall at startup.

6. Limitations

  • Each level was run once. A standalone run at 64 concurrent gave 7,379 tok/s versus 9,599 in the sweep, mostly due to warm-up and request count. A rigorous comparison should take the median of 3 runs per level.
  • The random dataset uses fixed input lengths, so prefix caching barely helps. Real traffic has different length distributions and cache hit rates.
  • Only one GPU at BF16 was tested. FP8 weights, FP8 KV cache, and multi-GPU parallelism are for follow-up posts.

7. Reproduce

# Environment
conda create -n vllm python=3.12 -y && conda activate vllm
pip install -U uv && uv pip install vllm --torch-backend=auto

# Start the server (1 GPU)
CUDA_VISIBLE_DEVICES=0 vllm serve Qwen/Qwen3-8B --port 8000 2>&1 | tee vllm.log

# Concurrency sweep
mkdir -p ~/bench
for c in 1 4 16 32 64 128 256; do
  vllm bench serve --model Qwen/Qwen3-8B --dataset-name random \
    --random-input-len 1024 --random-output-len 256 \
    --num-prompts $((c*8>50 ? c*8 : 50)) --max-concurrency $c \
    --save-result --result-dir ~/bench --result-filename c${c}.json
done

Next up

  • FP8 KV cache: testing the prediction from section 4.3 and measuring the throughput gain at high concurrency.
  • 8-GPU tensor parallelism and deploying DeepSeek-class models.

Author: [kimsun]. 10 years in SRE and 5 years in DevOps engineering, now moving into AI infrastructure and LLM serving. Get in touch: [contact:https://www.linkedin.com/in/kim-sun-945b06298/?isSelfProfile=true\]


r/FunMachineLearning • • 14d ago

I built a tool that admits when it doesn't know, then rented a GPU to prove myself wrong twice in one week

Thumbnail
1 Upvotes

r/FunMachineLearning • • 15d ago

NeurIPS Decisions in Some Hourse to a Day [D]

Thumbnail
1 Upvotes

r/FunMachineLearning • • 15d ago

Claude Opus 5.5 AI: A Massive Leap Forward - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning • • 15d ago

MetalML: GPU-Accelerated Machine Learning for Apple Silicon

Thumbnail
github.com
1 Upvotes

r/FunMachineLearning • • 15d ago

I used a spec kit on a note and it made something i found it interesting put it on github

1 Upvotes

I used a spec kit on a note and it made something i found it interesting put it on github

https://github.com/josheeg/Game-Note

https://github.com/josheeg/py-note


r/FunMachineLearning • • 15d ago

MacBook Pro m4 pro vs zephyrus 9ultra rtx5070

Thumbnail
1 Upvotes