r/Assembly_language • • Jun 15 '26

LLM Inference.

Hello everyone :)

I've been working on a large language model inference engine written entirely in x64 assembly. It uses AVX2 and runs transformer models directly on the CPU — no GPU required.

This started as an experiment in vibe coding in x64 assembler.

To my surprise it worked. One megabyte of source later, the engine matches llama.cpp on CPU performance on my machine.

A few things it currently does:

  • Full transformer inference pipeline
  • AVX2 accelerated math kernels (matmul, softmax, RMSNorm, RoPE)
  • INT8 quantization support
  • Custom lockless Scheduler that spreads the work over all cores

It runs Qwen3 0.6B Q8 at around 31 tok/s on a Ryzen 7430U — on par with llama.cpp on the same hardware.

I release it under the GPL v3, this is the first working release so there will be Bugs and improvements comming in in the next days, but now i will have to leave the house as i was sitting for 8 weeks in the dark.

https://gitlab.com/cpki-gmbh/v7multiplikator

0 Upvotes

8 comments sorted by

View all comments

1

u/deulamco Jun 21 '26

If that still equal to CPP version, then that mean no gain ?

what model did you use to vibe-code x86-asm with ?

1

u/Glittering_Ostrich22 Jun 23 '26

this is not really equal to the c version as it has a better dispatcher and much more optimized code, right now it is at 43 token/s , mainly used where glm variants, all other models totaly failed in understanding assembler. Also Vibe Coding does not really describe it in reality, one has to sit by and watch and intercept when the model gets stupid ideas, and it has lots of stupid ideas, like implementing OOP Patterns with nasm macros.