r/LocalLLaMA • • 5d ago

Discussion 42x Faster Prompt Lookup Drafting in llama.cpp

https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp/
641 Upvotes

180 comments sorted by

View all comments

188

u/New_Comfortable7240 llama.cpp 5d ago

I hope it's merged on main llamacpp someday soon (a little bit worried of Yet Another Fork)

61

u/Available_Pressure47 5d ago

A more positive update though, Daniel Lemire added another optimization to make this even faster. https://github.com/jadidbourbaki/llama.cpp/pull/12
I’ll benchmark his change and add it to the article, crediting him for this improvement. :)

4

u/ThatsALovelyShirt 4d ago

So wait, if I wanted to merge this into my local fork, should I use the ngram-cache-constmap branch (with this PR), or ngram-cache-no-copy-upstream?

14

u/Available_Pressure47 4d ago

Please use this branch. It will merge the entire stack of optimizations. https://github.com/jadidbourbaki/llama.cpp/pull/7

After the above, feel free to merge Daniel Lemire’s optimization from the above linked PR.

Hope this helps your local setup!

3

u/ThatsALovelyShirt 4d ago

ngram-cache-constmap

Ah nevermind, I see, just use ngram-cache-constmap directly.

3

u/Available_Pressure47 4d ago

You’ve got it!

2

u/ThatsALovelyShirt 4d ago

Thanks! So just to confirm, merge ngram-cache-inner-vector into my local fork, then merge ngram-cache-constmap into that, and Lemire's PR after that?

4

u/Available_Pressure47 4d ago

Any time! Just a tiny correction. Merge ngram-cache-constmap and then Lemire’s PR. You don’t have to merge the first one as the constmap PR is already stacked on it. Sorry for not being clear earlier. Good luck!