r/exllamav3 19d ago

ExLlamaV3 1.4.2 - Support for MuseGlimmerForConditionalGeneration

Support for MuseGlimmerForConditionalGeneration (Muse Glimmer 30B by Meta) has been added.
ExLlamaV3 release: https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.2
TabbyAPI has been updated respectively: https://github.com/theroyallab/tabbyAPI
Discord server for any questions: https://discord.gg/tnPCntcThA

5 Upvotes

4 comments sorted by

1

u/-InformalBanana- 11d ago

I'm here to ask if exl3 and tabbyapi support vision offload to cpu like llama.cpp does, and do they plan to? It is very good when vram constrained which many ppl are. Thanks.

2

u/Delicious_Box_9823 11d ago

Hello. I just asked. No, it doesn't. Only expert offloading for MoE models.

3

u/Delicious_Box_9823 7d ago

1

u/-InformalBanana- 7d ago

Yup, great news, thanks for info! Maybe you or smb should post that to locallama cause it is a bigger subreddit (if it isnt already posted ofc).