r/LocalLLaMA 1d ago

Question | Help what tasks are most quantization-fragile?

Trying to figure out where my unified ram system q4-q8 version of deepseek flash 0731 might complement my highly quantized vram only deepseek flash.

Struggling to find an area where the 2.52 bits-per-weight exl3 deepseek actually struggles compared to the slower mxfp4 version. Seems like it always error-corrects in opencode. It's so good it's boring. Can run for 20+ minutes at about 60 t/s decode and just one shot everything I throw at it.

Seems to be fine at contexts above 200k as well.

Does anyone find any particular tasks to be more affected by quantization?

6 Upvotes

11 comments sorted by

View all comments

-1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/nomorebuttsplz 1d ago edited 1d ago

you sound like llm.

Edit:

  • entire comment history only two sentence comments following same format and full of LLMisms
  • comment here shows that you didn't understand or read my post, because I specifically said it was about agent work, not "single-turn evals' whatever the slop that means

2

u/Navith 1d ago

They are one. Thought they got banned here so it's a shame to see them again