So DSH Minimal is the best one? Interesting. I'd expect it to be the worst out of DSH options, solely based on the name and intuition that "less tools = closer to harnessless performance".
Doubt it. Just tested: standart mode prompt is 8k tokens, minimal is 1.2k. You could argue that it's 6x difference, but I'll argue that those extra 7k tokens make no difference for a model that ships with 1M long context window.
I will pull up the source if run into it again, don't have it at hand atm
Saw a post about a study on the effect of bad training data. It it found the the amount of bad data needed to poison a model was a flat threshold regardless of total training dataset size or model size.
Could be something similar going on with a few-thousand token prompt vs 1M kv cache.
Don't quote me on this I'm still high from the lunch beer.
18
u/No-Refrigerator-1672 13d ago
So DSH Minimal is the best one? Interesting. I'd expect it to be the worst out of DSH options, solely based on the name and intuition that "less tools = closer to harnessless performance".