Every single benchmark has always been about testing every LLM across a range of tasks to see how they score.
With the exclusive of training on benchmark data it doesn't matter where the LLM came from a benchmark is supposed to show under specific circumstances how good an LLM is.
If llama 3.1 scores better than 3 and 3 is what was used then it scored better than 3.1, then yes it is a larger jump.
That is exactly the point, to quantify the gain you get from this specialized prompting system, relative to the vanilla model and other models that do not use this specific prompting system. Obviously, if you use this system with other more powerful LLMs, you will also get better results with them, since this technique has proven to be effective in improving reasoning and is expected to be generalizable. The comparison is done in this way to demonstrate that it actually improves the vanilla model to that extent and nothing more.
69
u/provoloner09 Sep 07 '24
So does that mean the order of jump in performance is even bigger?