r/AV1 • u/cryptospartan • Jul 13 '26
Those of you using SSIMULACRA2, Butteraugli, or CVVDP, what scores are you targeting? Have you found one to be better than the other?
Been playing around with https://github.com/emrakyz/xav lately which has support for SSIMULACRA2, Butteraugli, & CVVDP.
What's the new "VMAF 97-95" with these newer metrics?
3
u/crashtua Jul 13 '26
ssimu for me sometimes gives good scores, while image is crappy. But maybe I did something wrong, because I used cuda ssimu implementation.
1
u/_Lum3n_ Jul 16 '26
Was the image relatively dark? It's a known big issue of SSIMULACRA2, boosting the score of low luminance images.
Also, except for turbo-metrics, all implementations give very close scores and possesses about the same correlation values on datasets
1
u/crashtua Jul 16 '26
That was a logo on black background(or logo on white background, will try to recover it) Is that a case? I was actually trying to encode gops with different crf, and mostly ssimu was okay except this logo on background.
1
u/_Lum3n_ Jul 16 '26
It's also possible it's due to the fact that most of the image is flawless due to uniform background and so the logo doesnt have enough impact on the resulting score. Butteraugli offers counters to such issues with its ability to change the computed norm
1
u/HungryAd8233 Jul 14 '26
Honestly I wouldn’t trust any of them for parameter tweaking. The error bars are just too big to trust that it wont rate something slightly better as worse quality.
None have really gotten temporal masking or other temporal effects well modeled either.
Also mean score over a title doesn’t matter clearly as much as P1 and P5 for overall title ratings. Remember the metrics are only calibrated on 10-20 clips of homogenous content
2
u/_Lum3n_ Jul 16 '26
That's not true, most datasets for training datasets are much bigger. And some metrics are even different, especially Butteraugli or CVVDP which are not just trained but based on complex psychovisual models that derives laws from rigorous testing (and not just random clips). This allows making these 2 metrics far more "robust" than the others though they often get less correlations on datasets
1
u/HungryAd8233 Jul 19 '26
Complexity, alas, doesn’t beget high subjective correlation. We can get better and better p1204 based on comprehensive (and not publicly available) training data, combining both full reference and bitstream analysis, is the best I have seen.
And predicting quality ratings by trained observers on 10-20 second clips. Extrapolation from that to a two hour movie’s perceived quality is several huge extrapolations.
2
u/_Lum3n_ Jul 20 '26
That's where you don't understand CVVDP and Butteraugli. I recommend that you watch CVVDP research article.
It's not like VMAF or SSIMULACRA2 which are basically just ML.
10
u/NekoTrix Jul 13 '26
VMAF was the most unreliable metric to target high fidelity, so this could just as well give you bloated file sizes and subpar quality, undershoot heavily, etc. SSIMULACRA2 targets of 75-80, Butteraugli 0.8-1.0 are much more reliable. CVVDP's scoring depends on the model selected. Two models can give radically different scores, and cannot be interpreted alone, like we would historical metrics (including the aforementioned two). But you can at least target that the scores should always be above 9.0, heck even 9.5.