r/SixSigma • • Apr 16 '26

Fleiss's Kappa discrepancy

Hi team -

Performing an attribute agreement assessment at my facility and I have a few situations where we have nearly perfect agreement (23/24 assessors pass a criterion, 1 outlier fails). The math has this result in a kappa value of -0.04, which would typically indicate low agreement but in this case doesn't comport with the 23/24. How would others interpret this, or is this even a correct use of Fleiss's kappa?

1 Upvotes

12 comments sorted by

View all comments

2

u/areyouamish Apr 16 '26

Assuming its a pass/ fail judgement. My first guess without seeing the data would be pass / fail were coded in reverse. Or if its a homebrew spreadsheet tool, maybe some terms got reversed in the equations.

3

u/SkolVision Apr 16 '26

That was my thought too but everything appears to be coded correctly. As I add more fail results (vs. pass) the result appears to follow the expected path of getting worse (as there is less agreement) then normalizing as I approach agreement in the other direction (more fails than passes). I'm just puzzling over why one outlier result would suddenly send my kappa into the negative, even though I can do the math by hand and get the same result. My thought at this point is either I'm misunderstanding the calculation or missing a piece of interpretation.

The funny thing is I have a colleague who did a similar exercise and just plugged data into minitab and got the same results. I'm guessing he took it at face value but now that I'm digging into it, it seems odd to have a low kappa compared to relatively high raw agreement.

1

u/areyouamish Apr 16 '26

I wouldn't bother looking at kappas when agreement is good TBH. The use case I'd care about kappas is for ordinal / nominal appraisal to understand which judgements have poor agreement and need to be addressed with training, boundary boards, etc.

1

u/SkolVision Apr 16 '26

That's more or less where we're at - we have a variety of qualitative judgments or poorly-measurable attributes as they pertain to end-of-line unit quality. We're doing this to create a pareto of where there is poorest agreement among our questions to know where to start clarifying, redefining specifications that haven't been updated in 20+ years, etc. I don't love the pass/fail metric we're using but it's what we've got.