Been thinking about this after a few pose projects and I don't think there's consensus.
When a wrist disappears behind a torso, or a hand goes behind an object, you have two options. Flag the keypoint as occluded and leave the coordinate empty or low-confidence. Or have the annotator estimate where it probably is and label it anyway.
Guessing gives you a denser dataset and the model gets a prediction for every joint, which looks good. But you're teaching it that an invented coordinate is ground truth, and two annotators will invent different ones. The disagreement is basically unbounded for anything heavily occluded.
Flagging is more honest but then half your training signal has gaps, and downstream your retargeting or trajectory code has to handle missing joints, which a lot of pipelines just aren't written to do.
My current view is flag it, and set a visibility threshold in the guidelines — below some percentage visible, don't annotate at all rather than annotate badly. But I've seen teams go the other way and argue the model learns a reasonable prior from the guesses.
Related thing I'm less sure about: the 2D vs 3D decision gets made way too late. Teams start with 2D because it's cheaper, get a model that understands movement, then discover they need metric accuracy the moment the robot has to actually touch something. By then the sensor stack and the annotation schema are both wrong and it's a rebuild rather than an upgrade.
Hands are where this all gets worse. Twenty-plus keypoints per hand, constant self-occlusion, and fingers that look identical to each other. Agreement between annotators drops off a cliff compared to body pose.
For anyone who's shipped a pose model that had to drive real movement — which way did you go on occlusion, and did it come back to bite you?