Person Detection with Distance, Gesture, and Player Status Prediction
A single-image YOLO model with person boxes, a camera-depth regression head, and independent hand-raise and player-role classifiers, supported by reviewed teacher labels and edge export.

One person detection, three additional attributes
I extended the person detector to report more than location. From one RGB image, the model predicts each person’s box and detection confidence together with camera-depth, hand-raise probability, and tennis-player-role probability.
Unlike the earlier foreground/background detector, this model uses a single image. It combines several supervised tasks on shared visual features.
Model heads and losses
The standard box and person-class branches remain intact. An additional scalar head uses a sigmoid to predict normalized depth, which is converted to meters using the model’s configured minimum and maximum. It uses direct regression with Smooth L1 supervision, rather than the distribution bins used by the ball-size model.
Two independent attribute logits predict hands-raised and player-role probabilities. Their sigmoid outputs are trained with binary cross-entropy and retained alongside the box through postprocessing. Player role is a per-detection classification; it does not itself assign a persistent identity.
Training-data workflow
The SAM-3D-Body labeling pipeline provides boxes and estimated camera depth. Geometry-based gesture labels and ReID-assisted human review add the hand and player targets.
The maintained dataset generator can retain valid boxes when individual attributes are missing, representing unknown targets separately so they can be masked during training and validation. Capture-grouped splits keep related frames and their variants together. Distance supervision is also masked when augmentations invalidate metric scale.
Depth targets normally use the camera’s forward-axis coordinate, distance_z. They should not be described as an interchangeable Euclidean range measurement. The training, export, and inference configurations must use the same normalization range.
Inference and deployment
The project includes box-and-attribute visualizations, validation of the individual tasks, TorchScript and ONNX export, calibration preparation, and RDK X5 inference paths. The joint tennis pipeline runs person inference on the HD camera and combines the predicted depth with camera calibration to map detections into the shared court frame.
The output supports player selection and gesture-aware control while keeping confidence and attribute probabilities available to the consuming application.
The subsequent person pose and racket-hand project adds 17 body keypoints and categorical racket-hand recognition, using those outputs to derive a free-hand-raised gesture.


