Enhanced Robotics (Tenniix Official)Company & role details

Person Pose Estimation with Distance and Racket-Hand Recognition

A single-image YOLOv8 pose model combining person boxes, 17 body keypoints, camera-depth regression, and racket-hand recognition, with pose-derived free-hand gestures and RDK X5 export and inference.

Illustrated summary of single-image person perception with 17 pose keypoints, distance regression, four racket-hand classes, and derived free-hand gesture logic.
Project summary illustration
My contribution
Extended YOLOv8 pose with distance and racket-hand heads, dataset conversion, augmentation safeguards, validation, gesture postprocessing, and edge export and inference.
Outcome
A unified person-perception task that exposes pose, depth, and racket-hand probabilities per detection, with a free-hand gesture signal and matching desktop and edge inference paths.

Project overview

I built the person_pose_attr task at Enhanced Robotics to give the tennis robot a richer description of each visible person: their body pose, estimated camera depth, and which hand holds the racket. The implementation extends YOLOv8 pose so these outputs share one model and one RGB image input.

The project covers model heads and losses, annotation conversion, training safeguards, validation, visualization, and edge export and inference. It builds on the earlier person-attribute work, adding explicit body keypoints and racket-side recognition. These also support a gesture rule that checks whether a free hand is raised.

Model outputs

Each detected person carries four groups of learned outputs:

Output Representation Purpose
Person detection Bounding box and confidence Locate each person in the image
Body pose 17 COCO keypoints, each with image coordinates and confidence Describe the face, shoulders, arms, hips, knees, and ankles
Distance One normalized scalar, decoded to meters Estimate camera-axis depth using the configured training range
Racket hand Four class probabilities: left, right, both, none Identify which anatomical hand or hands hold the racket

The pose is 2D, with distance predicted separately. Depth labels default to the camera’s forward-axis coordinate; the dataset converter also supports Euclidean distance when explicitly configured. Dataset normalization, model settings, and deployment decoding must agree on the chosen quantity and range.

The extended training objective combines the existing detection and keypoint losses with Smooth L1 distance regression and categorical cross-entropy for racket hand. Inference retains all four racket-hand probabilities alongside the selected class.

Dataset preparation and training safeguards

The annotation converter accepts COCO17 joints or the MHR70 layout produced by SAM-3D-Body. It explicitly maps MHR70 joints into COCO order, including the wrist indices, then exports boxes, normalized distance, racket-hand labels, and image-normalized keypoints. Capture-grouped train/validation splits keep frames from the same recording together; grouping by venue is also supported.

Incomplete annotations are handled per task. Missing distance targets are masked, and an unknown racket-hand annotation is ignored for racket classification while the sample still contributes to its valid detection, pose, and distance objectives. Unknown is a training label, separate from the four learned classes.

Horizontal flips swap both the left/right keypoints and the left/right racket label. This preserves the anatomical meaning of the supervision. Distance supervision is masked for augmentations that invalidate metric scale, while compatible objectives continue training.

Free-hand gesture logic

The inference layer combines pose confidence and racket-hand classification to derive a hand-raise signal:

  • A racket in the left hand allows the right hand to trigger the gesture; a racket in the right hand allows the left.
  • With both hands holding the racket, neither hand is free. With no racket hand identified, either arm can qualify.
  • A qualifying arm has its wrist above the elbow and head reference, with its elbow above the shoulder.
  • Missing or unreliable keypoints can produce an uncertain result, which stays distinct from a confirmed negative.

The head reference uses visible face landmarks, with a shoulder-line fallback when face points are unavailable. This is pose-based postprocessing, and its result is exposed alongside the model predictions. It can supply a gesture-aware control application with the raised-hand state and the relevant arm.

Validation and edge inference

Image, directory, and video inference render boxes, confidence-filtered skeletons, depth, racket side, and the free-hand gesture. JSON output preserves named keypoints, detection confidence, metric distance, all racket-hand probabilities, and gesture uncertainty for downstream processing.

Validation covers detection and pose quality, distance error, and racket-hand classification, including per-class precision/recall, macro F1, and a confusion matrix. Racket-hand metrics use matched detections at IoU 0.50 and exclude unknown annotations, so their interpretation remains tied to detection matching.

The deployment tooling includes ONNX export, post-training quantization calibration preparation for Horizon RDK X5, an onboard BIN-model runner, and TorchScript export and inference. The raw deployment contract has 15 tensors: box, person class, keypoints, distance, and racket-hand outputs at each of three feature scales. Corresponding decoders restore those outputs and apply the gesture logic.

Outcome

The completed task brings person detection, 2D pose, depth estimation, and racket-hand recognition into a shared model, with data preparation and inference tools that preserve those meanings through export. The derived free-hand signal adds useful context for interpreting player gestures, while the full pose and class probabilities remain available to the application.

This project records the implemented capabilities; it does not report a measured accuracy or device-latency result.