SAM-3D-Body Auto-Labeling for Person Detection
A SAM-3D-Body labeling workflow that produces person boxes, camera-depth estimates and hand-raise labels, followed by ReID-assisted role review, outlier handling and YOLO export.

From body estimates to training data
I built the dataset workflow around SAM-3D-Body to convert tennis images and video frames into reviewed person annotations. The output supports both person detection and the later depth, gesture, and player-role model.
What the teacher produces
For each detected person, the pipeline stores an image-space box, normalized box coordinates, camera translation, and body keypoints. It records camera-axis depth as distance_z and the translation-vector magnitude as distance_l2; these are distinct measurements and the export configuration selects which to supervise.
Hand-raise labels are derived from body geometry. The rules combine 2D and 3D wrist position, torso-relative upward direction, head clearance, elbow angle, and landmark visibility. Either arm can produce a positive label, and uncertain cases are recorded for review. These are teacher-generated labels, not directly measured ground truth.
Identity-assisted review and cleanup
An OSNet ReID workflow carries track identities between frames and lets the annotator select the active player. The GUI supports correcting identity switches, changing player status, deleting erroneous boxes, and drawing missed people.
Dataset statistics identify implausible depth values. The cleanup tool uses configurable limits, defaulting to 7–40 m in the reviewed script, and moves an entire image/JSON pair when a listed person is outside those limits. This is a dataset-quality rule, not a geometric test that proves a person is inside the court.
Dataset outputs
The workflow exports boxes and normalized depth, hand-raise, and player-role attributes for YOLO training. Its normalization range must match the consuming model; it is configured separately from cleanup. Visualization tools review both image overlays and the recovered 3D labels.
This project supplies the training-data side of the person-attribute detector. The lightweight deployed detector learns from the reviewed annotations without requiring SAM-3D-Body at runtime.

