Interactive SAM3 Auto-Labeling for Tennis Ball Detection
A PyQt/SAM3 annotation workflow with multi-frame prompts, mask propagation, review and correction loops, and ISAT export in the original image coordinates.

Reducing repeated annotation work
I adapted the SAM3 video predictor into an interactive annotation workflow for tennis footage. An annotator can mark a ball on selected frames, propagate masks through a clip, inspect failures, and correct them before exporting training labels.
Prompt, propagate, review, correct
The PyQt interface accepts positive and negative point prompts on multiple frames, with box-prompt support as well. Prompts retain their frame indices so the predictor can use corrections at different points in the sequence.
After propagation, the preview interface lets the annotator add, change, or remove corrections and rerun the predictor. A single-frame preview provides a quicker check before propagating again. The workflow therefore supports repeated review, rather than assuming one click on the first frame will correctly label an entire video.
Complementary image views
The annotation script supports a cropped ball view and a resized full-frame view. In the current implementation, the ball crop is 1280 × 640 with configurable offsets, while the resized view limits the longest side to 640 pixels. Processing uses temporary frame copies.
The crop path rejects masks touching its boundary. The resized path skips balls fully contained within the corresponding crop region. These rules let the two views cover complementary parts of the scene while reducing partial and duplicate labels.
Export and downstream training
Accepted masks become ISAT polygon annotations with bounding boxes, area, object grouping, and task notes. Coordinates are transformed back to the original image using the resize scale or crop offset. Both ball views use the ball category, with task notes distinguishing their origin.
Batch mode reuses the loaded model across clip directories and supports resuming from a selected directory. Separate dataset converters turn these reviewed annotations into temporal detection and ball-size training samples.

