Temporal Tennis Ball Detection with Ball Diameter Prediction
A three-frame ball detector with explicit diameter supervision, size-aware postprocessing, reviewed temporal datasets, and ONNX/TorchScript deployment paths.

From the size head to diameter prediction
This project completes the diameter-specific workflow built on the joint detection and size-regression branch. The model predicts a ball’s location, frame identity, confidence, and apparent diameter from a three-frame input.
The input contains nine RGB channels. Output classes 0, 1, and 2 identify the corresponding frame within the triplet; the diameter is a separate continuous attribute attached to each detection.
Explicit diameter labels
The dataset converter reads reviewed ball polygons and estimates an equivalent-circle radius from their area, with fallbacks for incomplete geometry. The default diameter target is twice that radius. It is stored separately from the localization box in a six-column label containing frame class, normalized size, center, width, and height.
Normalization uses the average image dimension, with corresponding conversions during preprocessing and training. Fixed-size localization boxes remain independent of the ball-size target. Triplets may contain detections in one, two, or all three frames, and fully empty triplets can be retained as background examples.
Training and postprocessing
A distribution-regression head learns the size attribute using DFL and Smooth L1 losses. The inference decoder converts its expectation into pixels using the feature stride. Prediction, non-maximum suppression, and result objects carry diameter independently of class scores.
Downstream geometry converts diameter to radius where needed. Keeping that factor of two explicit matters: using a diameter as a radius would directly distort size-based depth estimates. Crop, resize, and camera-motion transforms must also preserve the correspondence between size and image coordinates.
Deployment and review loop
ONNX export uses a wrapper that rearranges three vertically stacked RGB frames into the model’s nine channels, alongside calibration-data preparation and edge inference. TorchScript paths support temporal video inference and integration with the multi-branch camera pipeline.
Visualization and validation tools inspect localization and size errors. The broader workflow saves video inference results for corner-case review, annotation correction, and dataset regeneration. Its diameter prediction supplies a geometric measurement; metric depth and 3D trajectory reconstruction are downstream stages.

