Enhanced Robotics (Tenniix Official)Company & role details

Cropped-Ball Radius Estimation for Tennis Depth Inference

A dedicated CNN that estimates ball radius in crop-space pixels using distribution regression, with single-frame and paired-frame inputs for downstream camera-geometry calculations.

Illustrated summary of cropped-ball radius estimation with single or paired image crops, a CNN distribution head, and downstream camera-geometry depth inference.
Project summary illustration

Why a separate size model

A localization box is a poor substitute for the visible radius of a tiny or motion-blurred tennis ball, particularly when the detector is trained with fixed-size boxes. I built a dedicated cropped-ball regressor to estimate that radius as a separate measurement.

Input and model architecture

The model supports a single RGB crop or a six-channel pair containing the current crop and a reference-frame crop from the same image location. The training entry point defaults to a 96 × 96 crop and three channels; six-channel operation is an explicit configuration.

A compact YOLOv8n-style backbone uses convolution, C2f, and SPPF blocks, followed by global average pooling and a linear distribution head. Softmax probabilities over radius bins are reduced to an expected value and scaled into crop-space pixels. Training combines Distribution Focal Loss with Smooth L1 supervision on the decoded radius.

Label preparation and training

The preprocessing pipeline reads ISAT segmentation polygons, rasterizes the ball mask, estimates its center with a distance transform, and derives a radius from configurable boundary-distance statistics. It then crops the image and normalizes the radius label by the crop size. Paired-frame datasets crop both images at the same location.

Preprocessing can generate samples at different image scales while recalculating the geometry. Training augmentations preserve radius under translation, flips, and right-angle rotations; appearance augmentations vary color, noise, and blur. Validation reports pixel error and saves predicted-versus-labeled radius overlays.

From pixels to scene geometry

The network outputs radius, not metric distance. Downstream code combines the radius with camera intrinsics, the crop-to-image scale, and an assumed physical tennis-ball size to estimate depth. Correctly undoing crop resizing is therefore part of the integration.

TorchScript export supports the model’s input-channel configuration. The dedicated ONNX edge-export script implements the single-frame RGB path. The later multi-branch ball pipeline can also use this model as an optional radius-refinement step after merging detections.