Enhanced Robotics (Tenniix Official)Company & role details

Temporal 3-Frame Tennis Ball Detection System

A YOLO-based temporal detector that stacks three RGB frames into nine channels and returns frame-indexed ball detections, with full-frame and center-crop inference branches.

Three RGB frames feed a nine-channel temporal YOLO model. Output classes 0, 1, and 2 assign detections to their respective input frames.
Project schematic · illustrative design

Problem and approach

A tennis ball may occupy only a few pixels, blur during flight, or blend into court markings. I adapted the YOLO detection pipeline to process three frames together so the model can use appearance and motion context when locating the ball.

Three frames in, three sets of detections out

The input concatenates three RGB images along the channel dimension, producing a nine-channel tensor. The detector uses classes 0, 1, and 2 to associate each predicted ball with the first, second, or third frame in the triplet. These classes identify time positions; they are not three different types of ball.

The training pipeline uses fixed-size boxes around annotated centers. This makes the detector primarily a ball-localization stage: its box dimensions should not be treated as a measurement of the physical ball’s apparent size. Radius estimation is handled separately in the cropped-ball model and later integrated in the ball-size regression model.

Inference and coordinate handling

Two branches process the same triplet: a resized full-frame view supplies broad coverage, and a center crop preserves more detail in the central region. Predictions are mapped back to the original image coordinates, grouped by their frame index, and merged with non-maximum suppression.

Optional KLT feature tracking compensates for camera motion before inference. When enabled, the inverse homography returns detections from the aligned frame back to the source image. Video and image-directory runners support both overlapping and non-overlapping triplets.

Engineering contribution

The work spans temporal image loading, frame-indexed labels, model training and validation, branch merging, visualization, and export. The edge adapter accepts three vertically stacked RGB frames and rearranges them into the nine-channel model input. This establishes the detection component used by the later tennis perception pipelines.