Enhanced Robotics (Tenniix Official)Company & role details

Grid-Guided Tennis Court Keypoint Detection

A six-channel court detector that pairs RGB with a separate synthetic grid, predicts line–grid intersection landmarks, and reconstructs court lines for camera-pose estimation.

Illustrated summary of RGB and synthetic-grid inputs, court-line intersection detection, line reconstruction, and camera-pose estimation.
Project summary illustration

A spatial reference for court lines

I developed a court detector that learns where court lines cross a regular image-space grid. These repeated intersections provide structured landmarks even when a useful physical court corner is outside the camera view.

RGB plus a separate grid

The model receives six channels: the original RGB image and a three-channel synthetic grid. Grid lines are not painted over the camera image. The grid is a fixed reference input, not a second video frame, even though the implementation calls the task court2f_grid.

The dataset converter intersects annotated court lines with vertical and horizontal grid lines. It assigns distinct classes to the left, center, right, and service-line intersection families and represents each target with a fixed-size box centered on the intersection. Extended class configurations also include net-post endpoints and center/net reference points.

From detections to court geometry

At inference, the centers of detected boxes become image-space points. The reconstruction stage groups them by line type and fits lines through those groups. Post landmarks and intersections constrain the net, sidelines, service line, and center line.

The resulting line structure and reference points feed the camera-pose stage in the dual-camera scene pipeline. This is more specific than predicting only the visible corners of a complete court.

Export options

The repository provides two ONNX input adapters. One embeds a constant grid into the graph and exposes an ordinary three-channel RGB input. The other accepts a vertically stacked RGB image and grid, then concatenates them into six channels inside the graph. Calibration images and runtime preprocessing must match the chosen adapter.

TorchScript and edge inference paths preserve the same RGB/grid ordering. The workflow includes annotation conversion, dataset visualization, training, validation, line reconstruction, and export preparation.