Dual-Camera 3D Tennis Scene Understanding
A calibrated HD/4K tennis-perception pipeline that combines court pose, multi-branch temporal ball detection, size-derived ball depth, and person-depth predictions in shared 3D court coordinates.

- My contribution
- Temporal model integration, multi-branch ball merging, court-pose estimation, depth-based projection, and scene visualization.
- Outcome
- HD and 4K model outputs mapped into a shared court frame, with 3D ball and player estimates, JSON results, and scene visualizations.
Project overview
I developed the perception pipeline that combines the HD and 4K camera streams into a court-centered scene representation. It connects specialized court, ball, and person models with timestamp alignment and calibrated camera geometry.
The work brings together the diameter-predicting ball model, person-attribute model, and grid-guided court model.
Model inputs and camera alignment
Court inference runs on the 4K view; person inference runs on the HD view. Ball detection uses temporal triplets from both cameras. Timestamp sidecars align the streams, with configurable timing/frame offsets and frame-index fallback when timestamps are unavailable.
Three ball branches balance coverage and detail:
| Branch | Input and role |
|---|---|
| P1 | Resized full 4K frames for wide coverage |
| P2 | A central 4K crop for a larger apparent ball; crop modes are configurable |
| P3 | An aligned HD triplet whose detections are projected into the 4K view using calibration and estimated depth |
Predictions are returned to the source coordinates, with optional KLT compensation for camera motion. Duplicate suppression uses box overlap and radius-scaled center separation, ranked by confidence with small branch preferences. The maintained pipeline also offers the cropped-ball model as an optional post-merge radius refinement.
Recovering camera pose
The court model supplies grid-intersection detections, reconstructed lines, and available net-post/reference landmarks. These produce 2D-to-3D correspondences against the known court model.
A PnP stage estimates the 4K camera pose, with RANSAC where applicable, refinement, reprojection-error checks, physical-position checks, and temporal smoothing. The HD camera pose follows from the calibrated transform between cameras. Rejected estimates are handled through the pipeline’s pose-state logic rather than accepted unconditionally.
Lifting detections into 3D
Ball depth is estimated from apparent image radius and an assumed physical ball diameter, using undistorted camera coordinates. The default physical diameter is configurable at 0.067 m. The resulting camera-space point is transformed into the court frame. This implementation uses size-based depth and calibrated cross-camera projection; it does not require a matched pair of ball detections for stereo triangulation.
For people, the pipeline uses the bottom-center of the detected box and the learned camera-axis depth. It projects that point through the HD camera model and transforms it into court coordinates, retaining the hand and player-role attributes when available.
Outputs and inspection
The runners can process the camera videos or reuse saved detection JSON. Outputs include camera pose, court geometry, ball and person coordinates, branch provenance, timestamped scene JSON, bird’s-eye-view video, and 3D visualization. A history of ball positions provides the displayed trajectory.
Engineering scope
The system integrates training outputs, TorchScript inference, temporal alignment, coordinate transforms, branch merging, and visual debugging. Geometry depends on camera calibration, landmark quality, apparent ball size, and the learned person depth. Small ball-size errors can cause substantial depth changes, so these outputs are estimates for downstream analytics rather than a measured guarantee of line-call accuracy.
The page’s diagrams explain the implementation. The related edge-tennis project retains its separate published demonstration.
Evidence & demonstration



