Project attempt to partially recreate a monocular cone perception pipeline from this paper
BRT Cone Pose Dataset was used as basis for the dataset, which is based on images from FSOCO dataset.
This project was developed in 6 steps:
-
Understand the FSOCO/BRT labels and bounding-box coordinates.
-
Train a YOLO model to detect and classify cones.
-
Prepare 80 x 80 cone crops and keypoint targets.
-
Train a custom PyTorch CNN to predict eight cone keypoints.
-
Connect both models and transform keypoints back to the original image.
-
Use OpenCV PnP to estimate each cone's position relative to the camera.
The PnP results are approximate: the dataset provides neither the original camera calibration nor matching physical 3D cone points. Therefore assumed camera intrinsics and approximate cone dimensions.