Abstract:Stereo vision-based distance measurement often encounters disparity discontinuities and depth distortions in complex driving scenes, primarily caused by weak textures, occlusions, and abrupt boundary changes. To address these issues, this study proposes a vehicle stereo vision depth estimation optimzation method, called FPDE, combining segmentation perception. The proposed method first applies SGBM to obtain an initial disparity map, and subsequently utilizes high-quality instance masks generated by YOLOv8-seg to constrain geometric boundaries and correct mismatches. A distance-transform-based structural enhancement strategy is further introduced to construct a boundary-to-center incremental depth model, enabling hierarchical spatial modeling within object regions. In addition, a percentile-based robust normalization scheme is adopted to suppress extreme values and stabilize the visual rendering of depth maps. Experiments on the KITTI dataset verify that FPDE reduces MAE and RMSE by 21.9% and 13.9%, respectively, and achieves 94.87% accuracy under the $\delta < 1.25 $ metric while maintaining real-time performance at approximately 45 f/s. Comparative experiments demonstrate that YOLOv8-seg outperforms YOLOv5-seg and YOLOv7-mask in terms of both accuracy and speed. These results confirm that FPDE significantly improves boundary consistency and internal depth structure, providing reliable support for vehicle-mounted stereo distance measurement in intelligent transportation and autonomous driving applications.