Abstract:With continuous advances in artificial intelligence algorithms and rapid breakthroughs in high-precision sensor technology, visual simultaneous localization and mapping (vSLAM) has become a core technology across various fields. This study provides an overview of classic vSLAM methods, classifies them, and reviews recent research progress. Given the varying characteristics of environmental features and perception approaches in vSLAM, feature extraction methods are categorized into direct and feature-based approaches, each discussed in detail. The feature-based methods are further categorized into point, line, edge, and hybrid feature extraction methods, which are analyzed. In addition, the methods of integrating depth information with direct methods and feature-based methods are explored, along with an introduction to the fundamental ideas of each method, their advantages and disadvantages, and a summary of representative algorithms for each method. Key vSLAM components, including feature matching, pose estimation, map updating, and loop closure detection, are introduced, with a systematic discussion of their development. Based on existing research achievements, new development trends in the field of vSLAM are summarized. The currently commonly used public datasets in the vSLAM domain are introduced, alongside a list of performance evaluation metrics for vSLAM systems and representative algorithms for each method, followed by a comparative analysis of experimental results. Finally, the current challenging issues in the field of vSLAM are summarized.