###
计算机系统应用英文版:2026,35(6):133-143
本文二维码信息
码上扫一扫!
Swin-DPEC: 面向室内复杂场景的自监督单目深度估计并行双分支网络
(成都信息工程大学 计算机学院, 成都 610225)
Swin-DPEC: Parallel Dual-branch Network of Self-supervised Monocular Depth Estimation for Indoor Complex Scenes
(School of Computer Science, Chengdu University of Information Technology, Chengdu 610225, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 43次   下载 81
Received:November 11, 2025    Revised:December 02, 2025
中文摘要: 当前主流的室内自监督单目深度估计框架多采用CNN (convolutional neural network)或ViT (vision Transformer)作为单一骨干网络. CNN受限于局部感受野, 难以刻画长距离依赖, 导致弱纹理区域深度估计模糊; ViT虽具备全局建模能力, 但在高分辨率特征阶段易丢失边缘与细节, 造成物体边界失真. 为此, 本文提出一种并行双分支网络 Swin-DPEC, 通过融合Swin-Transformer的全局上下文建模能力与所设计的双路增强卷积(dual-path enhanced convolution, DPEC)模块的局部细节提取能力, 实现优势互补. 网络的全局支路基于Swin-Transformer构建, 利用层级窗口注意力机制捕获多尺度长程上下文, 为低纹理区域提供几何先验; 局部支路通过设计的DPEC模块强化对边缘、角点等高频结构的感知. 为实现全局与局部支路特征的动态融合, 本文提出交叉注意力特征融合(cross-attention feature fusion, CAFF)模块, 以像素级交叉注意力机制自适应聚合全局语义与局部细节. 在NYUv2与ScanNet数据集上的实验结果表明, Swin-DPEC相较于现有自监督方法取得更优性能. 相对误差分别达到0.123、0.138, 准确率分别达到85.1%、81.1%, 且生成的深度图具有更好的全局结构准确性与局部细节保真度.
Abstract:Current mainstream indoor self-supervised monocular depth estimation frameworks predominantly adopt either CNN (convolutional neural network) or ViT (vision Transformer) as a single backbone. CNN is constrained by limited receptive fields, which hinder the modeling of long-range dependencies and result in blurred depth estimation in weakly textured regions. Although ViT exhibits strong global modeling capability, edge and fine-grained details are easily lost at high-resolution feature stages, leading to object boundary distortions. To address these limitations, a parallel dual-branch network, termed Swin-DPEC, is proposed, which integrates the global context modeling capability of Swin-Transformer with the local detail extraction ability of a dual-path enhanced convolution (DPEC) module to achieve complementary advantages. The global branch of the network is constructed based on Swin-Transformer and exploits a hierarchical window attention mechanism to capture multi-scale long-range context, thereby providing geometric priors for low-texture regions. The local branch enhances the perception of high-frequency structures, such as edges and corners, through the designed DPEC module. To enable dynamic fusion of global and local features, a cross-attention feature fusion (CAFF) module is introduced, which adaptively aggregates global semantics and local details via a pixel-wise cross-attention mechanism. Experimental results on the NYUv2 and ScanNet datasets demonstrate that Swin-DPEC achieves superior performance compared with existing self-supervised methods, with relative errors of 0.123 and 0.138 and accuracy rates of 85.1% and 81.1%, respectively. Furthermore, the generated depth maps exhibit improved global structural accuracy and enhanced local detail fidelity.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金 (62172061); 四川省重大科技专项“揭榜挂帅”项目(2025ZDZX0011)
引用文本:
葛庚,沈艳.Swin-DPEC: 面向室内复杂场景的自监督单目深度估计并行双分支网络.计算机系统应用,2026,35(6):133-143
GE Geng,SHEN Yan.Swin-DPEC: Parallel Dual-branch Network of Self-supervised Monocular Depth Estimation for Indoor Complex Scenes.COMPUTER SYSTEMS APPLICATIONS,2026,35(6):133-143