[1]KOU Qiqi,CHEN Feiyu,ZHANG Huaqiang,et al.Indoor self-supervised monocular depth estimation based on image edges similarity[J].CAAI Transactions on Intelligent Systems,2026,21(3):713-726.[doi:10.11992/tis.202505005]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 3
Page number:
713-726
Column:
学术论文—机器感知与模式识别
Public date:
2026-05-05
- Title:
-
Indoor self-supervised monocular depth estimation based on image edges similarity
- Author(s):
-
KOU Qiqi1; CHEN Feiyu2; ZHANG Huaqiang2; CHENG Deqiang2; HAN Chenggong2
-
1. College of Computer Science and Technology, China University of Mining and Technology, Xuzhou 221116, China;
2. College of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China
-
- Keywords:
-
self-supervision; monocular depth estimation; image edges similarity; feature aggregation; pose optimization; indoor scenes; shape priors; contextual consistency
- CLC:
-
TP391.4
- DOI:
-
10.11992/tis.202505005
- Abstract:
-
In this paper, we propose a self-supervised depth estimation network model based on image edge similarity to address the issue of inaccurate depth inference in indoor monocular depth estimation due to complex structures, severe edge overlapping, and large rotational components. First, we introduce an image edge similarity loss function as a shape prior constraint to mitigate performance degradation caused by occlusions and overlaps. Second, we design an adaptive feature aggregation module to fuse multi-scale features while maintaining contextual consistency, thereby enhancing semantic associations in weakly related scenes. Finally, we propose a rotation optimization module that refines the rotational components in pose estimation by weighted fusion of the vectors from different paths, reducing rotation errors. Experimental results show that our method achieves depth prediction accuracies of 82.9% and 78.0% on the NYU Depth V2 and ScanNet datasets, respectively, outperforming existing state-of-the-art methods. The proposed method can recover depth maps with rich details and clear, smooth edges, effectively improving depth estimation in indoor scenes.