Multi-Modal RGB–Depth Image Segmentation Using Feature Fusion

Authors

  • Noor Safa Department of Electrical Engineering, College of Engineering, University of Al-Mustansiriyah, Iraq Author
  • Zainab Majeed Abid Electrical Engineering Department, College of Engineering, Mustansiriyah University, Baghdad, Iraq Author

DOI:

https://doi.org/10.31272/ajece.42

Keywords:

Multi-modal image segmentation; Cross-attention mechanism; RGB-D fusion; Deep learning; Computer vision; Feature fusion

Abstract

Image segmentation remains a challenging task, particularly in complex environments where visual information from RGB images alone is often insufficient. Factors such as poor lighting, occlusions, and background clutter can significantly degrade segmentation performance. To address these limitations, multi-modal approaches that incorporate additional data sources, such as depth information, have gained increasing attention. However, effectively combining different modalities in a meaningful way is still an open problem. In this paper, a novel multi-modal segmentation framework based on a cross-attention mechanism was presented that enables more effective interaction between RGB and depth features. Instead of relying on simple fusion strategies, the proposed method allows each modality to guide the feature selection process of the other, leading to more informative and discriminative representations. The network follows a dual-branch design, where features are first extracted independently and then fused through the proposed attention module. This paper evaluates the proposed approach on standard benchmark datasets and compares it with several baseline methods, including single-modality and early-fusion models. The results show consistent improvements in segmentation accuracy, particularly in challenging scenarios.

References

He, Q., Wu, M., Zhang, P., et al., “Multimodal image segmentation with dynamic adaptive window and cross-scale fusion”, Applied Sciences, Vol. 15, No. 19, 2025.

Zhang, M., Zhang, Y., Liu, S., et al., “Dual-attention transformer-based hybrid network for multi-modal medical image segmentation”, Scientific Reports, Vol. 14, 2024.

Qin, F., Liang, Y., Yang, C., et al., “Medical image segmentation network based on multi-scale cross-attention and Wavelet Transform”, Journal of King Saud University- Computer and Information Sciences, Vol. 37, No. 97, 2025.

Seungik, L., et al., “CrossFormer: Cross-guided attention for multi-modal object detection”, Pattern Recognition Letters, Vol. 179, pp. 144-150, 2024.

Song, K., Zhang, Y., Bao, Y., et al., “Self-enhanced mixed attention network for multi-modal image segmentation”, Sensors, Vol. 23, No. 14, 2023.

Simonyan & Zisserman, 2015 – Very Deep Convolutional Networks (VGG).

Vaswani et al., Attention is All You Need, 2017.

U-Net: Convolutional Networks for Biomedical Image Segmentation.

Goodfellow et al., “Deep Learning”, 2016.

Website: NYU Depth V2 Dataset.

Website: SUN RGB-D Dataset.

Fully Convolutional Networks for Semantic Segmentation.

V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation.

Nahian Siddique et al., “U-Net and its variants for medical image segmentation: theory and applications”, IEEE Access, vol. 9, pp. 82031–82057, 2021.

Liang-Chieh Chen et al., “Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation (DeepLabV3+)”, ECCV, 2018.

V. Badrinarayanan, A. Kendall, R. Cipolla, “SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation”, IEEE TPAMI, 2017.

C. Hazirbas et al., “FuseNet: Incorporating Depth into Semantic Segmentation via Fusion-Based CNN Architecture”, ACCV, 2016.

Downloads

Published

2026-08-30