一种改进MaxViT的海岛瞬时水边线精确分割模型
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

P715.7

基金项目:

上海市教委AI赋能科研计划(A1?3405?25?000307)


An improved MaxViT-Based model for accurate segmentation of island instantaneous waterline
Author:
Affiliation:

Fund Project:

Artificial Intelligence Promoting Research Paradigm Reform and Empowering Discipline Leap Plan

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    为了解决高分辨率无人机遥感影像中海岛水边线分割过程中易出现的断裂、误分割与漏分割等问题,本研究提出一种面向高分辨率无人机遥感影像的海岛瞬时水边线分割模型MF-MaxViT(Moga-Semantic FPN-MaxViT)。该模型以MaxViT(Multi-axis vision transformer)为主干网络,主要改进包括:引入局部窗口注意力与全局栅格注意力,实现局部细节与全局语义特征的协同提取;在MaxViT Block中集成Moga(Multi-order gated aggregation)模块,通过特征分解、多阶上下文特征提取及特征聚合,增强模型对多尺度上下文信息的表征能力;引入Semantic FPN(Semantic feature pyramid network)结构,实现多尺度特征融合,提升深层语义与浅层边界信息的协同表达能力。实验结果表明,在自主构建的高分辨率无人机遥感数据集上,MF-MaxViT的平均交并比(Mean intersection over union,mIoU)达94.40%,平均像素准确率(Mean pixel accuracy,mPA)与总体准确率(Overall accuracy,OA)分别为97.15%和97.14%,性能优于DeepLab v3+、CrossFormer、U-Net和TransUNet等主流模型。此外,在2个公共卫星遥感数据集(中国沿海区域数据集和YTU-WaterNet数据集)上,MF-MaxViT亦展现出良好的泛化能力。研究结果表明,该模型能够在高分辨率海岛遥感影像中实现精确且稳定的水边线分割,为海岛遥感动态监测提供了一种高效、可靠的技术方案。

    Abstract:

    To address the problems of fragmentation, misclassification, and omission that often occur in high-resolution UAV remote sensing-based island waterline segmentation, this study proposes MF-MaxViT(Moga-Semantic FPN-MaxViT), a dedicated model for instantaneous island waterline segmentation from UAV remote sensing imagery. The model employs MaxViT as the backbone and integrates both the Moga module and Semantic FPN, with key improvements as follows: Multi-axis attention mechanism: incorporating block attention and grid attention within MaxViT to enable the joint extraction of fine-grained local details and global semantic features; Moga(Multi-order gated aggregation) module integration: embedding the Moga module within MaxViT blocks to decompose features, extract multi-order contextual information, and perform feature aggregation, thereby enhancing the representation of multi-scale context; Semantic feature pyramid: employing a Semantic FPN(Semantic feature pyramid network) structure for multi-scale feature fusion, which strengthens the collaborative representation of deep semantic and shallow boundary information.Experimental results demonstrate that MF-MaxViT achieves an mIoU(Mean intersection over union,mIoU) of 94.40% on a self-constructed UAV remote sensing dataset, representing an improvement of 0.95% over the original MaxViT. The model also attains an mPA(Mean pixel accuracy,mPA) of 97.15% and an OA(Overall accuracy,OA) of 97.14%, outperforming mainstream models including DeepLab v3+, CrossFormer, U-Net, and TransUNet. In addition, MF-MaxViT also demonstrates strong cross-dataset generalization capability on two public satellite remote sensing datasets, the China Coastal Region dataset and the YTU-WaterNet dataset.In summary, MF-MaxViT provides a high-precision, robust approach for UAV-based instantaneous island waterline segmentation, offering reliable technical support for dynamic waterline monitoring and ecological risk assessment applications.

    参考文献
    相似文献
    引证文献
引用本文

王振华,任宇,孔茹,吴静,杨峰,隋家军,宋刚成.一种改进MaxViT的海岛瞬时水边线精确分割模型[J].上海海洋大学学报,2026,35(5):1269-1283.
WANG Zhenhua, REN Yu, KONG Ru, WU Jing, YANG Feng, SUI Jiajun, SONG Gangcheng. An improved MaxViT-Based model for accurate segmentation of island instantaneous waterline[J]. Journal of Shanghai Ocean University,2026,35(5):1269-1283.

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-06-03
  • 最后修改日期:2025-10-19
  • 录用日期:2025-11-27
  • 在线发布日期: 2026-09-08
  • 出版日期: 2026-09-30
文章二维码
关闭