多尺度辅助判别的遥感场景分类网络及病虫害分类应用

Multi-Scale Auxiliary Discriminative Remote Sensing Scene Classification Network and Its Application in Pest and Disease Classification

  • 摘要: 遥感图像场景分类中,由于地物特征复杂导致的类内多样性与类间相似性问题,传统的特征提取与简单的多尺度融合机制难以满足高精度分类的需求。为此,本文提出了一种多尺度多信息辅助判别网络MSIADNet (Multi-Scale Multi-Information Assisted Discriminatiom Network)。首先,设计了多尺度选择注意力模块,通过动态调整感受野,使模型能够在浅层特征中精准提取并融合不同尺寸的场景目标;其次,提出了多信息辅助聚合模块,利用可学习的残差编码主动从深层特征中挖掘易被忽略的次要特征,并与主要特征协同建模,为模型提供更全面的辅助判别信息。实验表明,MSIADNet在UCM、AID和NWPU三个遥感基准数据集上分别取得了99.86%、97.23%和95.00%的最高准确率。此外,为了验证所提方法在应对复杂背景时的跨领域泛化能力,本文将该网络作为应用拓展引入到自建的农作物病虫害分类任务中,同样取得了94.89%和94.35%的优异准确率。实验结果充分证明了MSIADNet不仅在遥感场景分类主任务上性能卓越,且具备出色的应用扩展能力。

     

    Abstract: Objectives: Remote sensing scene classification encounters persistent bottlenecks rooted in intra-class diversity and inter-class similarity, where traditional static feature extraction mechanisms fail to model the intricate spatial topologies of complex surface targets. Existing methodologies predominantly focus on capturing highly salient primary features while discarding surrounding, fine-grained secondary features, such as peripheral structures around urban complexes or dispersed background architectures, thereby compromising classification robustness in highly cluttered environments. To break these limitations, a multi-scale multi-information assisted discriminative network, termed MSIADNet, is formulated with the primary objective of dynamically bridging macro-scene contexts with micro-structural details through joint primary-secondary feature collaboration, thereby significantly suppressing false positives in fine-grained scenario categorization. Furthermore, to evaluate cross-domain generalization and structural adaptability, the structural efficacy of the proposed model is extended and validated on a fine-grained agrarian crop pest and disease classification task. Methods: The architectural framework optimizes feature extraction, aggregation, and semantic cross-validation through three cohesive advanced components, beginning with a multi-scale selective attention module that departs from traditional parallel feature concatenation by introducing a cross-guided attention mechanism with asymmetric large and small convolution kernels to dynamically allocate optimal receptive fields for varying-sized land cover patches. To address the limitations of static codebooks that overfit to dominant regions, a learnable residual encoding module embeds a codebook and a smoothing factor directly into the backpropagation framework for joint optimization, actively mining hidden secondary features from deeper layers while restricting the codebook size to 32 to optimize parameters. These auxiliary representations are then deeply aggregated with core multi-scale macro-scene features through a primarysecondary feature collaboration mechanism, which explicitly incorporates the spatial topology and contextual correlation between primary and secondary land covers into final decision criteria to break conventional singleregion decision boundaries. To bypass the domain chasm between macro-satellite observations and micro-crop images, a multi-stage progressive training strategy executes source domain pre-training on standard large-scale benchmarks to solidify macro-spatial representations, followed by target domain fine-tuning on a self-built agricultural dataset to smoothly adapt the learnable residual encoder to fine-grained lesion textures under strict multiseed statistical protocols using five sequential random seeds. Results: Extensive empirical benchmarking across three standard remote sensing scene datasets demonstrates the superior performance of the proposed network. On the UCM dataset under standard 5:5 and 8:2 splitting protocols, the model achieves peak overall accuracies of 99.20%and 99.86%, respectively, outperforming state-of-the-art architectures including CDLNet and SCViT, while guaranteeing commanding overall accuracies of 95.87% and 97.23% on the AID dataset under 2:8 and 5:5 partitions, and securing top-tier performance at 93.41% and 95.00% on the challenging NWPU dataset under 1:9 and 2:8 splits. Quantitative parameterizations confirm that the architecture maintains a highly streamlined computational footprint, retaining total parameters at 60.26 M and computational complexity at 11.21 MACs. In the extended crop pest and disease identification task, the model successfully registers outstanding overall accuracies of 94.89% and 94.35%, where detailed confusion matrices and Grad-CAM feature visualizations explicitly confirm that the network sharply mitigates category confusion among highly similar biological targets, such as beet armyworm, Asian corn borer, and striped rice borer, by effectively isolating cluttered foliar backgrounds. Conclusions: The implementation of MSIADNet establishes a robust paradigm for scene classification by demonstrating that the explicit synergy between multi-scale selective attention and dynamic residual encoding can effectively counteract deep intra-class variations and high inter-class similarities. Furthermore, the successful transition from macro-scale remote sensing benchmarks to micro-scale agricultural pathology proves that the learned structural representations and background-isolation capabilities possess profound cross-domain scalability.

     

/

返回文章
返回