小样本下基于MAE自监督与交叉注意力机制的尾矿灰分软测量方法

CHEN Guozhen1, WANG Ranfeng1, ZHANG Xiong1, ZHANG Linkai1, ZHANG Shuxin1

  • 摘要: 针对煤炭浮选过程中尾矿图像灰分标签依赖人工采样化验获得,化验周期长、成本高,而现有基于机器视觉的灰分软测量模型通常需要大量带标签样本支撑,工业现场短时间内可获得的带标签样本规模与模型训练需求不匹配,导致模型稳定性和泛化能力不足的问题,提出一种小样本下基于MAE自监督与交叉注意力机制的尾矿灰分软测量方法。该方法利用工业现场容易采集的无标签尾矿图像构建基于掩码自编码器(Masked Autoencoder,MAE)的图像块遮蔽重建任务,使编码器在不依赖灰分标签的条件下学习尾矿图像的全局结构表征;在灰分软测量阶段,将预训练MAE编码器作为全局表征提取分支,并结合ResNet-34局部特征提取分支,通过交叉注意力机制实现全局表征与局部特征的信息交互,再将映射后的人工特征与交互后的深度特征进行门控融合,实现全局表征、局部特征与人工特征的协同建模。实验结果表明,所提方法在测试集上表现出较好的尾矿灰分软测量性能,R2、MSE、RMSE和MAE分别达到0.9007、3.13、1.77和1.46,优于CNN、ViT-Tiny和SVR等对比模型。研究结果说明,MAE自监督预训练能够有效利用现场无标签图像信息,缓解小样本下灰分软测量模型易过拟合和特征学习不足的问题,交叉注意力机制能够增强全局表征与局部特征的融合效果,进一步提升模型的灰分软测量性能,为尾矿灰分图像软测量提供了方法参考。

     

    Abstract: During coal flotation, ash-content labels for tailings images are obtained through manual sampling and laboratory assays, which are time-consuming and costly. Existing machine-vision-based soft-sensing models for ash content generally require large numbers of labeled samples. However, the number of labeled samples that can be obtained within a short period at industrial sites does not match the requirements of model training, resulting in insufficient model stability and generalization. To address this issue, a soft-sensing method for tailings ash content under small-sample conditions is proposed based on masked autoencoder (MAE) self-supervised pretraining and a cross-attention mechanism. The method uses readily available unlabeled tailings images collected in industrial settings to construct an image-patch masking and reconstruction task based on the masked autoencoder (MAE), enabling the encoder to learn global structural representations of tailings images without relying on ash-content labels. During the soft-sensing stage, the pretrained MAE encoder serves as the global-representation extraction branch, while ResNet-34 serves as the local-feature extraction branch. A cross-attention mechanism is used to enable information interaction between the global representation and local features. The mapped handcrafted features are then fused with the interaction-enhanced deep features through a gated fusion mechanism, enabling the joint modeling of global representations, local features, and handcrafted features. Experimental results show that the proposed method achieves good tailings ash soft-sensing performance on the test set, with R2, MSE, RMSE, and MAE of 0.9007, 3.13, 1.77, and 1.46, respectively, outperforming CNN, ViT-Tiny, SVR, and other comparison models. These results indicate that MAE-based self-supervised pretraining can effectively exploit unlabeled images collected on site, alleviating overfitting and insufficient feature learning in ash soft-sensing models under small-sample conditions. The cross-attention mechanism enhances the fusion of global representations and local features, further improving the model's ash soft-sensing performance and providing a methodological reference for image-based soft sensing of tailings ash content.

     

/

返回文章
返回