Abstract:
During coal flotation, ash-content labels for tailings images are obtained through manual sampling and laboratory assays, which are time-consuming and costly. Existing machine-vision-based soft-sensing models for ash content generally require large numbers of labeled samples. However, the number of labeled samples that can be obtained within a short period at industrial sites does not match the requirements of model training, resulting in insufficient model stability and generalization. To address this issue, a soft-sensing method for tailings ash content under small-sample conditions is proposed based on masked autoencoder (MAE) self-supervised pretraining and a cross-attention mechanism. The method uses readily available unlabeled tailings images collected in industrial settings to construct an image-patch masking and reconstruction task based on the masked autoencoder (MAE), enabling the encoder to learn global structural representations of tailings images without relying on ash-content labels. During the soft-sensing stage, the pretrained MAE encoder serves as the global-representation extraction branch, while ResNet-34 serves as the local-feature extraction branch. A cross-attention mechanism is used to enable information interaction between the global representation and local features. The mapped handcrafted features are then fused with the interaction-enhanced deep features through a gated fusion mechanism, enabling the joint modeling of global representations, local features, and handcrafted features. Experimental results show that the proposed method achieves good tailings ash soft-sensing performance on the test set, with R2, MSE, RMSE, and MAE of 0.9007, 3.13, 1.77, and 1.46, respectively, outperforming CNN, ViT-Tiny, SVR, and other comparison models. These results indicate that MAE-based self-supervised pretraining can effectively exploit unlabeled images collected on site, alleviating overfitting and insufficient feature learning in ash soft-sensing models under small-sample conditions. The cross-attention mechanism enhances the fusion of global representations and local features, further improving the model's ash soft-sensing performance and providing a methodological reference for image-based soft sensing of tailings ash content.