CGZRDM:煤矸石零样本识别蒸馏模型

CGZRDM: Zero-shot recognition distillation model for coal gangue

  • 摘要: 针对井下煤矸石分选场景中,现有煤矸目标检测模型存在误检率高、不具备零样本识别能力的问题,提出一种煤矸石零样本识别蒸馏模型(Zero-Shot Recognition Distillation Model for Coal Gangue, CGZRDM)。首先,利用工业相机采集并人工标注井下煤矸石模拟图像,构建目标检测数据集;基于标注信息提取煤矸石感兴趣区域(Region of Interest, ROI),构建图像分类数据集。然后,使用目标检测数据集训练YOLOv8n,得到煤矸石检测模型。再后,在CLIP-ViT-Base-Patch16的图像编码器中引入纹理增强模块(Texture Enhancement Module, TEM)构建煤矸石零样本识别模型作为教师模型;采用CNN-Transformer混合网络替换教师模型的图像编码器,构建轻量级学生模型CGZRDM。最后,使用图像分类数据集训练教师模型;利用改进的基于特征的视觉变换器知识蒸馏方法(Feature-based Knowledge Distillation for Vision Transformers, ViTKD),将训练好的教师模型的知识迁移至学生模型。生产环境中,使用煤矸石检测模型对实时采集的井下煤矸石图像进行定位,输出煤矸石ROI位置信息;根据位置信息提取煤矸石ROI;利用蒸馏后的学生模型对煤矸石ROI进行分类,输出煤矸石ROI类别信息。实验结果表明:蒸馏后的学生模型在闭集测试集和零样本集上的Top-1准确率分别为93.2%和86.4%,较CBAM-ViT和CLIP-Adapter等对比模型分别平均提升2.7和0.2个百分点;在RK3588边缘设备上推理速度达38.2 FPS。该工作整合YOLO实时检测与CLIP零样本识别能力,为井下煤矸石分选提供了一种高精度、强泛化的解决方案。

     

    Abstract: To address the issues of high false positive rate and lack of zero-shot recognition capability in existing coal gangue object detection models for underground coal gangue sorting scenarios, a Zero-Shot Recognition Distillation Model for Coal Gangue (CGZRDM) is proposed. First, industrial cameras are used to capture and manually annotate simulated underground coal gangue images to construct an object detection dataset. Based on the annotation information, Regions of Interest (ROI) of coal gangue are extracted to build an image classification dataset. Second, the object detection dataset is used to train YOLOv8n, yielding a coal gangue detection model. Then, a Texture Enhancement Module (TEM) is introduced into the image encoder of CLIP-ViT-Base-Patch16 to construct a coal gangue zero-shot recognition model as the teacher model. A CNN-Transformer hybrid network replaces the image encoder of the teacher model to build the lightweight student model CGZRDM. Finally, the teacher model is trained on the image classification dataset, and an improved Feature-based Knowledge Distillation for Vision Transformers (ViTKD) method is employed to transfer the knowledge from the trained teacher model to the student model. In the production environment, the coal gangue detection model performs localization on real-time captured underground coal gangue images, outputting the position information of coal gangue ROI; the ROI are extracted based on the position information; and the distilled student model classifies the coal gangue ROI, outputting the category information. Experimental results show that the distilled student model achieves Top-1 accuracies of 93.2% on the closed-set test set and 86.4% on the zero-shot set, outperforming comparative models such as CBAM-ViT and CLIP-Adapter by an average of 2.7 and 0.2 percentage points, respectively. The inference speed reaches 38.2 FPS on the RK3588 edge device. This work integrates the real-time detection capability of YOLO with the zero-shot recognition capability of CLIP, providing a high-accuracy and strong-generalization solution for underground coal gangue sorting.

     

/

返回文章
返回