基于改进EfficientNet的煤矸音频分类方法

宋庆军; 焦守悦; 姜海燕; 宋庆辉; 郝文超

doi:10.13272/j.issn.1671-251x.2024090013

基于改进EfficientNet的煤矸音频分类方法

Coal gangue audio classification method based on improved EfficientNet

摘要

摘要: 针对煤矸音频特征提取过程中设备运行噪声干扰严重及单一提取方法易导致信息丢失的问题，提出了一种基于改进EfficientNet的煤矸音频分类方法。采用基于Mel频谱和Gammatone倒谱系数的特征提取方法，有效捕捉矸石声音中的低频信息和细节特征。选择EfficientNet−B0作为骨干网络，并对其进行以下改进：将原有的多尺度通道注意力模块换成卷积块注意力模块，得到卷积注意力特征融合（CAFF）模块，通过网络自学习为不同空间位置的特征分配不同的权重信息，生成新的有效特征；在原有的MBConv模块中并行嵌入频域通道注意力（FCA）模块，加强特征图的表达能力，从而提高整个网络的性能。实验结果表明：引入CAFF模块后，模型准确率提升了0.61%，F₁得分提升了0.52%，且模型收敛更快，说明CAFF模块有效提升了模型对频谱特征的捕捉能力；引入FCA模块后，准确率提升了0.45%，F₁得分提升了0.62%，说明模块的叠加可以进一步提高模型的泛化能力和处理复杂特征的能力；改进EfficientNe模型的准确率为91.90%，标准差为0.108，显著优于同类对比音频分类模型。

Abstract: To address the issues of severe interference of equipment operating noise and information loss caused by single extraction methods during coal gangue audio feature extraction, a coal gangue audio classification method based on improved EfficientNet is proposed. The method adopted a feature extraction approach combining Mel spectrogram and Gammatone frequency cepstral coefficients to effectively capture low-frequency information and detailed features in gangue audio. EfficientNet-B0 was selected as the backbone network, and the following improvements were made: the original multi-scale channel attention module was replaced with a convolutional block attention module, resulting in the Convolutional Attention Feature Fusion (CAFF) module. This module allowed the network to autonomously assign different weight information to features in different spatial positions, generating new effective features. Additionally, a Frequency-domain Channel Attention (FCA) module was embedded in parallel within the original MBConv module, strengthening the representation ability of feature maps and thereby improving overall network performance. The experimental results demonstrated that after introducing the CAFF module, the model's accuracy improved by 0.61%, the F₁ score increased by 0.52%, and convergence was faster, indicating that the CAFF module effectively enhanced the model's ability to capture spectral features. After integrating the FCA module, accuracy improved by 0.45%, and the F₁ score increased by 0.62%, showing that combining these modules further enhanced the model's generalization ability and its ability to process complex features. The improved EfficientNet model achieved an accuracy of 91.90%, with a standard deviation of 0.108, significantly outperforming other comparable audio classification models.

HTML全文

参考文献(23)

施引文献

资源附件(0)