DFCRN:一种煤矿场景下基于多尺度复数域的语音增强模型

DFCRN: A Multi-Scale Complex-Domain Speech Enhancement Model for Coal Mine Scenarios

  • 摘要: 传统卷积循环网络在频谱表示上取得进展,但其感受野与跨帧长程依赖建模不足,且传统前馈序列记忆网络(FSMN)在复数域的非线性表达能力有限,难以应对煤矿场景中的强机械噪声、多源非平稳干扰与传输失真,不能有效识别说话人语音及其内容,从而影响安全生产。为此,本文提出扩张频率卷积循环网络(DFCRN)模型,在复数域引入频率扩张卷积编码-解码器(FDC-ED)以多尺度扩大频率感受野,并采用堆叠的复数扩张前馈序列记忆神经网络(CD-FSMN)并行建模时序与频谱记忆,从而保留相位信息并增强对非线性失真的恢复能力。通过在通用数据集(VoiceBank+DEMAND)、DNS-2020及构建的煤矿调度语音数据集(Coal-single)上实验,验证了DFCRN 在主客观指标(如 CSIG/COVL、FWSEGSNR 、SRMR 及若干 MOS 类无参考评估)上均显著优于典型模型(如 GTCRN,AnyEnhance_lightweight等),尤其在真实煤矿噪声下对语音清晰度与可懂度提升明显。

     

    Abstract: Conventional convolutional recurrent networks have achieved notable progress with spectral representations; however, they are limited by insufficient receptive fields and inadequate modeling of long-range inter-frame dependencies. Moreover, traditional feedforward sequence memory networks (FSMNs) exhibit limited nonlinear expressive power in the complex domain, making them ineffective in coal-mine environments characterized by strong mechanical noise, multi-source non-stationary interference, and transmission distortions. The speaker's voice and its content cannot be recognized effectively, thus affecting safety production. To address these challenges, we propose a Dilated Frequency Convolutional Recurrent Network (DFCRN) Model. Specifically, DFCRN introduces, in the complex domain, a Frequency-Dilated Convolutional Encoder–Decoder (FDC-ED) to expand the frequency receptive field in a multi-scale manner, and employs stacked Complex-Dilated Feedforward Sequence Memory Networks (CD-FSMNs) to model temporal and spectral memory in parallel, thereby preserving phase information and improving the ability to recover from nonlinear distortions. Experiments conducted on the general benchmark VoiceBank+DEMAND, DNS-2020, and a newly constructed coal-mine dispatch speech dataset (Coal-single) demonstrate that DFCRN consistently and significantly outperforms representative models (e.g., GTCRN, AnyEnhance_lightweight) across both objective and subjective metrics, including CSIG/COVL, FWSEGSNR, SRMR, and several MOS-like no-reference evaluations. In particular, under real coal-mine noise conditions, DFCRN yields pronounced improvements in speech clarity and intelligibility.

     

/

返回文章
返回