基于融合注意力的矿井低照度图像增强模型

Fusion attention-based model for low-light image enhancement in mines

  • 摘要: 针对煤矿井下图像增强过程中易出现暗区、细节恢复不足与噪声同步放大的问题,提出了一种基于融合注意力的矿井低照度图像增强模型。采用编码−解码网络作为主干网络,在各阶段引入残差密集块(RDB),以缓解井下暗区弱纹理在深层特征传递过程中易衰减、细节难保留的问题;在解码阶段设计通道–空间融合注意力(CSA)模块,以增强设备边缘、人员轮廓等关键区域响应,并抑制背景噪声及粉尘散射干扰;构建由均方误差损失与感知损失组成的复合损失函数,以解决亮度恢复过程中结构保持与视觉自然度难以兼顾的问题,在保证像素一致性的同时提升结构保真度与视觉质量。基于自建CMUHL矿井数据集开展对比实验,结果表明:所提模型的峰值信噪比(PSNR)达30.544 dB,结构相似性(SSIM)达0.914,均优于FLOL,AnlightenDiff,Retinexformer等模型;模型参数量为4.37×106个,计算量为56.19 GFLOPs,在取得较优增强性能的同时兼顾了模型复杂度与部署可行性。真实场景验证结果显示,该模型在自然图像质量评价器(NIQE)指标上取得最低值,说明增强结果在统计自然度方面表现较好;盲/无参考图像空间质量评估器(BRISQUE)指标整体保持较低水平,与主观视觉效果一致,表明其在真实井下复杂光照环境下具有较好的稳定性和适用性。

     

    Abstract: To address the problems of dark regions, insufficient detail recovery, and simultaneous noise amplification that easily occur during underground mine image enhancement, this study proposed a fusion attention-based low-light image enhancement model for mines. An encoder-decoder network was adopted as the backbone network, and Residual Dense Block (RDB) was introduced at each stage to alleviate the attenuation of weak texture information in dark mine regions during deep feature transmission and the difficulty in preserving details. A Channel-Spatial Attention (CSA) module was designed in the decoding stage to enhance the responses of key regions such as equipment edges and personnel contours and to suppress interference from background noise and dust scattering. A composite loss function consisting of mean squared error loss and perceptual loss was constructed to address the difficulty of balancing structural preservation and visual naturalness during brightness recovery, thereby improving structural fidelity and visual quality while ensuring pixel consistency. Comparative experiments were conducted on the self-built CMUHL mine dataset. The results showed that the proposed model achieved a Peak Signal-to-Noise Ratio (PSNR) of 30.544 dB and a Structural Similarity Index (SSIM) of 0.914, outperforming FLOL, AnlightenDiff, Retinexformer, and other models. The model had 4.37×106 parameters and a computational cost of 56.19 GFLOPs, indicating that it achieved superior enhancement performance while balancing model complexity and deployment feasibility. Real-scene validation results showed that the proposed model achieved the lowest Natural Image Quality Evaluator (NIQE) value, indicating better statistical naturalness of the enhanced results. The Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE) values remained relatively low overall and were consistent with the subjective visual effects, indicating that the model had good stability and applicability in complex real underground mine lighting environments.

     

/

返回文章
返回