[1]SHEN Xueli,LU Chengxiang,CUI Yifeng,et al.Lightweight phase-preserving speech enhancement network with dynamic memory augmentation[J].CAAI Transactions on Intelligent Systems,2026,21(3):802-812.[doi:10.11992/tis.202506018]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 3
Page number:
802-812
Column:
人工智能院长论坛
Public date:
2026-05-05
- Title:
-
Lightweight phase-preserving speech enhancement network with dynamic memory augmentation
- Author(s):
-
SHEN Xueli; LU Chengxiang; CUI Yifeng; JIN Haibo
-
School of Software, Liaoning Technical University, Huludao 125105, China
-
- Keywords:
-
deep learning; deep neural networks; intelligent information processing; natural language processing; phase optimization; parameter optimization; memory-augmented networks; robustness
- CLC:
-
TP183;TN912.34
- DOI:
-
10.11992/tis.202506018
- Abstract:
-
To address phase distortion under low-signal-to-noise ratio (SNR) conditions and inadequate noise adaptability in complex acoustic scenes, this study proposes an enhanced speech enhancement network based on an explicit magnitude-phase framework. First, a memory-enhanced time-frequency transformer is designed. It utilizes a dynamic memory matrix and a gated fusion mechanism to improve modeling of impulsive noise. Second, a sparse retrieval mechanism reduces the scale of parameter interaction, thereby significantly reducing model parameters. Finally, a task-uncertainty-driven dynamic loss weighting strategy is developed to jointly optimize anti-wrapping phase restoration, complex spectral reconstruction, and perceptual quality. Compared with the baseline model, the proposed model achieves a 9.7% reduction in parameters while delivering a 2.3% higher wideband perceptual evaluation of speech quality (WB-PESQ) at –5 dB SNR on the VoiceBank+DEMAND dataset and a 1.33% performance gain on the Domain Name System (DNS) Challenge dataset, demonstrating its effectiveness in phase fidelity and noise robustness.