Yang Zhang, Wenbin Jiang, Zhen Wang, KaiYing Wu, Wen Zhang, and Fei Wen
Abstract
Generative speech enhancement (SE) methods, typically implemented with diffusion or flow matching, exhibit strong generalization to unseen acoustic conditions. However, they often underperform discriminative approaches in matched, in-domain settings. Recent attempts to close this gap have relied on pretrained discriminative models or cascaded multi-flow architectures, which add inference steps and increase computational cost. To address these limitations, we propose HyFlowSE, a flow matching SE framework with hybrid generative–discriminative learning. HyFlowSE leverages neural ordinary differential equations (ODEs) for end-to-end training and jointly optimizes generative and discriminative objectives within a single model. Experiments on popular benchmark datasets show that HyFlowSE, with only 5.2 M parameters, outperforms other generative SE methods across nearly all evaluation metrics, with especially pronounced gains at low signal-to-noise ratios.