Time-shifted anechoic targets, a two-stage distortion-perception framework, and curated data improve universal speech enhancement and achieve state-of-the-art results.
Vanilla LoRA matches variant performance within 1-2% when learning rates are tuned, and differing optimal rates stem from Hessian eigenvalue variations.
Discrete diffusion models learn data support before frequencies because reverse edits scale by validity first and coefficients second; absorbing diffusion prioritizes validity-improving moves over uniform diffusion's trichotomy.