Comparing Training Objectives for Neural Post-Filtering of Coded Music

Alexander Heinmüller

This page contains demonstration material supporting the following paper:

  1. Alexander Heinmüller, Andreas Brendel, Pablo M. Delgado, and Jürgen Herre
    Comparing Training Objectives for Neural Post-Filtering of Coded Music
    In International Workshop on Acoustic Signal Enhancement (IWAENC), 2026.
    @inproceedings{ObjectivesPostFiltering,
    author = {Alexander Heinm\"uller and Andreas Brendel and Pablo M.\ Delgado and J\"urgen Herre},
    title = {Comparing Training Objectives for Neural Post-Filtering of Coded Music},
    booktitle = {International Workshop on Acoustic Signal Enhancement (IWAENC)},
    year = {2026},
    }

Abstract

Neural post‑filters are an effective tool for improving the reconstruction quality of audio codecs, and hence, numerous models have been proposed in the literature. We compare several training paradigms--discriminative training, generative adversarial networks (GANs), score-based diffusion (SD), and conditional flow matching--for post-filtering speech and music, and we evaluate the resulting quality improvements across different perceptual audio codecs. We show that GANs can achieve music quality similar to that of SD models at a fraction of the computational complexity. Since the characteristics of a music signal are relevant for the performance of a codec, we further investigate how neural post‑filters behave across different signal types.

Audio Examples

This demonstration material was generated with items from the publicly available ODAQ dataset [1].

For each item, the following conditions are included:

  • Reference: Item at reference quality
  • Baseline: Coded with mp3 at 16kbps without post-filtering
  • Discriminative: Post-filtered by a discriminatively trained network
  • GAN: Post-filtered by a generative adversarial network
  • Diffusion: Post-filtered by a score-based diffusion model
  • Flow Matching: Post-filtered by a flow matching model

Music A

Music B

Violin

Choir

Snaps

References

[1] Torcoli, M., Wu, C. W., Dick, S., Williams, P. A., Halimeh, M. M., Wolcott, W., & Habets, E. A. (2024, April). ODAQ: Open dataset of audio quality. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)