Reconstruction-Centered Audio Representation Learning for Music (REALM)

Logo_DFG Teaser_REALM Logo_FAU Logo_IIS

The aim of the REALM project is to develop multi-view audio representations that learn from incomplete and heterogeneous music data to support robust, interpretable, and controllable music analysis, generation, and transformation. The project is funded by the German Research Foundation. On this website, we summarize the project's main objectives and provide links to project-related resources (data, demonstrators, websites) and publications.

Project Description

Reconstruction-Centered Audio Representation Learning for Music

This project is situated in the field of Music Information Retrieval (MIR), spanning music analysis tasks such as transcription, beat and harmony estimation, and expressive performance modeling, as well as generative transformations such as timbre transfer and re-instrumentation. Many recent advances are driven by audio representation learning, as the resulting embeddings largely determine what musical properties can be captured, transferred, and controlled. The goal is to learn representations that are musically meaningful, computationally accessible, and robust across tasks, datasets, and recording conditions. However, two challenges remain central. First, music datasets are often heterogeneous and incomplete, for example with scores that lack performance detail or audio that lacks reliable symbolic descriptions. As a result, supervised learning is limited. Second, there is a tension between enforcing invariances to timbre, instrumentation, or recording conditions and preserving enough information to support downstream tasks such as transcription and resynthesis. Learning consistent and reusable representations under these constraints remains an open problem.

To address these challenges, we adopt a reconstruction-centered, multi-view approach to representation learning. We train encoder–decoder models that map audio excerpts, together with aligned symbolic or feature-based views, into a compact latent embedding and reconstruct musically relevant targets (e.g., spectrograms, F0 trajectories, beats, chords, instrument activity), while handling missing modalities through masking and modality dropout. Reconstruction is central because it makes information preservation explicit: the embedding must retain the information required to recover musical content. This enables analysis by synthesis, in which we perturb or edit embeddings, resynthesize audio using deterministic and generative decoders (e.g., regression and diffusion/flow models), and evaluate the results with musically and perceptually grounded criteria.

Building on this foundation, we learn semantic–acoustic embeddings that align audio with symbolic descriptions at local and global time scales, and we develop controllable latent-space transformations for edits such as transposition, re-instrumentation, and performance-related changes. We also establish music-informed evaluation protocols that test fidelity, robustness, and controllability, as well as downstream utility across representative MIR tasks. The outcome is a coherent and reproducible methodology for learning interpretable and controllable music representations from heterogeneous and partially aligned data, yielding unified representations that support analysis, generation, and musically meaningful manipulation within a single framework.

Projektbeschreibung

Repräsentationslernen für Musikaufnahmen mit Rekonstruktion als Leitprinzip

Dieses Projekt ist im Bereich Music Information Retrieval (MIR) angesiedelt. MIR umfasst Musikanalyseaufgaben wie Transkription, Beat-Tracking, Harmonie- und Interpretationsanalyse sowie generative Transformationen, etwa Timbre-Transfer oder Re-Instrumentierung. Viele aktuelle Fortschritte beruhen auf dem Lernen von Audio-Repräsentationen, da sie nachgelagerte Aufgaben vereinfachen. Um musikalisch aussagekräftige, algorithmisch gut nutzbare und über Aufgaben, Datensätze und Aufnahmebedingungen hinweg robuste Repräsentationen zu lernen, müssen zwei zentrale Herausforderungen gemeistert werden. Erstens liefern Musikdatensätze oft heterogene und teils lückenhafte Inhalte, etwa Partituren ohne Aufführungsdetails oder Musikaufnahmen ohne verlässliche symbolische Beschreibungen, was überwachtes Lernen erschwert. Zweitens stehen sich die gewünschte Invarianz gegenüber bestimmten Aspekten wie Timbre, Instrumentation oder Aufnahmebedingungen und die Bewahrung notwendiger Information für nachgelagerte Aufgaben wie Transkription oder Resynthese diametral entgegen. Es ist noch nicht hinreichend gelungen, konsistente und wiederverwendbare Repräsentationen unter diesen Bedingungen zu lernen.

Zur Adressierung dieser Probleme verfolgen wir einen rekonstruktionsbasierten multimodalen Ansatz für das Lernen von Musikrepräsentationen. Wir trainieren Encoder–Decoder-Modelle, die Musikaufnahmen zusammen mit zeitlich ausgerichteten symbolischen oder merkmalbasierten Darstellungen in ein kompaktes Embedding überführen. Daraus werden musikalisch relevante Zielgrößen (z. B. Spektrogramme, F0-Verläufe, Beats, Akkorde oder Instrumentenaktivität) rekonstruiert, wobei fehlende Modalitäten im Training durch Maskierung und Modalitäts-Dropout simuliert werden. Rekonstruktion bildet den Kern des Projekts, da das Embedding die nötige Information zur Synthese musikalischer Inhalte enthalten muss. Das ermöglicht Analyse durch Synthese. Embeddings werden gezielt manipuliert, mit verschiedenen Decodern (regressions-, Diffusion/Flow-, modellbasiert) in Audiosignale zurückgeführt, die anhand musikalischer und perzeptiver Kriterien bewertet werden.

Auf dieser Basis lernen wir semantisch–akustische Embeddings, die Musikaufnahmen mit symbolischen Beschreibungen auf lokalen und globalen Zeitskalen konsistent verknüpfen, und entwickeln steuerbare Transformationen im latenten Raum (z.B. Transposition oder Re-Instrumentierung). Ergänzend etablieren wir musikinformierte Evaluationsprotokolle, die Rekonstruierbarkeit, Robustheit, Struktur und Kontrollierbarkeit der Embeddings sowie deren Nutzen in repräsentativen MIR-Aufgaben prüfen. Ergebnis ist eine kohärente und reproduzierbare Methodik zum Lernen interpretierbarer und steuerbarer Musikrepräsentationen aus heterogenen, nur teilweise aufeinander ausgerichteten Daten, die Analyse, Generierung und musikalisch sinnvolle Manipulation in einem gemeinsamen Rahmen vereint.

Projected-Related Activities

Projected-Related Resources and Demonstrators

The following list provides an overview of the most important publicly accessible sources created in the REALM project:

Projected-Related Publications

The following publications reflect the main scientific contributions of the work carried out in the MPW project.

    Projected-Related Ph.D. Theses

      Links