AURA: Unified Multimodal Framework for Conversational Music Editing
Abstract
AURA is a multimodal conversational framework that uses a large language model to interpret dialogue and reference audio, then injects distilled concept tokens into a frozen MusicGen model for precise, progressive music editing with minimal trainable parameters.
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large language model to interpret the complete dialogue history, an optional image, and reference audio, distilling the editing intent into compact concept tokens. A concept-to-audio module injects these tokens and frame-aligned reference features into a frozen MusicGen backbone, enabling precise edits while preserving unaffected content. AURA optimizes only 91M parameters while retaining 1.9B frozen backbone parameters. Experiments on Slakh2100 and MoisesDB demonstrate substantial improvements in edit correctness and content preservation over existing instruction-guided methods, including a 4-5 times reduction in FAD for out-of-domain addition and removal.
Get this paper in your agent:
hf papers read 2609.14344 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
OpenRB-Lab/AURA-Chat-Edit
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper