Abstract
Automatic music mixing aims to combine multitrack recordings into a balanced and coherent musical piece. Because the content of different songs and the subjective preferences of mixing engineers jointly shape the final outcome, a practical system should deliver well-balanced mixes while allowing for controllable stylistic variation. However, most existing methods treat automatic mixing and mixing style control as separate tasks, making it difficult for a single system to produce high-quality mixes while remaining editable and style-aware. To address this limitation, this paper presents Diff2Mix, a generative automatic mixing system based on diffusion models and a differentiable mixing console. This system offers two levels of optional user control: a reference audio enables overall production style control, and the differentiable mixing console provides explicit audio effects parameters for interpretability and fine-grained optimization. We demonstrate our system’s competitive performance through both objective and subjective evaluations in terms of mixing quality and control ability.
Audio Examples
This page provides supplementary audio examples for the Diff2Mix paper, demonstrating automatic music mixing and reference-guided mixing.
The automatic mixing table compares generated mixes with baseline systems and human references. The reference-guided mixing table shows how each song is mixed under different reference production styles.
BibTeX
@inproceedings{zong2026diff2mix,
title = {Diff2Mix: Controllable Music Mixing via Diffusion Models and Differentiable Audio Effects},
author = {Zong, Yisu and Shi, Jinjie and Reiss, Joshua},
booktitle = {Proceedings of the 27th International Society for Music
Information Retrieval Conference (ISMIR)},
year = {2026}
}