SliderSound: Transferring Continuous Semantic Controls from Procedural to Pretrained Neural Sound Effects Models

Y. Zong and J. Reiss. Submitted to ICASSP 2027.

Abstract

Sound effects design involves creating and modifying sounds to meet production requirements and creative intentions. Text-to-audio models can support this process by generating high-quality material from natural language descriptions, but precise and continuous control over individual sound properties remains challenging. Procedural audio models provide such control through explicit and meaningful parameters, although their simplified synthesis methods can limit sound realism. To address these limitations, this paper presents a method for transferring continuous semantic controls from procedural models to a pretrained neural sound effects model, enabling parameter-based adjustment of sounds generated from text. We train a control adapter using supervision from paired procedural sounds to learn the changes caused by parameter adjustments and apply corresponding changes to neural outputs. Our evaluations across multiple sound effect categories demonstrate the effectiveness of the proposed method.

Read the manuscript (PDF) ยท All publications