Abstract
Text-to-Music diffusion models are increasingly used in real-world applications, yet deployment remains challenging: generations can collapse to limited patterns even with diverse initial noise and prompts, and inference-time diversity control often harms text alignment and fidelity by distorting key prompt cues established in early denoising.
To address this, we propose Padding-Annealed Diffusion Sampling, which perturbs only a padding-indexed subspace while keeping non-padding conditioning fixed, enabling controlled exploration with reduced semantic drift.
However, in a text-unaware VAE latent space, such exploration is less likely to stay within genre-faithful neighborhoods, limiting genre-consistent diversity. We therefore introduce Text-Aware Latent space that aligns local neighborhoods with text-implied genre structure, promoting genre-consistent diversity.
Together, the two techniques form a unified pipeline that, compared to prior methods that perturb the full conditioning, achieves a better text alignment--diversity trade-off: at comparable text alignment, it delivers 15.4% higher diversity with a relatively small fidelity drop, and further improves within-genre diversity by 71.6%.
Audio Examples
Baseline: Limited diversity despite high audio quality, with occasional genre mismatch.
CADS: Enhanced diversity but degraded audio quality, suffering from semantic drift.
PADS (Ours): Achieves rich diversity and high audio quality.
PADS-TAL (Ours): Achieves rich diversity and high audio quality while maintaining strict genre consistency.
The audio samples below were produced by a generative model as demonstration outputs. Unauthorized reproduction or use is prohibited.
Task 1: [Samples with single Prompts & Fixed Initial Noise]
Notice: All samples are generated twice using the same initial seed; however, the initial seed differs across prompts.
Prompt
Baseline
PADS (Ours)
CADS
tropical house, drums, 105 bpm, dance
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
timpani, soundtrack, 125 bpm, soothing, ambient
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
groove, 125 bpm, dance, upbeat, happy, musical instrument
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
piano, 100 bpm, double bass, light, calm
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Task 2: [Samples with single Prompts & Random Initial Noise]
Baseline
PADS-TAL (Ours)
CADS
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Baseline
PADS-TAL (Ours)
CADS
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Baseline
PADS-TAL (Ours)
CADS
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Baseline
PADS-TAL (Ours)
CADS
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Task 3: [Samples with various Prompts & Random Initial Noise]
Pop
Prompt
Baseline
PADS-TAL (Ours)
CADS
110 bpm, bass guitar, hip hop, pop, drums
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
pop, electric guitar, bass guitar, synthesizer, 110 bpm, funky, plucked string instrument
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
inspiring, pop, guitar, 110 bpm, musical instrument, acoustic guitar, dance
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Electronic
Prompt
Baseline
PADS-TAL (Ours)
CADS
upbeat, drum machine, electric piano, 125 bpm, dubstep
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
upbeat, drum and bass, bass, musical instrument, aggressive, instrumental, sampler, electronic, cinematic, 125 bpm
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
synthesizers, inspiring, 125 bpm, summer, spray, waterfall, electronic
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
Electronic Pop
Prompt
Baseline
PADS-TAL (Ours)
CADS
110 bpm, passionate, boing, emotional, soulful, electronic pop, love
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
saturday night, boing, electronic pop, double bass, 110 bpm, happiness, groovy, birthday, the synthesizer, piano, inside
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
funky, guitar, electronic pop, 110 bpm, positive, musical instrument, groovy, outside, boing, upbeat
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
New Age
Prompt
Baseline
PADS-TAL (Ours)
CADS
emotional, new age, piano
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
115 bpm, tranquil, new-age music, moving, piano, strings, melancholic, lullaby, emotional
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
musical instrument, emotional, piano, moving, peaceful, reflective, relaxed, calm, sentimental, electric piano, new age
Your browser does not support the audio element.
Your browser does not support the audio element.
Your browser does not support the audio element.
The audio samples above were produced by a generative model as demonstration outputs. Unauthorized reproduction or use is prohibited.