Description
Make-An-Audio represents a significant advancement in text-to-audio generation, leveraging prompt-enhanced diffusion models to overcome challenges like limited high-quality datasets and the complexity of modeling long audio sequences. The system introduces a novel pseudo prompt enhancement technique, utilizing a distill-then-reprogram approach. This method effectively alleviates data scarcity by incorporating weakly-supervised data, including language-free audios.
Furthermore, Make-An-Audio employs a spectrogram autoencoder to predict self-supervised audio representations rather than raw waveforms. This architectural choice, combined with robust contrastive language-audio pretraining (CLAP) representations, enables the model to achieve state-of-the-art performance in both objective and subjective evaluations. The system demonstrates remarkable controllability through classifier-free guidance, allowing users to fine-tune generated audio outputs.
Beyond basic text-to-audio, Make-An-Audio showcases impressive generalization capabilities for X-to-Audio tasks, embodying the principle of "No Modality Left Behind." This unlocks the ability to generate high-definition, high-fidelity audio based on user-defined modality inputs. The system supports personalized text-to-audio generation, enabling the injection of unique objects into new scenes and transformation across different styles. Examples include generating a "baby crying" from an initial "thunder" sound, resulting in realistic and faithful audio. It also supports audio inpainting and image-to-audio generation, expanding its creative potential.
Make-An-Audio Highlights
Prompt-enhanced diffusion model for text-to-audio generation
Pseudo prompt enhancement using distill-then-reprogram approach
Spectrogram autoencoder for audio representation prediction
Contrastive language-audio pretraining (CLAP) representations
Classifier-free guidance for controllability
Generalization for X-to-Audio tasks (e.g., image-to-audio, video-to-audio)
Personalized text-to-audio generation with object injection and style transformation
Audio inpainting capabilities
Generates high-definition, high-fidelity audio
Addresses data scarcity in audio generation
Getting Started with Make-An-Audio
Access Model: Obtain access to the Make-An-Audio model.
Integrate via API: Connect to the model's API for programmatic use.
Provide Text Prompt: Input desired text descriptions for audio generation.
Specify Modality Input (Optional): For X-to-Audio, provide relevant image or video input.
Configure Guidance: Utilize classifier-free guidance for fine-tuning audio characteristics.
Generate Audio: Initiate the audio generation process.
Refine Output: Adjust parameters for personalized audio or inpainting tasks.
Make-An-Audio's Use Cases
- Text-to-Speech Synthesis
- Sound Effect Generation
- Audio Inpainting
- Image-to-Audio
- Video-to-Audio
- Personalized Audio Content
- Music Generation
- AI Research


