Description
Make-A-Video is a state-of-the-art AI system designed for text-to-video generation, building upon advancements in text-to-image technology. The system learns by analyzing images with descriptions to understand visual concepts and how they are described, alongside unlabeled videos to grasp the dynamics of motion. This dual approach allows Make-A-Video to translate textual imagination into visual narratives.
Users can bring their creative visions to life by providing simple text prompts, generating whimsical and one-of-a-kind videos. The system supports various styles, from surreal and realistic to stylized, and can depict a wide range of scenes, from fantastical creatures to everyday scenarios. Examples include a dog in a superhero outfit, a teddy bear painting, or a robot dancing.
Beyond generating videos from text, Make-A-Video offers capabilities to add motion to static images or fill in the motion between two existing images. It also allows for the creation of video variations based on an original input video, enhancing creative possibilities. The system aims to be the new state of the art in video generation, demonstrating significant improvements in text representation and video quality compared to previous benchmarks, as validated by user studies.
Meta AI is committed to advancing AI responsibly with Make-A-Video. Steps are taken to mitigate the creation of harmful, biased, or misleading content through data filtering and by adding watermarks to all generated videos to identify them as AI-created. While currently a work in progress, the goal is to eventually make this technology publicly available, with a focus on safe and intentional release through continued analysis and testing.
Make-A-Video Highlights
Generates videos from text prompts
Learns from images with descriptions and unlabeled videos
Supports surreal, realistic, and stylized video outputs
Enables creation of videos from single images
Fills in motion between two images
Generates video variations from an input video
Watermarks AI-generated content for identification
Aims for improved text representation in videos
Aims for higher video quality compared to previous SOTA
Incorporates responsible AI development practices
Getting Started with Make-A-Video
Access model: Inquire about future releases or access opportunities.
Input prompt: Provide a text description of the desired video content.
Generate video: The AI system processes the prompt to create a video.
Refine and vary: Create variations of the generated video or add motion to static images.
Integrate responsibly: Utilize the technology with awareness of its AI-generated nature and responsible use guidelines.
Make-A-Video's Use Cases
- Text-to-Video Creation
- Image Animation
- Video Variation Generation
- Content Ideation
- Artistic Expression
- Prototyping Visuals



