Description
TwelveLabs is an enterprise video intelligence platform and API powered by multimodal, video-native foundation models. It lets developers build applications with semantic video search, multimodal video understanding, and video embeddings, processing visual, audio, speech, and on-screen text together so users can search, analyze, and reason across video content.
With semantic search you can query entire video libraries in natural language to locate specific actions, scenes, dialogue, and even human emotions across hours or years of footage without tags. The platform also segments content by scene and pacing, generates text such as summaries, chapters, and highlights, supports compliance and brand-safety scanning, and produces embeddings for building search, classification, and recommendation features. Its infrastructure ingests multimodal data through a single pipeline at around 60x real-time speed, indexing an hour of video in about a minute at 10,000+ hours per day.
TwelveLabs is built on two foundation models: Marengo, a multimodal embedding model that turns video into spatiotemporal embeddings findable by content across many languages, and Pegasus, a video language model that reasons continuously over the full temporal arc of an asset up to about two hours. It integrates via API, SDKs, and MCP, and is used across media and entertainment, advertising, sports, government, and security.
The platform is SOC 2 Type II certified with encrypted data handling and can deploy across cloud, private cloud, on-prem, or air-gapped environments. It offers a free tier for testing (under 10 hours of indexing) plus Developer and Enterprise plans, with pay-as-you-go pricing.
TwelveLabs's Core Features
Semantic natural-language search across entire video libraries
Multimodal understanding across vision, audio, speech, and on-screen text
Video embeddings for search, classification, and recommendations
Text generation from video (summaries, chapters, highlights)
Automatic content segmentation and compliance scanning
High-speed ingestion at ~60x real-time, 10k+ hours per day
Marengo embedding model and Pegasus video language model
API, SDK, and MCP access with flexible cloud, on-prem, and air-gapped deployment
How to use TwelveLabs?
Index your video: Ingest your library through the API or SDK.
Search or analyze: Query moments in natural language or request summaries and insights.
Generate embeddings: Turn video into vectors for search, classification, or recommendations.
Integrate: Build features into your app via API, SDK, or MCP.
Deploy: Run in the cloud, private cloud, on-prem, or an air-gapped environment.
TwelveLabs's Use Cases
- Media archive search
- Highlight and clip generation
- Contextual ad placement
- Compliance and brand safety
- Security and incident review







