TwelveLabs closed a $100 million Series B on July 1, 2026, bringing total funding to approximately $207 million. The round was co-led by NEA and NAVER Ventures, with participation from Amazon, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital, and Red Bull Ventures.
The company is building what it calls a full-stack agentic video intelligence system — not a search tool layered on top of video, but a complete architecture combining perception, knowledge, and reasoning to make video a first-class input for machine intelligence.
Why Video, Why Now
Text has been the dominant substrate of AI for the past several years: language models trained on text, evaluated on text, deployed to produce text. TwelveLabs CEO Jae Lee argues this is a fundamental limitation: “The substrate of machine intelligence is recorded reality in motion, not language.”
The numbers support the framing. Video represents an estimated 90% of the world’s data, yet almost none of it is searchable, structured, or usable by enterprise systems. A media company may have a million hours of archived footage; a hospital may have decades of surgical recordings; a manufacturer may have years of production-floor video. All of it is effectively opaque — stored but not understood.
TwelveLabs is building the infrastructure layer to change that.
The Technology: Marengo 3.0 and Pegasus 1.5
The company’s core platform rests on two models:
Marengo 3.0 is described as “the world’s most powerful video embedding model." It understands every sound, every spoken word, and every motion across time — producing dense vector representations of video content that can be searched semantically, not just by keyword. This is the retrieval layer: given a query, Marengo can surface the relevant moments across a video library.
Pegasus 1.5 operates at a higher level of abstraction. It takes video as input and produces structured data: scene boundaries, named entities, temporal segments, and semantic context. The output is something a downstream system — or a human — can actually reason with. Pegasus converts moving images into the kind of data that feeds databases, triggers workflows, and informs decisions.
Together, the two models form a perception and structuring pipeline that turns raw video into organized, queryable intelligence.
Rodeo is the first application built on top of this stack. Launched in June 2026, Rodeo is TwelveLabs’ first direct product offering — moving the company from pure API infrastructure toward serving end users who want to interact with their video archives without writing code.
The AWS Partnership
Amazon joined the Series B as an investor, and the deal comes with a strategic layer: AWS is now TwelveLabs’ preferred cloud provider under a multiyear commitment, specifically for video inference workloads running on AWS Trainium chips. Trainium is Amazon’s custom silicon for training and inference, positioned as a cost-effective alternative to Nvidia GPUs for AI workloads at scale.
This partnership signals two things. First, Amazon is betting that video AI inference is going to be a major compute category — large enough to warrant building a dedicated customer relationship. Second, TwelveLabs gets access to infrastructure economics that smaller, GPU-dependent competitors cannot easily replicate.
Funding History
| Round | Date | Amount | Key Investors |
|---|---|---|---|
| Series A | Jun 2024 | $50M | NEA, NVIDIA NVentures |
| Series A Extension | Dec 2024 | $30M | Databricks, SK Telecom, Snowflake Ventures, HubSpot Ventures, In-Q-Tel |
| Corporate Minority | Oct 2025 | $3M | — |
| Series B | Jul 2026 | $100M | NEA, NAVER Ventures, Amazon, Radical Ventures, Index Ventures |
Prior funding rounds brought in major names across AI infrastructure and developer tooling. The Databricks/Snowflake participation in the December 2024 extension reflects TwelveLabs’ positioning in the data stack: video intelligence as a data source that feeds the same warehouses and analytics pipelines enterprises already run.
In-Q-Tel’s participation — the CIA’s technology investment arm — flags government and intelligence as a serious target market for video AI.
Use Cases and Industries
TwelveLabs targets six sectors where large video archives already exist but remain largely unprocessed:
- Media and entertainment: content search, clip discovery, highlight generation
- Government: surveillance footage analysis, archival search, evidence processing
- Advertising: brand monitoring, ad placement, content categorization
- Security: real-time incident detection, forensic review
- Sports: play-by-play indexing, athlete tracking, broadcast production
- Automotive: dashcam analysis, driver behavior data, training data for self-driving
In each case, the core value proposition is the same: turn passive video storage into active, queryable intelligence.
Company Growth
TwelveLabs has grown to approximately 178 employees as of mid-2026, up from around 58 a year earlier — roughly a 3× increase in headcount in 12 months. The company operates from San Francisco and Seoul, and will use Series B capital to add offices in New York and London to address global customer demand.
The dual presence in San Francisco and Seoul reflects the company’s origins: TwelveLabs was founded by Korean entrepreneurs building in the US, and NAVER Ventures’ role as a Series B co-lead — NAVER is South Korea’s dominant internet platform — keeps that connection institutionalized.
What Isn’t Clear Yet
Market size versus pipeline: Government, media, sports, and automotive are named as target markets, but these are very different buying contexts. A government contract for surveillance footage analysis has a completely different sales cycle than a media company licensing a video search API. It is unclear whether TwelveLabs is pursuing all of these simultaneously or is concentrating on one beachhead first.
Revenue figures: The funding round does not include any ARR or revenue disclosure. For a company at Series B with $207M raised and 178 employees, some signal of commercial traction would help calibrate the valuation. The absence is not unusual for an infrastructure company still expanding its customer base — but investors reading this round without revenue data are pricing in significant continued growth.
Rodeo product-market fit: Rodeo is described as TwelveLabs’ first “application-layer” product. But it launched only weeks before the Series B closed. Whether it finds a mass-market audience or remains a demonstration of what the underlying APIs can do is still to be determined.
Bottom Line
TwelveLabs is making a bet that the next decade of AI will need to make sense of the world’s video — not just its text. The technology stack (Marengo 3.0 + Pegasus 1.5) is technically specific enough to be credible, and Amazon’s participation as both investor and strategic partner is a strong external validation signal. Tripling headcount in a year and opening offices in New York and London suggests this is not a slow burn.
The question is market sequencing. Government surveillance, Hollywood archives, and car dashcam data are all solvable problems for the same core technology — but they are different businesses. How TwelveLabs allocates this $100M across those channels will determine whether it becomes an AI category leader or spreads thin across too many verticals simultaneously.
Sources: GlobeNewswire press release | Bloomberg | PYMNTS | finsmes | Eastern Herald | TwelveLabs Series A blog