Twelve Labs, a startup building artificial intelligence that can understand video the way language models understand text, has raised $100 million in a Series B funding round as it bets that the next frontier of AI lies in motion rather than words.
The company announced the funding on its blog on July 1, 2026. The round was led by NEA and NAVER Ventures, with participation from Amazon alongside Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital, and Red Bull Ventures. For the wider AI industry coverage tracking where venture capital is flowing, the deal underscores a growing conviction among investors that video is the next major data type AI must master.
'The World Does Not Happen in Text'
Co-founder and chief executive Jae Lee framed the company's mission in stark terms. "Five years ago, we began with a simple observation: The world does not happen in text. It happens in motion," he wrote in the announcement.
In an interview with Bloomberg News, Lee expanded on the idea, arguing that video "is the most similar signal data that we receive as humans to learn about the world." He contrasted that with the dominant AI architectures of the moment, noting that "that's different from even the latest frontier models such as Fable 5 and Mythos, which are still language models."
The argument rests on a gap in how AI currently works. The last decade of AI development, Lee wrote, "made text programmable," but video has yet to receive the same treatment. He described the world's video as "mostly dark matter to machines," sitting in archives, on drones, and in satellite feeds, still accessed largely "through filenames, folders, captions, transcripts, and human memory."
Making Video Usable by Agents
Twelve Labs' stated goal is to close that gap. "Our goal is to make every second of video addressable, searchable, and usable by agents," Lee wrote, pointing to the rise of AI agents, autonomous systems that can reason and take action, as the key driver of demand.
If language models made documents searchable and code generators made software faster to build, a video-understanding model could make the enormous and largely untapped archive of recorded video accessible to automated systems. That has applications across media, security, autonomous vehicles, and enterprise, any field where decisions depend on understanding what is happening inside a video stream.
A Crowded but Distinct Category
Twelve Labs occupies a niche that sits between two of the most visible categories in AI. On one side are video generators, tools that create new video from text prompts. On the other are large language models that process text. Twelve Labs is focused instead on video understanding, interpreting existing video so that machines can search it, summarize it, and act on it.
PYMNTS noted that video generators have been part of a new generation of consumer software categories forming around AI, alongside AI companions, conversational search, and prompt-based coding tools. Twelve Labs' approach targets the supply side of that trend, building the underlying models that could make video a first-class data type for the agents and applications being built on top of it.
Why Investors Are Backing Video AI
The roster of backers in the round points to the strategic interest video AI commands. Amazon's participation is notable given the company's AWS cloud division is a major provider of AI infrastructure and its own investments in video and media services. NAVER Ventures, the investment arm of the South Korean internet giant, and Radical Ventures, a firm focused on AI, signal global interest in the category.
The round also reflects a broader pattern in AI funding through mid-2026, where investors are directing capital toward startups pursuing data modalities beyond text. While text-based models have absorbed the lion's share of attention and investment, video and other forms of rich, real-world data are seen as the next arena where meaningful breakthroughs, and returns, may be found.
What the Funding Enables
Twelve Labs did not detail specific spending plans, but the scale of the raise, $100 million in a single round, suggests significant investment in compute, talent, and model development. Training video-understanding models is far more computationally demanding than training text models, given the sheer size of video files, which raises the cost of progress.
For Lee, the bet is that the payoff justifies that cost. As he put it, "the richest record of reality is still largely outside the semantic layer that modern AI systems use." Twelve Labs' new funding is a wager that bringing video inside that layer will be one of the defining AI advances of the coming years.
Sources: PYMNTS, Bloomberg News, Twelve Labs company blog.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →

