Jev, the non-LLM "decision model" from startup TypeSafe AI, has completed Pokémon Red — fighting its way from Pallet Town to the Hall of Fame in 37 hours and 40 minutes of play at a total compute cost of about $1.65, according to the project's website.

The run, built by developer Christian Mathiesen, was streamed live with tokens and costs on display and open-sourced on GitHub. It shot to the front page of Hacker News over the weekend, drawing hundreds of upvotes and a lively discussion about what it says about the future of AI agents. Tom's Hardware covered the achievement, noting that a non-LLM engine succeeded where traditional chatbots had stalled for months — and that Claude Opus 5 coached the model through its dead ends.

The Run by the Numbers

The statistics from the project's own tracking show how unusual this run was for an AI system:

  • Total time: 37 hours, 40 minutes
  • Decisions made: 16,150
  • Input tokens: roughly 39.2 million
  • Total Jev cost: about $1.65
  • Typical decision time: 0.4 seconds
  • Opponents fought: approximately 1,660
  • Team wipes: 16
  • Elite Four attempts: 15
  • Toughest stretch: Victory Road

The final Hall of Fame team: Charizard at level 83, backed by Graveler (62), Nidoqueen (45), Beedrill (44), Haunter (39), and Primeape (29).

The economics are the headline. Sixteen thousand decisions for less than the price of a coffee is a striking contrast with LLM-based game-playing agents, which can burn through tens of dollars in tokens over long sessions. Mathiesen said he chose Pokémon after concluding that Jev could decide quickly, but not yet quickly enough to play Doom.

Why Pokémon Red Is Hard for AI

Pokémon Red has become an unlikely benchmark for agentic AI. The 1996 Game Boy classic demands long-horizon planning: grinding levels, managing a party of six, navigating maze-like caves without a map, and backtracking when a key item was missed. Language-model agents have historically wandered in circles, got stuck in caves, and burned enormous budgets on redundant actions.

Jev takes a fundamentally different approach. Rather than generating text, the model returns calibrated decisions directly — a design TypeSafe AI has marketed as a "System One" alternative to the deliberate, language-heavy reasoning of large models. According to Tom's Hardware's earlier coverage of the model, the company claims Jev is roughly 193 times faster and 445 times cheaper than LLM alternatives on decision tasks.

The catch, the developer acknowledged, is that Jev needed help. Per Tom's Hardware, Claude Opus 5 coached the decision model through its dead ends — a hybrid approach in which a large language model provides strategic guidance while the fast, cheap model executes thousands of individual decisions. The project page wryly notes that Jev "only knew where to go next because it had a guide."

What It Means for AI Agents

The result is a data point in a growing argument that the future of AI agents is not one giant model doing everything, but a division of labor: fast, cheap, specialized models handling high-frequency decisions, with slower reasoning models stepping in only when the situation gets ambiguous.

For robotics, gaming, and real-time control systems — domains where decisions must arrive in milliseconds and budgets are measured in cents — that architecture is appealing. A warehouse robot, a game-playing NPC, or a trading system cannot afford a 2-second GPT-style round trip for every micro-action.

Skeptics on Hacker News pointed out the run's caveats: the game is decades old and thoroughly documented, the coaching model did the strategic heavy lifting, and 15 Elite Four attempts suggests plenty of trial and error. Those are fair points — but they miss the more interesting signal. Sixteen thousand good decisions at $0.0001 each is exactly the cost profile autonomous agents need if they are ever to run at scale.

TypeSafe AI, founded by a former OpenAI RLHF researcher, has positioned Jev as a complement to language models rather than a replacement. The Pokémon project — code, stream, and cost breakdown all public — is the most vivid demonstration yet of what a decision-first stack can do.

The full source code is available on GitHub, and the completed run can be watched on YouTube, including every wipe on Victory Road.

Stay Ahead of AI

Follow AI Buzz Wire for the latest AI developments.

Read more AI news →