A small but striking project is circulating on Hacker News this week: openTPU, an open-source AI accelerator whose hardware design was itself produced by AI agents. The repository, published under the Apache 2.0 license on GitHub, describes what its authors call "an open-source AI accelerator, developed by AI" — and unlike most demos in this genre, it runs real production models with their real weights on a physical card that enthusiasts can study end to end.

The project asks two questions, in the words of its own README: how far can AI agents go at hardware design, and can they build the chip that runs their own inference? The second question is the provocative one. An AI system designing the silicon — or in this case, the FPGA bitstream — that generates its own tokens is a loop that captures exactly where the industry's imagination is heading. For more context on this story, see our ongoing breaking AI news.

What Actually Runs on the Card

The hardware target is unglamorous by data-center standards: an Inspur YPCB-00338 card built around a Xilinx Kintex-7 xc7k480t FPGA with two DDR3 channels and a 17.1 GB/s peak memory bandwidth. That is a far cry from an HBM-stacked accelerator, which makes the results more interesting, not less.

According to the project's published measurements, the design runs ten modern open models with their real weights. LFM2.5-230M decodes at 59 tokens per second in int8 and 85.8 tokens per second in 4-bit. Qwen3-0.6B manages 21.6 tokens per second, while larger models scale down as expected: SmolLM3-3B at 5 tokens per second int8, Phi-4-mini (3.8B) at 4, and Qwen3.5-4B at 5.9 tokens per second in 4-bit. Gemma 4 E2B and E4B also run, keeping their per-layer embedding tables resident on the card.

The team went further with mixture-of-experts models that exceed the card's 4 GiB of DRAM. LFM2.5-8B-A1B — 8.5 billion parameters with 1.7 billion active — decodes at 10.6 tokens per second by streaming experts from host storage, with 98.5 percent of expert uses hitting on-card slots and only 5.2 MB transferred per token. Qwen3.5-35B-A3B runs at 3.95 tokens per second while streaming 153 MB per token over PCIe. Every configuration, the README states, matches the project's simulator token for token — a bit-exactness claim that hardware hackers will recognize as the hard part of the job.

Built Like a Real Silicon Project

What separates openTPU from typical hobby FPGA demos is the completeness of the stack. The monorepo contains the SystemVerilog hardware design, a custom instruction set, a bit-exact simulator, a kernel language with its own compiler, and the host software that drives the PCIe card, including an otpu-smi monitoring utility in the spirit of nvidia-smi and an otpu-chat demo.

The production image runs a four-column systolic matrix unit and a stream engine at 133.33 MHz, with LiteDRAM memory controllers calibrated by a small CPU inside the memory core — the card calibrates both DDR3 channels in 12 seconds with no host involvement. The current build, deployed since October 1, decodes several models 8 to 10 percent faster than the previous image while pushing DRAM utilization to 91-94 percent of peak.

The project also documents a 4-bit weight format built on FP4 values with two-level block scales at 4.25 bits per weight, keeping the language-model head in int8 for accuracy. The format cuts bytes per token by roughly a third and lifts decode speed by 40 to 45 percent on some models.

The README credits the project's approach to lessons from an earlier repository, auto-arch-tournament, which explored AI-driven hardware architecture search. openTPU is the application of that method to a complete accelerator, with the agents' design validated against a simulator before ever touching hardware.

Why It Matters

The performance here will not trouble Nvidia. A Kintex-7 with DDR3 is a decade-old class of hardware, and the models it runs are small. But that is the point the project makes implicitly: the entire accelerator — RTL, instruction set, simulator, compiler, and host stack — is legible to one person, in one repository, and it was designed at least in part by AI agents rather than a hardware team.

Three currents converge in that idea. First, open-source AI hardware has long been stymied by the sheer breadth of expertise required; an agent-driven design flow lowers that barrier. Second, inference demand is pushing researchers toward exotic special-purpose hardware, and automated design search is a natural fit for exploring that space. Third, the self-hosting angle — AI building the machine it runs on — has become a signal the community watches closely, judging by the project's rapid rise on Hacker News, where it gathered more than 150 points within hours.

The repository was created on September 24 and received its latest push on October 2, with measured results dated as recently as October 1 — a development pace worth watching in its own right. For anyone who wants to understand how an AI accelerator works from a matrix multiply in Python down to the wires, the project's own framing is accurate: the whole thing lives in one small monorepo you can read end to end.

Whether agent-designed hardware becomes a serious niche or remains a research curiosity, openTPU has demonstrated something concrete: the tools that let AI models write software have now produced a working, verified hardware system that runs those same models. The loop is closed, at least at FPGA scale.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →