The National Security Agency is spending billions of dollars this year testing advanced AI models, with computing power driving most of the expense, according to classified estimates described to Congress.
The Washington Sun reported the figures, citing two sources familiar with the classified estimates. The disclosure has sharpened a debate on Capitol Hill over who should pay for oversight of increasingly capable AI systems — and just how expensive serious government scrutiny has become.
For ongoing coverage of AI oversight and regulation, see the latest AI policy news.
Compute Is the Biggest Bill
According to one of the sources cited by The Washington Sun, computing power is the single largest expense in the NSA's AI testing operation. Staffing costs add to the total, because the agency has to compete with the pay packages offered by frontier AI labs for a small pool of experts capable of evaluating cutting-edge models.
The sums involved are drawing attention because earlier public estimates of government AI oversight were dramatically lower. The Congressional Budget Office pegged the cost of a bipartisan bill to create a center for AI risks at roughly $20 million per year. Lawmakers now expect that full-scale AI oversight could cost tens of billions of dollars annually, according to the sources.
Why Testing Frontier Models Costs So Much
The economics of AI testing mirror the economics of using AI at scale. Consumer subscriptions to AI products are heavily subsidized: analysis firm SemiAnalysis has estimated that a maxed-out $200-per-month ChatGPT Pro subscription is worth up to $14,000 at API list prices, while Anthropic's similarly priced Claude Max plan is worth up to roughly $8,000.
Business and government customers, by contrast, pay closer to real costs. OpenAI has shifted new enterprise contracts to token-based billing, and at Anthropic, an estimated 75 to 85 percent of revenue comes from usage-based contracts, per SemiAnalysis. Agentic workflows — exactly the kind of autonomous behavior safety testers most need to probe — can consume up to a thousand times more tokens than a standard chat conversation.
Providers are experimenting with ways to tame those costs for large customers. According to The Information, OpenAI now lets some big clients pay only for completed tasks rather than raw tokens, and has introduced GPT-6 Sol and Luna, two models priced at half the usual API rate. But those discounts do not apply to the most powerful frontier systems, which remain the most expensive to run — and providers increasingly gate their most capable models behind closed cybersecurity programs, adding another layer of cost and complexity for any outside evaluator.
The spending pressures run in both directions. The Financial Times has reported that OpenAI plans to spend about $856 billion on computing power through 2030, while Anthropic is preparing for an IPO expected by November at the latest. SemiAnalysis estimates the gross margin on Anthropic's API business exceeds 80 percent — margins that help fund the frontier development race the NSA is now trying to independently evaluate.
Who Pays for Oversight?
The question of funding has already split into competing proposals. Nat Purser of the AI Verification and Evaluation Research Institute has proposed a levy on AI developers that would help finance independent government testing — an approach that would formalize what industry already does voluntarily.
Notably, The Washington Sun reported that Anthropic and OpenAI have signaled they might be willing to pay for expanded government oversight. That position aligns with both companies' public support for third-party safety evaluations, and would spread costs across the industry rather than concentrating them in the federal budget.
The Broader Oversight Landscape
The NSA spending revelation lands amid a broader restructuring of how the U.S. government evaluates AI systems. The agency has previously advocated for voluntary testing access to commercial models, and national security agencies have grown increasingly central to AI policy as frontier systems gain cyber and biological capabilities that regulators struggle to assess independently.
The price tag also illustrates an uncomfortable reality for policymakers: the cost of meaningful oversight scales with the cost of the systems being scrutinized. A budget built for $20 million-a-year reviews is ill-equipped for models whose training and evaluation runs cost hundreds of millions of dollars each.
As Congress weighs reauthorization of AI oversight programs, the classified estimates give lawmakers a concrete — and sobering — data point about what serious evaluation actually costs.
Stay Ahead of AI
Government spending on AI oversight is set to become one of the defining policy fights of the coming year. For more breaking AI news, from regulation to frontier model launches, follow AI Buzz Wire.
Read more AI news →