Qualcomm's chief executive says the biggest AI companies are asking him for something current smartphones cannot come close to doing: running 100-billion-parameter models on the phone itself, continuously, by 2028.

Speaking in an interview on October 8 that was reported by Fortune and picked up this week by outlets including TOKENPOST and Startup Fortune, Cristiano Amon said multiple AI companies — whose names he declined to give — are pushing for smartphones that can keep a 100-billion-parameter model running around the clock. For readers following the shift of AI from the cloud into pockets and wrists, it is the clearest demand signal yet, and our AI hardware news section tracks where that push goes next.

The Demand Comes From AI Labs, Not Qualcomm's Roadmap

The number matters because of where it comes from. Amon was not announcing Qualcomm's own product target; he was relaying what AI labs are telling the world's largest mobile chipmaker they want to build. Today's flagship chips top out far below that figure — the Snapdragon 8 Elite Extreme Gen 6, Qualcomm's highest-end mobile platform, can run mixture-of-experts models above 30 billion parameters on-device, with a context window of roughly 32,000 tokens.

A 100-billion-parameter model running nonstop is more than triple that ceiling, and it changes the engineering problem in kind, not just in degree. As Startup Fortune's report notes, the request is not for a bigger version of the assistant people already use. It is for agentic AI: software that watches, reasons and acts on a user's behalf instead of waiting for a prompt, which means the model has to stay resident in memory and keep doing inference in the background all day without cooking the battery or the chassis.

Why Memory Is the Bottleneck

The gap between 30 billion and 100 billion parameters is fundamentally a memory problem. On-device model weights have to live in RAM, and tripling the parameter count while also keeping inference running continuously multiplies memory bandwidth and power demands in ways current flagship hardware was never designed to absorb.

Amon acknowledged that memory chip supply for smartphones is under a "temporary constraint" — the same shortage that has driven up phone prices across the industry this year — but he framed it as pressure that will force more memory-efficient chip designs rather than simply more memory. That tracks with Qualcomm's recent silicon choices: the 8 Elite Extreme Gen 6 pairs its Hexagon NPU with a new Element Accelerator and a larger shared memory pool specifically to keep latency and power draw in check for on-device mixture-of-experts inference.

The Chips That Have to Bridge the Gap

Qualcomm's current flagship generation — the Snapdragon 8 Elite Gen 6 and 8 Elite Extreme Gen 6, built on a 2nm process — is slated to power 2027's flagship phones from Honor, iQOO, Motorola, OnePlus, OPPO, Redmi, RedMagic, vivo and Xiaomi, with devices like the Samsung Galaxy S27 class expected at the top of the range. If the AI industry's 2028 wishlist is to be met, it is this chip cycle and the one after it that have to deliver the memory architecture, thermals and battery efficiency to keep a frontier-scale model alive in a shirt pocket.

Amon also said, according to Fortune's reporting, that Qualcomm is working with nearly all of the major AI players on undisclosed hardware projects, including device categories beyond phones. That is consistent with the company's recent posture of positioning Snapdragon silicon as the default inference layer for edge AI wherever it lands — phones, cars, PCs and glasses.

What Continuous On-Device AI Would Actually Mean

If AI companies get what they are asking for, the change would be structural rather than cosmetic. A model that runs all day in the background can observe context, anticipate needs and act without a round trip to a data center — which implies lower latency, better privacy, and no per-query cloud cost. It also implies phones with larger, faster, more power-hungry memory than today's flagships carry, in a market where memory pricing is already stretched.

There is healthy skepticism to apply here. Chip executives have made bold on-device AI predictions before, and 2028 is far enough out that the claim functions partly as a negotiating position with memory suppliers, phone makers and the AI labs themselves. But the direction of the request is notable: AI companies are not asking for faster cloud GPUs this time. They are asking for a phone that never stops thinking.

Whether Qualcomm's customers ship that device in 2028 or later, the bottleneck to watch is not parameters — it is memory bandwidth per watt. The company that solves that, in Redmond or San Diego or Taipei, will own the next definition of a flagship phone.

Stay Ahead of AI

Chip wars, model launches and edge AI breakthroughs — covered without the hype.

Read more AI news →