Skip to content
← All news
4 min read

Qualcomm's 2nm Snapdragon chips are built to run AI agents on the phone

Two flagship chips, a new NPU block for transformer workloads, and a claim that the top chip can run a 30-billion-parameter model. What that changes for apps that send user data to the cloud.

Qualcomm says its top phone chip can run a 30B-parameter model. The limit is memory, not silicon.

Qualcomm announced two flagship phone chips at Snapdragon Summit in Maui on September 22, 2026: the Snapdragon 8 Elite Extreme Gen 6 and the Snapdragon 8 Elite Gen 6. Both are built on a 2nm process, which Android Authority reports is TSMC's. Both are pitched around AI agents running on the phone itself.

The numbers, per 9to5Google and GSMArena: two prime CPU cores up to 5.0GHz and six performance cores up to 4.0GHz, the first 5GHz mobile CPUs. For the Extreme, Qualcomm claims 13 percent more CPU performance, 44 percent more GPU and 35 percent more NPU than last year's chip. Memory support goes up to 24GB of LPDDR5X or LPDDR6. Xiaomi launched the 18 Pro Max on the Extreme and the 18 Pro on the standard chip in China the next day.

The 30B claim, read carefully

The Hexagon NPU gets a new block Qualcomm calls the Element Accelerator, aimed at transformer and agent workloads, with context windows up to 32K tokens. The Extreme adds 50 percent more NPU shared memory, which Qualcomm says enables mixture-of-experts models of 30B+ parameters. That is Qualcomm's claim. Nobody has published independent numbers yet for tokens per second, battery cost, or how much quality survives the compression needed to fit such a model in phone memory.

Memory is also the constraint. Android Authority frames the launch against a RAM shortage, and Implicator cites a Counterpoint forecast of a 14 percent drop in 2026 smartphone shipments because of memory costs, the same squeeze that pushed OnePlus out of the US and Europe. A chip that can run a 30B model helps only in a phone that ships with enough RAM to hold one.

Why a build studio cares

On-device inference changes the first question in any data-flow review: where does the user's data go? A feature that runs on the phone never sends its prompt to a cloud API, which simplifies data residency. It also gets harder to verify. A cloud call shows up in network logs. A local model does not, and an app can still fall back to the cloud when the phone runs short of memory. When a vendor tells a client a feature is on-device, the check is simple: run it in airplane mode on the lowest-spec phone they support, then watch the network traffic when the phone comes back online.

Next step: read 9to5Google's rundown for the full spec list. If a vendor sold you an on-device AI feature and nobody has tested that claim, write to us at hello@gattyworks.com.

QualcommOn-Device AIHardwareMobileQualcommSnapdragonSnapdragonSummitXiaomiOnDeviceAIAIAgentsEdgeAIAndroidSemiconductorsMobileDev

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.
Book a call