Intel opened this year's Hot Chips conference with a deep dive into Crescent Island, the data center GPU it is betting on for the next phase of AI workloads: agentic inference. Presented by Chief Enterprise AI Systems Architect Sumit Mohan and Intel Fellow Hong Jiang, the talk centered on three metrics the company says define the agentic era - tokens per watt, large memory capacity, and an open software stack.

The timing matters. Nvidia used the same conference to detail its Vera Rubin NVL72 rack-scale system, and every accelerator vendor at Hot Chips 2026 is now pitching efficiency rather than raw peak performance. Intel's answer is a part that swaps the memory technology mainstream AI GPUs use for something cheaper per gigabyte, and it makes for striking headline numbers. For more context on this story, see our ongoing AI industry coverage.

A 480GB PCIe Card at 350 Watts

Crescent Island is a 350-watt, air-cooled PCIe GPU. Intel's branded card ships with 160GB of LPDDR5X memory, but the design allows ODM partners to build variants with up to 480GB on board - roughly triple the capacity of typical high-end accelerator cards, and enough to keep very large KV-caches resident for long-context and multi-session agentic workloads. The design covers data types from FP4 and MXFP4 up through FP64.

The memory choice is the strategic decision in the product. LPDDR5X trades bandwidth for capacity and cost, which fits an inference workload profile where enormous context windows and tool-call loops consume memory capacity more than raw compute. Intel frames the entire design around "tokens per watt" - delivering more useful output per unit of energy rather than winning peak-FLOPS comparisons.

Xe3P Silicon: 32 Cores, 256 XMX Engines

Under the hood, Crescent Island is Intel's Xe3P architecture, and Intel presented it as a generational leap over the Battlemage-based Xe2 generation. Where Xe2 offered 20 Xe cores and a 4-deep systolic XMX array, Crescent Island brings 32 Xe cores and a 16-deep systolic array, feeding 256 XMX engines in total.

The third-generation Xe Matrix Extensions add a three-way extended Xe matrix design with FP4 precision co-issue and FP64 support. Each Xe core carries a 1MB general register file and 512KB of L1 cache, with a 32MB unified L2 shared across the device. A media engine with four decoders and four encoders sits alongside, and the whole SoC connects over PCIe Gen5 x16 with support for scale-up across an open switch fabric. Intel lists an active idle power of 50 watts or less in the G0 state and roughly 10 watts in the low-power G8 state, achieved through distributed power domains, packet-based network-on-chip routers with aggressive clock gating, and independent DVFS rails for graphics and media.

RAS Features Borrowed From the Server Playbook

For a data center part, reliability engineering got as much stage time as throughput - a theme Nvidia also leaned into with Vera Rubin. Crescent Island carries ECC and parity across key memory structures, error checking on every hop of the internal interconnect fabric, dynamic page offlining, hard post-package repair, and PCI Express Advanced Error Reporting. Intel says the combination cuts down on silent data corruption, the failure mode that matters most for long-running unattended inference.

Intel also positioned the part inside a portfolio that stretches from Arc Pro workstation GPUs through the SambaNova SN50 RDUs it acquired, with Crescent Island as the enterprise data center anchor at the top. The company showed a roadmap lineage running back more than six years, casting the GPU as a mature third-generation design rather than a first attempt. The card was first unveiled at Computex in June; Hot Chips delivered the architectural detail behind it.

Why Agentic Inference Favors This Design

The workload argument is the most interesting part of Intel's pitch. Agentic AI - models that chain tool calls, hold long contexts open, and coordinate between CPU and GPU - taxes latency, memory capacity, and system throughput simultaneously, rather than the batch-throughput profile of chatbot serving. KV-cache capacity for long contexts and compressed-domain concurrent sessions is built into the SoC design explicitly.

That framing also explains the LPDDR5X bet. If the next wave of AI compute is memory-hungry agents rather than training runs, a 480GB card at 350 watts is a credible alternative to premium HBM-based parts - provided the software stack delivers. Intel's emphasis on an open stack and open switch fabric is aimed squarely at buyers wary of Nvidia's integrated platform lock-in.

Crescent Island won't displace the training accelerators that dominate AI headlines, and Intel did not disclose pricing or availability at the conference. But as a statement of where Intel thinks inference is going - fat memory, lean power, open fabric - it is one of the clearer architecture stories of this year's Hot Chips.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →