Arm C2 CSS diagram

Arm Slows Mobile CPU Introductions


Arm has squeezed more performance from its Lumex Ultra CPU; however, Arm’s new GPU advances the company’s mobile platform further by implementing significant hardware changes. At the same time, Arm has introduced the Neoverse N4 for servers and high-end embedded chips.

Because Arm’s smartphone-chip licensees update their flagship processors annually, Arm revises its designs (IP) at the same pace. It’s a rapid cadence for CPUs and GPUs, which requires months of architecture exploration for major changes, followed by months of implementation and validation. This year, Arm updated only its big Ultra CPU, eking out a small per-cycle throughput gain. The GPU, however, received a major upgrade. Neoverse customers have longer product cycles, and the N4 is the first new Neoverse release since the V3, N3, and E3 refresh a year and a half ago.

How Does the Arm Mali G2-Ultra NX GPU Advance Mobile Graphics?

The Arm Mali G2-Ultra NX GPU integrates CNN engines to accelerate frame interpolation, upscaling, and ray-tracing denoising. It also features an expanded register set and triangle-strip processing for ray tracing to boost gaming performance by 14%.

  • Neural stuff has come to Arm GPUs, fulfilling the plan disclosed last year.
  • Mali G2-Ultra NX is the name of the new neural-enabled GPU.
  • Neural algorithms include:
    • Super sampling—inferring a higher-resolution image from a lower-resolution one to reduce execution time, power, and needed computing resources to create the higher-res image.
    • Upscaling—inferring frames to raise frame rates without additional hardware and power.
    • Super sampling and denoising—inferring pixels in ray-traced scenes to raise quality.
  • Accelerating neural networks—to efficiently execute these algorithms, Arm has added convolutional neural network (CNN) engines to its shader cores. Principally, these operate on eight-bit integer (INT8) data.
  • Ray tracing improves in the new GPU. Last year’s G1 added hardware for traversing the boundary-volume hierarchy. The G2 streamlines storage and processing, eliminating the processing of redundant data by handling triangle strips (chains of triangles sharing vertices) instead of individual triangles.
  • The execution unit gains a bigger register set, resulting in a major change to the GPU instruction set. Game engines benefit from more registers by keeping more data local, reducing time- and energy-consuming spills to memory while increasing GPU resource utilization.
  • Gaming performance increases by 14%, and the benchmark uplift is even greater.

What Performance Gains Does the Arm Lumex C2-Ultra Mobile CPU Offer?

The Lumex C2-Ultra CPU achieves a 7% per-cycle performance (IPC) boost. Combined with a higher clock rate, single-thread performance increases by 15%. Dual SME2 units scale AI throughput up to 1.7× that of C1-based single-unit configurations.

  • No updates for the little and midsize cores (Lumex Nano and Pro CPUs) arrived this year. The little cores have been on an idiosyncratic multiyear upgrade cycle, but this is the first time that the midsize (formerly known as the big) core has skipped a year. However, Arm and its customers may apply the C2 name (e.g., Lumex C2-Pro) to these older CPUs to indicate they’re still generationally synced to the Ultra.
  • IPC gains (per-cycle performance boosts) are modest for the Lumex C2-Ultra, at approximately 7%. By contrast, preceding generations have achieved double-digit IPC gains from microarchitectures sharing the same quantity of decoders and execution resources.
  • CSS performance climbs by 15% on single-thread tests and 12%on multithread tests when comparing C2 to C1 configurations comprising two Ultra and six Pro cores. The gains reflect Ultra’s greater IPC and a higher clock rate enabled by newer process technology. CSS is Arm’s name for its hard macro products.
  • AI performance increases by 1.7× when comparing a C2 design featuring two SME2 units to a C1 implementation with only a single enabled unit.

What Improvements Does the Neoverse CSS N4 Offer Infrastructure Chips?

The infrastructure-targeted Neoverse N4 reaches 3.8 GHz and supports up to 128 cores per die when paired with the CMN S4 mesh. The Neoverse CSS N4 improves system performance with LPDDR6 and MRDIMM memory and PCIe 7.

  • Neoverse N4 details are sparse. The N4 (Dionysus) is the C1 Lumex-Pro (Gelas) enhanced for server and high-end embedded designs.
  • Clock rate climbs to 3.8 GHz compared with 3.5 GHz for the C1 Lumex-Pro, assuming a 3 nm CSS implementation in both cases. Arm may have optimized critical paths or could have simply constrained power less, reflecting the N4’s differing applications.
  • System-level features account for most of the advantages of the CSS N4 compared with previous Neoverse N-series CSS offerings:
    • CMN S4—a new core mesh network scales to 16 × 16 grids, up from 12 × 12 with the CMN S3. The bigger mesh and other enhancements double bisection bandwidth.
    • Core count grows to 128 in the maximally configured CSS. The CSS N2, by contrast, supported a maximum of 64 cores per die. (Arm did not commercialize a CSS N3.) However, the older CMN already supported 128 cores.
    • DRAM support in the new CSS is greatly improved, interfacing to LPDDR6 and MRDIMM memories in addition to regular DDR5.
    • PCIe 7 quadruples bandwidth compared with the CSS N2’s PCIe 5. Data-center NPUs such as Amazon Trainium are already employing PCIe 6, and the newer interface would support next-generation accelerators.

How Is MediaTek Implementing the Arm Lumex C2-Ultra and Mali G2 Architectures?

MediaTek integrates Lumex C2 CPUs and the Mali G2 GPU into its Dimensity 9600 Pro SoC, featuring two 4.55 GHz C2-Ultras and six C2-Pro cores. Pairing this layout with the Mali G2-Ultra NX GPU yields an 18% ray-tracing performance uplift.

MediaTek is Arm’s biggest mobile IP customer and an early adopter of Arm’s newest designs. The chipmaker expects the first phones employing the Dimensity 9600 Pro based on the new Lumex CPU and Mali GPU to launch this quarter. MediaTek previously licensed Arm’s soft IP, not CSS hard macros. It has demonstrated a remarkable ability to achieve the industry’s best power, performance, and area (PPA) with Arm CPUs. We expect the new Dimensity to continue this practice. Its configuration differs from the exemplary CSS, comprising two 4.55 GHz C2-Ultras, three 4.35 GHz performance-optimized C2-Pros, and three 3.1 GHz power- and cost-optimized C2-Pros. The differing C2-Pro configurations reflect MediaTek’s skills in backend design. The Dimensity 9600 Pro also integrates the Mali G2-Ultra NX GPU, achieving 18% higher scores than the earlier Dimensity 9500 on the 3Dmark Solar Bay Extreme benchmark, which exercises ray tracing.

Bottom Line

Every microarchitecture eventually reaches its maximum potential, beyond which incremental design changes yield only small performance gains. The Lumex Ultra may have reached this point. If so, we expect a major overhaul next year. At the same time, CPU performance is already sufficient for users other than twitch gamers seeking a game-console-like experience. For them, GPU improvements come first, and the G2 upgrades ray tracing and accelerates upscaling and frame interpolation. Desktop gamers, who always have the option to buy a more expensive, higher-power card, have derided these features. In smartphones, however, they should deliver a better perceived experience within a mobile power budget, but it’s up to game developers to manage their application to maintain visual quality.

For infrastructure customers, a Neoverse V4 was conspicuous by its absence. A few customers previously employed the Neoverse N series in server processors. However, we expect this application to exclusively employ V-series cores, despite their worse area efficiency, because of their greater instruction throughput. Although Arm is gaining share in data-center CPUs, and these chips are entering the market on an annual cadence, the company surprisingly hasn’t introduced a Neoverse V4. The N cores, however, will remain important for DPUs, embedded processors, and other chips that require a balance between performance and cost/power that favors the latter.


Posted

in

by


error: Selecting disabled if not logged in