NextSilicon Arbel chip with exposed die

NextSilicon Arbel Targets RISC-V Servers


NextSilicon, a startup focusing on processors for high-performance computing, is developing a server-class 64-core CPU chip called Arbel. A RISC-V design based on a core NextSilicon employed in its Maverick-2 data-flow compute accelerator, Arbel is due to ship in early 2028. A development chip is already in customers’ hands.

The RISC-V architecture has been mooted as an alternative to Arm Android-based smartphones, automotive electronics, and hyperscalers’ servers. It has modest traction in HPC, where the European Union is funding a project. Meanwhile, several companies license high-end RISC-V cores, but they have disclosed few customers.

Against a crowded RISC-V IP market, NextSilicon is entering the server-processor market with a complete 64-core design rather than a licensable core. Delivering a standalone general-purpose CPU risks spreading the startup thin alongside its HPC-focused Maverick-2 chip. However, that chip integrates a version of the Arbel CPU to handle scalar code. Leveraging that investment by redeploying the core in a standalone CPU provides NextSilicon a path to capture additional designs—provided the design delivers competitive performance.

NextSilicon Arbel Use Cases, RISC-V Architecture, and Performance

Arbel can complement the Maverick-2 computing accelerator by performing host functions, including preprocessing and transferring data. It can also be a server CPU, running diverse workloads. NextSilicon began developing the chip after having completed the core it’s based on. Already built into the Maverick-2 chip to execute serial code, the core required little adaptation for use in a server processor. NextSilicon developed a CPU from scratch because a licensed design would’ve required extensive rework to integrate with the Maverick-2 data-flow engines, which handle parallel code.

Arbel complies with the RVA23 RISC-V profile, the standard for RISC-V application-processing designs. NextSilicon estimates that an Arbel core will score 2.6 per GHz on SpecInt2017 and 3.4 per GHz on SpecFP2017. However, most licensors of RISC-V designs (IP) rate their cores on the older SpecInt2006. On the basis of NextSilicon figures, we estimate that the company’s CPU will score about 24 per GHz on SpecInt2006, putting its performance near that of the Akeana 5300, Arm Neoverse-v3, and Tenstorrent Ascalon-X. Implemented in a 5 nm process, Arbel targets 3.4 GHz. For comparison, the 3 nm Arm AGI processor operates at a 3.2 GHz base frequency and can boost to 3.7 GHz. The lower clock speed helps reduce power. NextSilicon rates Arbel at 250 W (TDP), below the Arm chip’s 300 W level.

NextSilicon Arbel Microarchitecture: Scalar Back End and Vector Units

Like most high-performance designs, the Arbel CPU implements a wide microarchitecture. The front end is 10 instructions wide, tying the Arm Lumex C1-Ultra and Nvidia Olympus (Vera) for the widest on the market. The scalar back end comprises 16 function units, as Figure 1 shows. Eight are integer units, including three supporting only basic ALU operations and five that also handle multiplication. To conserve area, only two of the latter support division, an uncommon and silicon-intense operation. The Lumex C1-Ultra and Olympus have similar scalar back ends. Like the Arm core, the Arbel CPU has three branch units. Whereas the Lumex C1-Ultra has four dual-function load/store units, Arbel separates the functions. Olympus also has separate load and store pipelines, and it has one additional branch and one additional load unit. A dual-thread design, Olympus may need the added hardware to sustain performance on both of its threads.

NextSilicon Microarchitecture Diagram
Figure 1. The NextSilicon Arbel CPU core implements a wide microarchitecture including 16 scalar and 3 vector units.

Accounting for its strong floating-point performance, each Arbel core has three 256-bit vector units. As with the scalar units, NextSilicon only endowed two of the units with dividers. Note that RISC-V has flexible mappings between logical and physical vector-unit widths; therefore, some implementations have different-width register files and data paths. In Arbel’s case, both physical vector registers and data paths (VLEN and DLEN) are 256 bits.

Each core includes a relatively large 2 MB private second-level (L2) cache, and the Arbel chip also has a 128 MB L3 cache. These capacities are in line with competing products. Prefetch engines help warm up L2 and L3 caches. Arbel will support DDR5 DRAM or LPDDR6, which JEDEC standardized in 2025 and should be commercially available in 2028. Memory expansion will be available through CXL 3.0 interfaces, and Arbel will also support PCIe Gen 6.

Bottom Line

Attitudes toward RISC-V for application processing have moved inversely to those toward Arm. Adoption of the Arm architecture by hyperscalers and new server-processor entrants Nvidia and Qualcomm has cooled interest RISC-V. At the same time, the number of developers of server-class RISC-V cores has grown. Therefore, no matter its performance, power, or cost, Arbel is entering the market at an inauspicious time. Maverick-2 is complementary, but not yet widely adopted. A single hyperscaler or HPC win, however, could make NextSilicon a viable CPU supplier for data-center and HPC customers.


Posted

in

by


error: Selecting disabled if not logged in