d-Matrix adopts Nvidia NVLink Fusion for its coming Raptor AI inference racks, placing the startup’s custom accelerators inside Nvidia’s MGX reference architecture. The companies expect the first Raptor-based systems to become available in the fourth quarter of 2027.

 

The arrangement gives d-Matrix a route into a rack ecosystem already familiar to data-center operators while preserving its focus on low-latency inference. It also shows how Nvidia is widening NVLink beyond all-Nvidia systems as cloud providers assemble specialized computing around different stages of an AI workload.

 

The announced system combines five main elements:

  • d-Matrix Raptor XPUs for latency-sensitive inference
  • Nvidia Vera CPUs and NVLink switches
  • BlueField-4 DPUs and ConnectX-9 SuperNICs
  • Spectrum-X Ethernet for scale-out networking
  • Astera Labs connectivity inside modular MGX trays

 

d-Matrix Adopts Nvidia NVLink Fusion for MGX Racks

d-Matrix’s September 10 announcement describes a multi-year product roadmap built around Nvidia’s rack-scale components. Its Raptor XPUs will connect through NVLink Fusion and fit into modular, cable-free MGX trays, with Astera Labs supplying custom connectivity technology.

 

Nvidia positions NVLink Fusion as a semi-custom infrastructure program for connecting third-party CPUs and accelerators to its rack-scale fabric. That matters because specialized chips still need processors, switches, networking, management software and a serviceable physical design before a data center can deploy them.

 

The agreement does not turn Raptor into an Nvidia GPU or establish that every accelerator can be exchanged without engineering work. It gives d-Matrix an integration path into a defined Nvidia architecture, reducing the number of surrounding rack components that the startup and its customers must design independently.

 

For Nvidia, the project extends its influence beyond systems built only from its own compute silicon. A custom accelerator connected through NVLink Fusion can still rely on Nvidia processors, networking and switches, keeping much of the rack architecture within Nvidia’s platform while allowing a partner to supply specialized compute.

 

How Raptor Splits Prefill From Decode

d-Matrix is targeting the decode phase of generative AI inference, when a model produces output tokens after processing the prompt. The company says this phase benefits from high memory bandwidth and predictable latency, especially for coding assistants, chatbots and voice agents that must respond interactively.

 

Its proposed heterogeneous rack can assign compute-heavy prefill work to GPUs and latency-sensitive decode work to Raptor XPUs. That separation is increasingly relevant as operators try to improve utilization rather than run every stage of a request on the same type of processor.

 

The architecture is a deployment proposal, not a published performance verdict. d-Matrix did not release customer benchmarks, rack pricing or financial terms with the announcement, and independent comparisons will be needed once production systems are available.

 

More on This Story

 

Raptor’s 3D DRAM Architecture

Raptor is the planned successor to d-Matrix’s production Corsair accelerator. The company describes Raptor as a two-story package that stacks a DRAM memory chip above an SRAM compute chip, bringing stored model data closer to the arithmetic used during inference.

 

That design attacks a central inference bottleneck: moving model weights between memory and compute for every generated token. More local bandwidth can reduce data movement, but the practical result will depend on software, model size, power use, yield and sustained performance under real workloads.

 

d-Matrix says technical details have been presented through IEEE and Hot Chips 2026 and that Raptor is being evaluated by hyperscalers and frontier AI laboratories. Those evaluation claims come from the company; no named customer commitment was announced.

 

Q4 2027 Availability Sets the Test

d-Matrix expects to tape out Raptor before the end of 2026 and targets initial Raptor-MGX rack availability in the fourth quarter of 2027. Tapeout marks completion of a chip design for manufacturing, not commercial shipment, leaving fabrication, packaging, validation and system qualification still to follow.

 

Reuters independently confirmed the design and availability timetable and reported that financial terms were not disclosed. The long lead time gives d-Matrix and its partners room to finish the silicon and rack integration, while giving customers a concrete window for planning tests.

 

The commercial test will be whether Raptor can deliver meaningful latency, throughput and efficiency gains without adding excessive operational complexity. NVLink Fusion can make the surrounding system more familiar, but d-Matrix must still prove its accelerator, software stack and split-compute approach under customer workloads.

 

If the schedule holds, the first racks will arrive more than a year after tapeout, reflecting the work required to convert new silicon into qualified infrastructure. The announcement establishes a credible integration route; measured results, named deployments and production volume will determine whether it becomes a significant alternative for AI inference.