Announcement

LLM Inference on New Silicon in Days, Not Years: d-Matrix Meets Infinity

Ishan Paidhungat
Ishan PaidhungatAuthor
Sravya Tirukkovalur
Sravya TirukkovalurContributor
Jeremy Nixon
Jeremy NixonContributor
Luke Bechtel
Luke BechtelContributor
April 30, 2026
Running Qwen3 on the Infinity d-Matrix Cloud, powered by the Corsair

Infinity has created Ignition, a research agent capable of implementing optimal inference on new hardware.

Ignition generates the entire inference stack for any chip. On traditional hardware, it already surpassed vLLM with a fully generated inference stack. Now we've done the same on completely new silicon: the d-Matrix Corsair.

Crossing the CUDA Gap

Groq and Cerebras were dormant for years. Then they implemented versions of CUDA that worked for their chips, and both surged 5x to 6x once inference at scale became possible. The difference was software.

~$5B $30B+

Groq

Language Processing Unit
Before

~$1B

2021 unicorn

After

$6.9B

Sep 2025 · +590%
Groq LPU
2M+ developers within 18 months of GroqCloud's 2024 launch

Cerebras

Wafer-Scale Chip · Inference Cloud
Before

~$4B

2021 Series F

After

$23B

Feb 2026 · +475%
Cerebras WSE
Catalyst: OpenAI partnership · 750MW ultralow latency compute

Groq went from ~$1B in 2021 to $6.9B after launching GroqCloud. Cerebras went from ~$4B to $23B after its OpenAI partnership made large-scale inference real. Enabling inference software bumped these chip providers' valuations by billions. In total, ~$5B became $30B+.

Infinity is building that software layer for the next generation of AI chips. Our collaboration with d-Matrix, an AI hardware startup which currently has a $2B valuation. What will their valuation be next?

The Corsair: From Silicon to Inference

d-Matrix sets the bar for what a well-engineered hardware platform looks like. Like Groq and Cerebras, they built silicon specifically for AI inference. The Corsair is built around 2 GB of high-bandwidth on-chip SRAM, a fundamentally different design choice from GPU memory hierarchies, purpose-built to eliminate the memory bottlenecks that dominate inference workloads.

d-Matrix Corsair chip

The Corsair has its own instruction set and programming model. Aviator, d-Matrix's SDK, includes a low-level kernel and model authoring toolchain, a runtime stack, and a Corsair-interpretable ISA specification. It exposes the chip's compute primitives, memory subsystem, and inter-chip communication fabric through a clean, well-documented interface.

Programs compile down to QJSON, a proprietary instruction format. This is a low-resource language and is not in the training dataset for any LLM. None of this exists anywhere outside d-Matrix; but that's no problem for Ignition.

10 Hours: Matrix Multiplication at Scale

Within 10 hours of getting access to the hardware, Ignition had tensor-parallel matrix multiplications running across all 32 compute units on the Corsair card, spanning both packages with cross-chip PCIe synchronization, at every shape required for large language model inference.

It didn't port an existing library. Using the chip's instruction set and runtime, it wrote all of the compute operations from scratch and figured out how to orchestrate them across the hardware. Tiling strategies, sharding across compute units, on-chip SRAM data residency, memory staging pipelines, cross-package data movement, multi-unit synchronization. All targeted at hardware it had never seen before.

Once enablement was in place, Ignition kept optimizing. An automated performance sweep across a variety of tiling strategies and resource configurations found the optimal approach for every target workload. The result: up to 92% Speed-Of-Light. Matrix multiplications running within 10% of the chip's empirical compute peak.

Table 1 · BF16 GEMM, d-Matrix Corsair · n = 16 · 1167 MHz
Shape (M×K×N)
Gangs
TFLOP/s
SOL
1024×1024×8192
16
606.0
55.0%
1024×4096×4096
16
730.7
66.4%
1024×8192×1024
16
586.5
53.2%
1024×8192×8192
16
959.3
87.1%
2048×1024×8192
16
776.3
70.5%
2048×4096×4096
16
842.1
76.5%
2048×8192×1024
16
753.7
68.4%
2048×8192×8192
16
1005.1
91.3%
4096×1024×8192
32
1627.6
73.9%
4096×4096×4096
16
914.7
83.1%
4096×8192×1024
16
877.7
79.7%
4096×8192×8192
32
2026.8
92.0%
8192×1024×8192
32
1318.1
59.8%
8192×4096×4096
16
957.0
86.9%
8192×8192×1024
32
1808.8
82.1%
16384×4096×4096
32
1929.0
87.6%
Table 1. Best per-shape BF16 GEMM throughput on the d-Matrix Corsair. Speed-of-Light (SOL) is the fraction of the chip's achievable compute peak that the kernel actually realizes, measured against an empirical ceiling of 68.83 TFLOP/s per gang (90% of theoretical peak). 100% SOL means saturating the hardware.

For context, real-world inference workloads on NVIDIA GPUs routinely leave significant peak performance on the table, even with mature software. Our agent hit 90%+ on hardware with no existing inference ecosystem.

10 Days: End-to-End LLM Inference

Matrix multiplications are the backbone of AI inference, but a working LLM needs much more: normalization, positional encodings, multi-head attention, activation functions, vocabulary projection, and the orchestration to chain them across dozens of transformer layers. Several of those operations live outside the matmul engine, on the chip's SIMD vector units.

Ignition got SIMD compute working end-to-end on the Corsair, generating vector kernels for every non-matmul operation in the pipeline. In 10 days, it had Qwen3, Qwen3.5, and Gemma4 running end-to-end on the Corsair. All transformer layers. Real weights. Real text in, real text out.

Our agent wrote every layer of the pipeline from scratch, including custom compute kernels where none existed. Each operation was validated against the original model's outputs, and composed into a full pipeline that produces the correct next token.

The system is live in internal preview as the Infinity d-Matrix Cloud, running inference requests on Corsair silicon.

Running Qwen3 on the Infinity d-Matrix Cloud, powered by the Corsair

What's Next for the Corsair?

On traditional hardware, our agent wrote an inference engine from scratch and autonomously optimized it to deliver up to 34.3% more tokens per second than vLLM. We are now using the same techniques to optimize inference for the Corsair.

With thanks to Sree Ganesan, Arash Fayyazi, Aseem Bathla, Max Sbabo, Ramya Ramachandran, Satyam Srivastava, and Sayantan Sarkar from d-Matrix for their guidance, tooling support, and collaboration throughout this work.

Ignite Your Hardware

Ignition can be deployed on any new chip architecture. It generates the entire inference stack, optimizes for peak throughput, and delivers production-ready software in days, not years.

Contact us