Infinity logo

Your Hardware. Frontier AI Performance.

Infinity delivers Day 1 model enablement for any AI accelerator — across every modality. Your hardware specs are the input; a production-ready inference library is the output.


Research

Research & Technical Deep Dives

How Infinity brings frontier inference to new silicon — announcements, case studies, and papers from the team.

Technical Deep Dive
Performant Full-Model Inference on d-Matrix Corsair: 20x Speedup

Infinity and d-Matrix worked hand in hand to optimize the performance of Qwen 3 by 20x tokens / second on a single Corsair card compared with Infinity’s original implementation. Context length is also extended by 16x.

July 2026

Read deep dive

Partnership Announcement
Infinity × d-Matrix: Partnering to Advance Performant Full-Model Inference

Infinity and d-Matrix have partnered to enable and optimize performant full-model inference on d-Matrix Corsair — bringing Qwen3 from its first tensor-parallel matrix operations to complete, stateful inference on a single Corsair card.

July 2026

Read announcement

Announcement
Igniting d-Matrix: LLM Inference on New Silicon, in Days, Not Years

Within 10 hours we had matrix multiplications at 90%+ of the chip’s theoretical peak. Within 10 days, Qwen3 was running end-to-end—every operation written from scratch.

April 2026

Read announcement

Case Study
Hacker News front page
Surpassing vLLM with a Generated Inference Stack

Infinity's infy optimization system wrote an inference engine from scratch and autonomously optimized it on Qwen3-8B. The resulting engine delivers up to 34.3% more tokens per second than vLLM when configured with identical parameters.

March 2026

Read research

Research Paper
ICLR 2026
OMEGA: Optimizing Machine Learning by Evaluating Generated Algorithms

In order to automate AI research we introduce a full, end-to-end framework, OMEGA: Optimizing Machine learning by Evaluating Generated Algorithms, that starts at idea generation and ends with executable code. Our system combines structured meta-prompt engineering with executable code generation to create new ML classifiers. The OMEGA framework has been utilized to generate several novel algorithms that outperform scikit-learn baselines across a robust selection of 20 benchmark datasets (infinity-bench). You can access models discussed in this paper and more in the python package: pip install omega-models.

March 2026

View paper


Careers

We're Hiring

We hire exceptional engineers and researchers who aspire to 100x impact. We currently have 7 open roles, on-site in San Francisco.