AI Inference Companies in 2026
AI inference companies build the products that run trained AI models in production and return outputs (tokens, images, embeddings) quickly and at low cost. The category covers inference APIs and platforms that serve open and custom models at low latency, inference engines, and chips or systems designed for inference rather than training.
AI Inference Companies to know
These companies were chosen because inference is their core product, either as an API or platform for serving models, an inference engine, or silicon and systems built for inference. They are selected using notability as of October 2026, using public-market or private valuation, the size of the latest funding round, reported revenue or adoption, and independent press coverage. Companies known mainly for general GPU rental (neoclouds), model hubs or broad MLOps were left to sibling directories. Together AI is the one exception: it stays here because it is a top player in both inference and neocloud.

Cerebras Systems
Wafer-scale AI chips and a cloud inference service built on them
Cerebras designs the Wafer-Scale Engine (WSE-3), an AI processor built on a single wafer. It sells CS-3 systems and offers inference and training cloud APIs, positioning wafer-scale hardware as a lower-latency alternative to GPU clusters for serving models. Customers named in IPO coverage include OpenAI, G42, MBZUAI and AWS.
- Headquarters
- Sunnyvale, California, USA

Fireworks AI
Inference and model-serving platform for open and customized models
Fireworks AI provides a platform to deploy, customize and serve generative AI models, with 200+ models across text, image and multimodal. It was founded by former Meta engineers, including CEO Lin Qiao, who previously led Meta's PyTorch team. In July 2026 it reported annualized revenue above $1B.
- Headquarters
- San Mateo, California, USA

Groq
LPU-based inference cloud running from its own data centers
Groq developed the Language Processing Unit (LPU), an ASIC designed for inference, and runs GroqCloud, an inference service used by developers. In December 2025 Nvidia signed a non-exclusive license to Groq's technology, in a deal valued at about $20B, and hired founder Jonathan Ross and president Sunny Madra. Groq stayed independent and runs its inference cloud from 13 data center locations, per TechCrunch.
- Headquarters
- Mountain View, California, USA

Together AI
Open-model inference, fine-tuning and GPU clusters from one cloud
Together AI hosts open-source models behind inference APIs and also offers fine-tuning and Nvidia GPU clusters, which it positions as a lower-cost alternative to closed models. Its founders include CEO Vipul Ved Prakash and researchers Ce Zhang, Chris Re, Tri Dao and Percy Liang. TechCrunch reported annual bookings above $1.15B in mid-2026, with customers including Cursor, Cognition and Decagon.
- Headquarters
- San Francisco, California, USA

Baseten
Inference platform for deploying and scaling production AI models
Baseten provides an inference platform and systems software for running open, custom and fine-tuned models in production, covering GPUs, autoscaling, observability and billing. It manages capacity across 18 clouds and 87 clusters and says it handles more than 1 billion inference calls a day. Customers include Abridge, Cursor, Lovable and OpenEvidence.
- Headquarters
- San Francisco, California, USA

Etched
Custom inference chips and full-stack inference clusters
Etched designs inference hardware: co-designed chips, racks and software sold as frontier inference clusters for running large language models. TechCrunch reported $1B in customer orders as of July 2026, including from Jane Street, along with a 10-megawatt data center in Silicon Valley and about 400 employees. In October 2026 TechCrunch reported that it was weighing funding offers at $40B to $50B.
- Headquarters
- San Jose, California, USA

SambaNova Systems
Dataflow AI chips and systems for enterprise inference
SambaNova builds Reconfigurable Dataflow Unit (RDU) inference chips, the SN40L and the SN50 (unveiled February 2026, shipping in H2 2026), along with inference systems and cloud services. TechCrunch reported that JPMorgan Chase chose SambaNova as an inference-infrastructure partner and that SoftBank is the first deployment partner for the SN50. Intel, which was earlier reported to be in acquisition talks with SambaNova, invested in the Series F and is co-developing inference products with it.
- Headquarters
- San Jose, California, USA

d-Matrix
Memory-centric compute platform for generative AI inference
d-Matrix builds Corsair, an inference accelerator that does compute in memory to reduce latency and energy use. It also offers JetStream networking accelerators and the Aviator software stack. In 2026 it acquired Wallaroo.ai (August) and adopted Nvidia NVLink Fusion rack-scale infrastructure (September).
- Headquarters
- Santa Clara, California, USA

DeepInfra
Inference cloud serving open models via OpenAI-compatible APIs
DeepInfra runs a GPU inference cloud from eight U.S. data centers. It serves 150+ open-source models through OpenAI-compatible APIs and also offers GPU instances for inference workloads. Its founding team previously built infrastructure for the imo messenger.
- Headquarters
- Palo Alto, California, USA
Inferact
Startup founded by the vLLM creators to commercialize the inference engine
Inferact was founded by the creators of vLLM, an open-source LLM inference engine incubated in 2023 in Ion Stoica's lab at UC Berkeley, to commercialize and develop the project. The company says vLLM supports 500+ model architectures and 200+ accelerator types. Simon Mo is CEO.
- Headquarters
- San Francisco, California, USA
AI Inference Companies at a glance
| Company | Headquarters | Focus |
|---|---|---|
| Cerebras Systems | Sunnyvale, California, USA | inference chips, inference API |
| Fireworks AI | San Mateo, California, USA | inference API, open models |
| Groq | Mountain View, California, USA | inference API, inference chips |
| Together AI | San Francisco, California, USA | inference API, open models |
| Baseten | San Francisco, California, USA | inference platform, model deployment |
| Etched | San Jose, California, USA | inference chips, ASIC |
| SambaNova Systems | San Jose, California, USA | inference chips, dataflow |
| d-Matrix | Santa Clara, California, USA | inference chips, in-memory compute |
| DeepInfra | Palo Alto, California, USA | inference API, open models |
| Inferact | San Francisco, California, USA | inference engine, open source |
Frequently asked questions
What is an AI inference company?
An AI inference company's core product runs trained AI models in production quickly and at low cost. It may be a hosted inference API (Fireworks AI, Together AI, Baseten, DeepInfra, GroqCloud), an inference engine (Inferact's vLLM), or chips and systems built for inference (Cerebras, Etched, SambaNova, d-Matrix).
Which AI inference companies are the largest as of 2026?
Cerebras went public on Nasdaq (CBRS) in May 2026, raising about $5.5B at $185 a share. The largest private rounds in 2026 included Etched ($700M at $21B, August), Fireworks AI ($1.505B at $17.5B, July), Baseten ($1.5B in tranches at $13B and $11B, June), SambaNova ($1B first close at $11B, July) and Together AI ($800M at $8.3B, July).
How do inference chip companies differ from inference API providers?
Chip companies design their own silicon and systems and often sell cloud access too. Examples are Cerebras (wafer-scale WSE-3), Etched (custom inference chips and clusters), SambaNova (RDU chips), d-Matrix (Corsair in-memory compute) and Groq (LPU). API providers such as Fireworks AI, Together AI, Baseten and DeepInfra mostly run models on GPUs and compete on software, model optimization and price.
What happened between Groq and Nvidia?
In December 2025 Nvidia signed a non-exclusive license to Groq's inference technology, in a deal valued at about $20B, and hired founder Jonathan Ross and president Sunny Madra. Groq stayed independent under CEO Doug Wightman, confirmed a $650M raise in June 2026 led by Disruptive and Infinitum, and runs GroqCloud from 13 data center locations.
How is AI inference different from model hosting or GPU clouds?
Inference specialists focus on the speed and cost of serving model outputs, using faster runtimes, custom chips or tuned APIs. Model-hosting platforms focus on deploying whatever model you bring, and GPU clouds and neoclouds rent raw compute. Some companies span more than one of these; Together AI, for example, offers inference APIs and GPU clusters, and TechCrunch describes it as a neocloud.
Which open-source project underpins many inference platforms?
vLLM, an open-source LLM inference engine incubated at UC Berkeley in 2023. Its creators founded Inferact, which raised a $150M seed round at an $800M valuation in January 2026, co-led by Andreessen Horowitz and Lightspeed, to commercialize it.
More AI company directories
- Voice AI CompaniesElevenLabs, Deepgram, PolyAI and more
- AI Agent CompaniesSierra, Harvey, Decagon and more
- AI Infrastructure CompaniesDatabricks, Scale AI, Surge AI and more
- Healthcare AI CompaniesOpenEvidence, Abridge, Tempus AI and more
- AI Coding CompaniesCursor (Anysphere), Cognition, GitHub and more
- AI GPU Cloud CompaniesCoreWeave, Lambda, Modal and more
About this directory: OpenCurious curates this list for researchers, founders, operators and buyers. Company descriptions and headquarters are based on publicly reported information as of and change often; figures a company reports about itself are its own claims. This is an editorial selection, not a paid ranking: no company paid to be listed, and companies appear in no particular order. Company names and logos are trademarks of their respective owners and are used only to identify each company. See our editorial policy. To suggest a company or a correction, email hello@opencurious.com.
The OpenCurious Newsletter
Stay curious. Get the best of the open web.
New directories, researcher tools, and open-source finds — straight to your inbox. No spam, ever.
- New directories
- Researcher tools
- Open-source finds
Occasional emails from OpenCurious. Unsubscribe anytime. Privacy Policy