OpenCurious Directory · Updated

AI Inference Companies in 2026

AI inference companies build the products that run trained AI models in production and return outputs (tokens, images, embeddings) quickly and at low cost. The category covers inference APIs and platforms that serve open and custom models at low latency, inference engines, and chips or systems designed for inference rather than training.

  1. Cerebras Systems
  2. Fireworks AI
  3. Groq
  4. Together AI
  5. Baseten
  6. Etched
  7. SambaNova Systems
  8. d-Matrix
  9. DeepInfra
  10. Inferact

AI Inference Companies to know

These companies were chosen because inference is their core product, either as an API or platform for serving models, an inference engine, or silicon and systems built for inference. They are selected using notability as of October 2026, using public-market or private valuation, the size of the latest funding round, reported revenue or adoption, and independent press coverage. Companies known mainly for general GPU rental (neoclouds), model hubs or broad MLOps were left to sibling directories. Together AI is the one exception: it stays here because it is a top player in both inference and neocloud.

  1. Cerebras Systems logo

    Cerebras Systems

    Wafer-scale AI chips and a cloud inference service built on them

    Cerebras designs the Wafer-Scale Engine (WSE-3), an AI processor built on a single wafer. It sells CS-3 systems and offers inference and training cloud APIs, positioning wafer-scale hardware as a lower-latency alternative to GPU clusters for serving models. Customers named in IPO coverage include OpenAI, G42, MBZUAI and AWS.

    Headquarters
    Sunnyvale, California, USA
  2. Fireworks AI logo

    Fireworks AI

    Inference and model-serving platform for open and customized models

    Fireworks AI provides a platform to deploy, customize and serve generative AI models, with 200+ models across text, image and multimodal. It was founded by former Meta engineers, including CEO Lin Qiao, who previously led Meta's PyTorch team. In July 2026 it reported annualized revenue above $1B.

    Headquarters
    San Mateo, California, USA
  3. Groq logo

    Groq

    LPU-based inference cloud running from its own data centers

    Groq developed the Language Processing Unit (LPU), an ASIC designed for inference, and runs GroqCloud, an inference service used by developers. In December 2025 Nvidia signed a non-exclusive license to Groq's technology, in a deal valued at about $20B, and hired founder Jonathan Ross and president Sunny Madra. Groq stayed independent and runs its inference cloud from 13 data center locations, per TechCrunch.

    Headquarters
    Mountain View, California, USA
  4. Together AI logo

    Together AI

    Open-model inference, fine-tuning and GPU clusters from one cloud

    Together AI hosts open-source models behind inference APIs and also offers fine-tuning and Nvidia GPU clusters, which it positions as a lower-cost alternative to closed models. Its founders include CEO Vipul Ved Prakash and researchers Ce Zhang, Chris Re, Tri Dao and Percy Liang. TechCrunch reported annual bookings above $1.15B in mid-2026, with customers including Cursor, Cognition and Decagon.

    Headquarters
    San Francisco, California, USA
  5. Baseten logo

    Baseten

    Inference platform for deploying and scaling production AI models

    Baseten provides an inference platform and systems software for running open, custom and fine-tuned models in production, covering GPUs, autoscaling, observability and billing. It manages capacity across 18 clouds and 87 clusters and says it handles more than 1 billion inference calls a day. Customers include Abridge, Cursor, Lovable and OpenEvidence.

    Headquarters
    San Francisco, California, USA
  6. Etched logo

    Etched

    Custom inference chips and full-stack inference clusters

    Etched designs inference hardware: co-designed chips, racks and software sold as frontier inference clusters for running large language models. TechCrunch reported $1B in customer orders as of July 2026, including from Jane Street, along with a 10-megawatt data center in Silicon Valley and about 400 employees. In October 2026 TechCrunch reported that it was weighing funding offers at $40B to $50B.

    Headquarters
    San Jose, California, USA
  7. SambaNova Systems logo

    SambaNova Systems

    Dataflow AI chips and systems for enterprise inference

    SambaNova builds Reconfigurable Dataflow Unit (RDU) inference chips, the SN40L and the SN50 (unveiled February 2026, shipping in H2 2026), along with inference systems and cloud services. TechCrunch reported that JPMorgan Chase chose SambaNova as an inference-infrastructure partner and that SoftBank is the first deployment partner for the SN50. Intel, which was earlier reported to be in acquisition talks with SambaNova, invested in the Series F and is co-developing inference products with it.

    Headquarters
    San Jose, California, USA
  8. d-Matrix logo

    d-Matrix

    Memory-centric compute platform for generative AI inference

    d-Matrix builds Corsair, an inference accelerator that does compute in memory to reduce latency and energy use. It also offers JetStream networking accelerators and the Aviator software stack. In 2026 it acquired Wallaroo.ai (August) and adopted Nvidia NVLink Fusion rack-scale infrastructure (September).

    Headquarters
    Santa Clara, California, USA
  9. DeepInfra logo

    DeepInfra

    Inference cloud serving open models via OpenAI-compatible APIs

    DeepInfra runs a GPU inference cloud from eight U.S. data centers. It serves 150+ open-source models through OpenAI-compatible APIs and also offers GPU instances for inference workloads. Its founding team previously built infrastructure for the imo messenger.

    Headquarters
    Palo Alto, California, USA
  10. Inferact logo

    Inferact

    Startup founded by the vLLM creators to commercialize the inference engine

    Inferact was founded by the creators of vLLM, an open-source LLM inference engine incubated in 2023 in Ion Stoica's lab at UC Berkeley, to commercialize and develop the project. The company says vLLM supports 500+ model architectures and 200+ accelerator types. Simon Mo is CEO.

    Headquarters
    San Francisco, California, USA

AI Inference Companies at a glance

AI Inference Companies: headquarters and focus
CompanyHeadquartersFocus
Cerebras SystemsSunnyvale, California, USAinference chips, inference API
Fireworks AISan Mateo, California, USAinference API, open models
GroqMountain View, California, USAinference API, inference chips
Together AISan Francisco, California, USAinference API, open models
BasetenSan Francisco, California, USAinference platform, model deployment
EtchedSan Jose, California, USAinference chips, ASIC
SambaNova SystemsSan Jose, California, USAinference chips, dataflow
d-MatrixSanta Clara, California, USAinference chips, in-memory compute
DeepInfraPalo Alto, California, USAinference API, open models
InferactSan Francisco, California, USAinference engine, open source

Frequently asked questions

What is an AI inference company?

An AI inference company's core product runs trained AI models in production quickly and at low cost. It may be a hosted inference API (Fireworks AI, Together AI, Baseten, DeepInfra, GroqCloud), an inference engine (Inferact's vLLM), or chips and systems built for inference (Cerebras, Etched, SambaNova, d-Matrix).

Which AI inference companies are the largest as of 2026?

Cerebras went public on Nasdaq (CBRS) in May 2026, raising about $5.5B at $185 a share. The largest private rounds in 2026 included Etched ($700M at $21B, August), Fireworks AI ($1.505B at $17.5B, July), Baseten ($1.5B in tranches at $13B and $11B, June), SambaNova ($1B first close at $11B, July) and Together AI ($800M at $8.3B, July).

How do inference chip companies differ from inference API providers?

Chip companies design their own silicon and systems and often sell cloud access too. Examples are Cerebras (wafer-scale WSE-3), Etched (custom inference chips and clusters), SambaNova (RDU chips), d-Matrix (Corsair in-memory compute) and Groq (LPU). API providers such as Fireworks AI, Together AI, Baseten and DeepInfra mostly run models on GPUs and compete on software, model optimization and price.

What happened between Groq and Nvidia?

In December 2025 Nvidia signed a non-exclusive license to Groq's inference technology, in a deal valued at about $20B, and hired founder Jonathan Ross and president Sunny Madra. Groq stayed independent under CEO Doug Wightman, confirmed a $650M raise in June 2026 led by Disruptive and Infinitum, and runs GroqCloud from 13 data center locations.

How is AI inference different from model hosting or GPU clouds?

Inference specialists focus on the speed and cost of serving model outputs, using faster runtimes, custom chips or tuned APIs. Model-hosting platforms focus on deploying whatever model you bring, and GPU clouds and neoclouds rent raw compute. Some companies span more than one of these; Together AI, for example, offers inference APIs and GPU clusters, and TechCrunch describes it as a neocloud.

Which open-source project underpins many inference platforms?

vLLM, an open-source LLM inference engine incubated at UC Berkeley in 2023. Its creators founded Inferact, which raised a $150M seed round at an $800M valuation in January 2026, co-led by Andreessen Horowitz and Lightspeed, to commercialize it.

All AI company directories →

About this directory: OpenCurious curates this list for researchers, founders, operators and buyers. Company descriptions and headquarters are based on publicly reported information as of and change often; figures a company reports about itself are its own claims. This is an editorial selection, not a paid ranking: no company paid to be listed, and companies appear in no particular order. Company names and logos are trademarks of their respective owners and are used only to identify each company. See our editorial policy. To suggest a company or a correction, email hello@opencurious.com.

The OpenCurious Newsletter

Stay curious. Get the best of the open web.

New directories, researcher tools, and open-source finds — straight to your inbox. No spam, ever.

  • New directories
  • Researcher tools
  • Open-source finds

Occasional emails from OpenCurious. Unsubscribe anytime. Privacy Policy