OpenCurious Directory · Updated

AI Model Hosting Companies in 2026

AI model hosting companies run the platforms where developers upload, deploy and serve machine learning models, whether those are their own custom models or open-source ones. Typical products include model hubs, managed or serverless model endpoints, containerized deployment tools and autoscaling GPU runtimes. The category is about getting a model into production as an API, not about building the fastest inference chips.

  1. Hugging Face
  2. Baseten
  3. Modal
  4. fal
  5. Replicate
  6. DeepInfra
  7. Ollama
  8. Northflank
  9. Cerebrium
  10. Novita AI

AI Model Hosting Companies to know

We picked platforms whose main job is hosting, deploying and serving models (custom or open-source) as production endpoints. We selected using funding, reported revenue or usage, and press coverage as of October 2026. We left out companies that fit better on sibling pages: inference-speed specialists such as Groq, Cerebras, Together AI and Fireworks AI are on the AI Inference page, and GPU fleet operators such as CoreWeave and RunPod are on the Neocloud or GPU Cloud pages. Every company listed was still operating under its own brand when we did the research, including Hugging Face (acquisition by Nvidia pending) and Replicate (owned by Cloudflare, still running its own API).

  1. Hugging Face logo

    Hugging Face

    Open model and dataset hub with managed Inference Endpoints and Spaces hosting

    Hugging Face runs the Hugging Face Hub, where developers share and host machine learning models and datasets, and maintains the open-source Transformers library. It sells Inference Endpoints, a managed service that deploys Hub or custom models to autoscaling production endpoints, plus Spaces for hosting model demos. On September 3, 2026, Nvidia announced an agreement to acquire the company and said Hugging Face would remain an open platform.

    Headquarters
    Brooklyn, New York, USA
  2. Baseten logo

    Baseten

    Platform for deploying and serving custom and open-source models in production

    Baseten sells infrastructure for deploying machine learning models as production APIs. Its products include dedicated deployments with autoscaling, Model APIs for open-source models, the open-source Truss packaging framework, and the Chains SDK for multi-model workflows. It also offers multi-node training for fine-tuning.

    Headquarters
    San Francisco, California, USA
  3. fal logo

    fal

    Serverless hosting and APIs for generative image, video, audio and 3D models

    fal (Features & Labels Inc.) runs a serverless platform that hosts generative media models, such as Flux, Kling and Veo, and serves them by API. It covers image, video, audio and 3D generation, and lets developers deploy their own models without managing GPUs. It also sells developer tools such as Workflows and Sandbox.

    Headquarters
    San Francisco, California, USA
  4. Replicate logo

    Replicate

    Model catalog and API for running and deploying containerized open-source models

    Replicate hosts a catalog of more than 50,000 containerized AI models that developers can run through an API, and lets them deploy their own models using Cog, an open-source packaging tool. Cloudflare acquired Replicate in November 2025 and plans to fold its catalog and custom-model deployment into Workers AI. Replicate's site and API still operate under its own brand.

    Headquarters
    San Francisco, California, USA
  5. DeepInfra logo

    DeepInfra

    Inference cloud for hosting open-source models and custom deployments

    DeepInfra runs an inference cloud that serves open-source text, embedding, speech and image models through APIs, and also supports custom model deployments on dedicated GPUs. Its founders previously built backend infrastructure for the imo messenger app. Its May 2026 Series B included NVIDIA, Samsung Next and Supermicro among the investors.

    Headquarters
    Palo Alto, California, USA
  6. Ollama logo

    Ollama

    Open-source tool for running models locally, plus hosted cloud models

    Ollama is open-source software for downloading and running large language models on local machines through a command line, desktop app and local HTTP API, with Python and JavaScript client libraries. In 2025 and 2026 it added hosted cloud models, web search support and coding-agent integrations. Ollama Cloud is a paid service for running hosted open models.

    Headquarters
    Palo Alto, California, USA
  7. Northflank logo

    Northflank

    Developer platform for deploying services, GPU workloads and models on any cloud

    Northflank is a developer platform for deployments, preview environments, sandboxes and GPU workloads. It runs on Kubernetes, either in its own cloud or in customers' cloud accounts. Teams use it to deploy and serve their own models alongside their application services.

    Headquarters
    London, UK
  8. Cerebrium logo

    Cerebrium

    Serverless GPU platform for deploying AI models and voice agents

    Cerebrium runs a serverless GPU platform that deploys containerized AI workloads, including LLMs, video models and voice agents, with autoscaling across several clouds and regions. It accepts custom code without requiring an SDK rewrite. Customers it names include LiveKit, Vapi, Deepgram and Tavus. The company started in Cape Town and is now based in New York.

    Headquarters
    New York, New York, USA
  9. Novita AI logo

    Novita AI

    Serverless model APIs, agent sandboxes and GPU instances for developers

    Novita AI offers one API to more than 200 LLM, image, audio, video and vision models. It also sells agent sandboxes and GPU compute, from on-demand instances to bare-metal clusters. Organizations it lists as users include Hugging Face, Quora and OpenRouter.

    Headquarters
    San Francisco, California, USA

AI Model Hosting Companies at a glance

AI Model Hosting Companies: headquarters and focus
CompanyHeadquartersFocus
Hugging FaceBrooklyn, New York, USAmodel hub, inference endpoints
BasetenSan Francisco, California, USAmodel deployment, dedicated endpoints
ModalNew York, New York, USAserverless GPU, model serving
falSan Francisco, California, USAgenerative media, serverless endpoints
ReplicateSan Francisco, California, USAmodel catalog, model deployment
DeepInfraPalo Alto, California, USAopen-source models, model APIs
OllamaPalo Alto, California, USAlocal models, open source
NorthflankLondon, UKdeployment platform, GPU workloads
CerebriumNew York, New York, USAserverless GPU, model deployment
Novita AISan Francisco, California, USAmodel APIs, serverless GPU

Frequently asked questions

What is an AI model hosting company?

It runs a platform where developers deploy and serve machine learning models, their own or open-source ones, as production APIs. Examples include Hugging Face Inference Endpoints, Baseten's dedicated deployments, Replicate's model catalog and Cog packaging tool, and Modal's serverless GPU functions.

Which AI model hosting companies raised the most money recently?

Baseten raised a $1.5B Series F in June 2026, in tranches at $13B and $11B valuations. Modal raised a $355M Series C at $4.65B in May 2026, and a larger round led by Accel has been reported but not confirmed. fal raised $140M at $4.5B in December 2025, led by Sequoia. DeepInfra raised a $107M Series B in May 2026.

How do model hosting platforms differ from AI inference companies?

Hosting platforms focus on deploying and running any model, including your own custom one, with autoscaling endpoints, packaging tools such as Truss and Cog, and model catalogs. Inference specialists focus mainly on serving popular models as fast and cheaply as possible, sometimes on custom chips. The two categories overlap, and companies such as Baseten and DeepInfra describe their own work in inference terms.

Is Hugging Face still independent?

Not for much longer. On September 3, 2026, Nvidia announced an agreement to acquire Hugging Face for $12.93B, expected to close in the first half of 2027 subject to regulatory approval. Nvidia said Hugging Face would remain an open platform and that Nvidia hardware would not be required to deploy on it.

Is Replicate still available after the Cloudflare acquisition?

Yes. Cloudflare acquired Replicate in November 2025 and said Replicate's existing APIs and workflows would keep working. It plans to bring Replicate's catalog of more than 50,000 models into Cloudflare Workers AI.

Can I host open-source models locally instead of in the cloud?

Yes. Ollama is open-source software for running models on your own machine through a command line and a local HTTP API. It also sells Ollama Cloud, a hosted service for open models.

All AI company directories →

About this directory: OpenCurious curates this list for researchers, founders, operators and buyers. Company descriptions and headquarters are based on publicly reported information as of and change often; figures a company reports about itself are its own claims. This is an editorial selection, not a paid ranking: no company paid to be listed, and companies appear in no particular order. Company names and logos are trademarks of their respective owners and are used only to identify each company. See our editorial policy. To suggest a company or a correction, email hello@opencurious.com.

The OpenCurious Newsletter

Stay curious. Get the best of the open web.

New directories, researcher tools, and open-source finds — straight to your inbox. No spam, ever.

  • New directories
  • Researcher tools
  • Open-source finds

Occasional emails from OpenCurious. Unsubscribe anytime. Privacy Policy