AI Model Hosting Companies in 2026
AI model hosting companies run the platforms where developers upload, deploy and serve machine learning models, whether those are their own custom models or open-source ones. Typical products include model hubs, managed or serverless model endpoints, containerized deployment tools and autoscaling GPU runtimes. The category is about getting a model into production as an API, not about building the fastest inference chips.
AI Model Hosting Companies to know
We picked platforms whose main job is hosting, deploying and serving models (custom or open-source) as production endpoints. We selected using funding, reported revenue or usage, and press coverage as of October 2026. We left out companies that fit better on sibling pages: inference-speed specialists such as Groq, Cerebras, Together AI and Fireworks AI are on the AI Inference page, and GPU fleet operators such as CoreWeave and RunPod are on the Neocloud or GPU Cloud pages. Every company listed was still operating under its own brand when we did the research, including Hugging Face (acquisition by Nvidia pending) and Replicate (owned by Cloudflare, still running its own API).

Hugging Face
Open model and dataset hub with managed Inference Endpoints and Spaces hosting
Hugging Face runs the Hugging Face Hub, where developers share and host machine learning models and datasets, and maintains the open-source Transformers library. It sells Inference Endpoints, a managed service that deploys Hub or custom models to autoscaling production endpoints, plus Spaces for hosting model demos. On September 3, 2026, Nvidia announced an agreement to acquire the company and said Hugging Face would remain an open platform.
- Headquarters
- Brooklyn, New York, USA

Baseten
Platform for deploying and serving custom and open-source models in production
Baseten sells infrastructure for deploying machine learning models as production APIs. Its products include dedicated deployments with autoscaling, Model APIs for open-source models, the open-source Truss packaging framework, and the Chains SDK for multi-model workflows. It also offers multi-node training for fine-tuning.
- Headquarters
- San Francisco, California, USA

Modal
Serverless Python cloud for running model inference, training and sandboxes on GPUs
Modal lets developers run Python functions as autoscaling cloud workloads on CPUs and GPUs, and the same platform hosts model inference endpoints. Its products cover inference for LLM, image, video and audio models, sandboxes for running untrusted code, and GPU training. The company said it passed $300M in annualized revenue when it announced its Series C in May 2026.
- Headquarters
- New York, New York, USA

fal
Serverless hosting and APIs for generative image, video, audio and 3D models
fal (Features & Labels Inc.) runs a serverless platform that hosts generative media models, such as Flux, Kling and Veo, and serves them by API. It covers image, video, audio and 3D generation, and lets developers deploy their own models without managing GPUs. It also sells developer tools such as Workflows and Sandbox.
- Headquarters
- San Francisco, California, USA

Replicate
Model catalog and API for running and deploying containerized open-source models
Replicate hosts a catalog of more than 50,000 containerized AI models that developers can run through an API, and lets them deploy their own models using Cog, an open-source packaging tool. Cloudflare acquired Replicate in November 2025 and plans to fold its catalog and custom-model deployment into Workers AI. Replicate's site and API still operate under its own brand.
- Headquarters
- San Francisco, California, USA

DeepInfra
Inference cloud for hosting open-source models and custom deployments
DeepInfra runs an inference cloud that serves open-source text, embedding, speech and image models through APIs, and also supports custom model deployments on dedicated GPUs. Its founders previously built backend infrastructure for the imo messenger app. Its May 2026 Series B included NVIDIA, Samsung Next and Supermicro among the investors.
- Headquarters
- Palo Alto, California, USA

Ollama
Open-source tool for running models locally, plus hosted cloud models
Ollama is open-source software for downloading and running large language models on local machines through a command line, desktop app and local HTTP API, with Python and JavaScript client libraries. In 2025 and 2026 it added hosted cloud models, web search support and coding-agent integrations. Ollama Cloud is a paid service for running hosted open models.
- Headquarters
- Palo Alto, California, USA

Northflank
Developer platform for deploying services, GPU workloads and models on any cloud
Northflank is a developer platform for deployments, preview environments, sandboxes and GPU workloads. It runs on Kubernetes, either in its own cloud or in customers' cloud accounts. Teams use it to deploy and serve their own models alongside their application services.
- Headquarters
- London, UK

Cerebrium
Serverless GPU platform for deploying AI models and voice agents
Cerebrium runs a serverless GPU platform that deploys containerized AI workloads, including LLMs, video models and voice agents, with autoscaling across several clouds and regions. It accepts custom code without requiring an SDK rewrite. Customers it names include LiveKit, Vapi, Deepgram and Tavus. The company started in Cape Town and is now based in New York.
- Headquarters
- New York, New York, USA

Novita AI
Serverless model APIs, agent sandboxes and GPU instances for developers
Novita AI offers one API to more than 200 LLM, image, audio, video and vision models. It also sells agent sandboxes and GPU compute, from on-demand instances to bare-metal clusters. Organizations it lists as users include Hugging Face, Quora and OpenRouter.
- Headquarters
- San Francisco, California, USA
AI Model Hosting Companies at a glance
| Company | Headquarters | Focus |
|---|---|---|
| Hugging Face | Brooklyn, New York, USA | model hub, inference endpoints |
| Baseten | San Francisco, California, USA | model deployment, dedicated endpoints |
| Modal | New York, New York, USA | serverless GPU, model serving |
| fal | San Francisco, California, USA | generative media, serverless endpoints |
| Replicate | San Francisco, California, USA | model catalog, model deployment |
| DeepInfra | Palo Alto, California, USA | open-source models, model APIs |
| Ollama | Palo Alto, California, USA | local models, open source |
| Northflank | London, UK | deployment platform, GPU workloads |
| Cerebrium | New York, New York, USA | serverless GPU, model deployment |
| Novita AI | San Francisco, California, USA | model APIs, serverless GPU |
Frequently asked questions
What is an AI model hosting company?
It runs a platform where developers deploy and serve machine learning models, their own or open-source ones, as production APIs. Examples include Hugging Face Inference Endpoints, Baseten's dedicated deployments, Replicate's model catalog and Cog packaging tool, and Modal's serverless GPU functions.
Which AI model hosting companies raised the most money recently?
Baseten raised a $1.5B Series F in June 2026, in tranches at $13B and $11B valuations. Modal raised a $355M Series C at $4.65B in May 2026, and a larger round led by Accel has been reported but not confirmed. fal raised $140M at $4.5B in December 2025, led by Sequoia. DeepInfra raised a $107M Series B in May 2026.
How do model hosting platforms differ from AI inference companies?
Hosting platforms focus on deploying and running any model, including your own custom one, with autoscaling endpoints, packaging tools such as Truss and Cog, and model catalogs. Inference specialists focus mainly on serving popular models as fast and cheaply as possible, sometimes on custom chips. The two categories overlap, and companies such as Baseten and DeepInfra describe their own work in inference terms.
Is Hugging Face still independent?
Not for much longer. On September 3, 2026, Nvidia announced an agreement to acquire Hugging Face for $12.93B, expected to close in the first half of 2027 subject to regulatory approval. Nvidia said Hugging Face would remain an open platform and that Nvidia hardware would not be required to deploy on it.
Is Replicate still available after the Cloudflare acquisition?
Yes. Cloudflare acquired Replicate in November 2025 and said Replicate's existing APIs and workflows would keep working. It plans to bring Replicate's catalog of more than 50,000 models into Cloudflare Workers AI.
Can I host open-source models locally instead of in the cloud?
Yes. Ollama is open-source software for running models on your own machine through a command line and a local HTTP API. It also sells Ollama Cloud, a hosted service for open models.
More AI company directories
- Healthcare Intelligence CompaniesIQVIA, Definitive Healthcare, Komodo Health and more
- AI Observability CompaniesArize AI, Braintrust, Langfuse and more
- AI Agent Infrastructure CompaniesLangChain, CrewAI, Browserbase and more
- AI Search CompaniesPerplexity, Glean, Exa and more
- Neocloud CompaniesCoreWeave, Nebius, Crusoe and more
- AI Inference CompaniesCerebras Systems, Fireworks AI, Groq and more
About this directory: OpenCurious curates this list for researchers, founders, operators and buyers. Company descriptions and headquarters are based on publicly reported information as of and change often; figures a company reports about itself are its own claims. This is an editorial selection, not a paid ranking: no company paid to be listed, and companies appear in no particular order. Company names and logos are trademarks of their respective owners and are used only to identify each company. See our editorial policy. To suggest a company or a correction, email hello@opencurious.com.
The OpenCurious Newsletter
Stay curious. Get the best of the open web.
New directories, researcher tools, and open-source finds — straight to your inbox. No spam, ever.
- New directories
- Researcher tools
- Open-source finds
Occasional emails from OpenCurious. Unsubscribe anytime. Privacy Policy