AI Observability Companies in 2026
AI observability companies make software that lets teams trace, evaluate, monitor and debug machine learning models and LLM applications, including AI agents, once they are running in production. Their platforms typically record each prompt, response and tool call, score outputs with code, LLM-as-a-judge or human review, and alert on quality, cost, latency or safety problems.
AI Observability Companies to know
These companies were picked for notability as of October 2026: recent funding and press coverage, reported adoption (open-source downloads, enterprise customers) and how central LLM/AI observability and evaluation is to each company's product. Companies whose main identity is agent frameworks, model hosting or general MLOps were left out in favor of sibling directories. One exception is Datadog, a general observability vendor, which is included because it ships a dedicated LLM and agent observability product to a very large customer base. Companies that have been acquired and shut down or folded into another product were excluded.

Arize AI
AI observability and LLM evaluation platform with the open-source Phoenix tracer
Arize AI makes Arize AX, an enterprise platform for observability and evaluation of both traditional ML models and generative AI applications. It also maintains Arize Phoenix, an open-source tracing and evaluation tool that was reported at two million monthly downloads in February 2025. The company was founded in January 2020 and is based in Berkeley, California.
- Headquarters
- Berkeley, California, USA

Braintrust
Evaluation and observability platform for AI applications and agents
Braintrust (Braintrust Data Inc.) offers tracing, automated evaluation with scorers and LLM-as-a-judge, and a prompt playground for teams building AI applications and agents. It runs on Brainstore, its own database for querying AI traces. Reported customers include Notion, Replit, Cloudflare, Ramp and Dropbox.
- Headquarters
- San Francisco, California, USA

Langfuse
Open-source LLM observability, evals and prompt management, now part of ClickHouse
Langfuse is an open-source (MIT-licensed core) platform for LLM tracing, evaluations and prompt management that can be self-hosted or used as a cloud service. ClickHouse acquired it in January 2026 and said it would keep the project open source. At acquisition it reported more than 20,000 GitHub stars and 23.1 million monthly SDK installs.
- Headquarters
- Berlin, Germany

Datadog
Cloud monitoring company with LLM and agent observability built into its platform
Datadog is a publicly listed monitoring and security company. Its Agent Observability (LLM Observability) product traces prompts, retrieval steps and tool calls, runs built-in and custom evaluators, and flags hallucinations, prompt injection attempts and PII exposure. Agent traces are linked to the application and infrastructure data that Datadog already collects.
- Headquarters
- New York, New York, USA

Galileo
AI evaluation, observability and guardrail platform built on its own evaluator models
Galileo builds an evaluation, observability and guardrailing platform for generative AI applications and agents. It scores outputs with its own compact evaluation models (the Luna family, including Luna-2) and can turn offline evals into production guardrails. Its October 2024 Series B announcement named Comcast and Twilio among its enterprise customers.
- Headquarters
- Burlingame, California, USA

Fiddler AI
AI observability and security control plane for ML, GenAI and agents
Fiddler AI offers a platform for monitoring, evaluating and governing ML models, generative AI applications and AI agents, which it calls the Fiddler AI Control Plane. The January 2026 Series C brought its total funding to $100 million. The company was founded in 2018 and is based in Palo Alto, California.
- Headquarters
- Palo Alto, California, USA

Raindrop
Monitoring platform that detects and triages AI agent failures in production
Raindrop monitors AI agents running in production. It traces messages and tool calls, detects recurring failure patterns, tracks custom signals and experiments, and includes a triage agent that investigates issues. Its Simulations feature tests agent changes before deployment, and listed customers include Vercel, Speak, Clay, Framer and AngelList.
- Headquarters
- San Francisco, California, USA

Comet
ML experiment tracking company behind the open-source Opik LLM observability tool
Comet started with ML experiment management and model monitoring and now develops Opik, an open-source platform for LLM and agent tracing, evaluation and production monitoring. The company says it serves over 150,000 developers and 450 enterprise, startup and academic teams, and reports $70 million in total funding.
- Headquarters
- New York, New York, USA

HoneyHive
Observability and evaluation platform for AI agents in production
HoneyHive provides tracing, online and offline evaluation, monitoring, alerting, annotation queues and prompt management for AI agents, including custom agents, coding agents and agents built on no-code platforms. The company names Commonwealth Bank as a customer.
- Headquarters
- Brooklyn, New York, USA

Arthur
AI monitoring, governance and runtime security platform for AI agents
Arthur (Arthur AI) began in 2018 as a model monitoring company and now offers a platform that discovers AI agents across cloud, on-premise and endpoint environments. Its products cover agent behavioral analytics, evaluation, runtime security and governance policy enforcement. Named customers include Axios and Expel.
- Headquarters
- New York, New York, USA
AI Observability Companies at a glance
| Company | Headquarters | Focus |
|---|---|---|
| Arize AI | Berkeley, California, USA | llm-observability, evaluation |
| Braintrust | San Francisco, California, USA | evaluation, llm-observability |
| Langfuse | Berlin, Germany | open-source, llm-observability |
| Datadog | New York, New York, USA | llm-observability, apm |
| Galileo | Burlingame, California, USA | evaluation, guardrails |
| Fiddler AI | Palo Alto, California, USA | ml-monitoring, llm-observability |
| Raindrop | San Francisco, California, USA | agent-monitoring, llm-observability |
| Comet | New York, New York, USA | open-source, llm-observability |
| HoneyHive | Brooklyn, New York, USA | agent-observability, evaluation |
| Arthur | New York, New York, USA | ml-monitoring, governance |
Frequently asked questions
What do AI observability companies do?
They trace, evaluate and monitor AI models and LLM applications in production. Their platforms record prompts, responses and tool calls, score outputs with code, LLM-as-a-judge or human review, and track quality, latency, cost and safety. Examples include Arize AI, Braintrust, Langfuse and Galileo.
Which AI observability tools are open source?
Langfuse keeps its core under the MIT license and can be self-hosted, and ClickHouse said it would stay open source after acquiring it in January 2026. Arize maintains the open-source Arize Phoenix, and Comet maintains the open-source Opik.
Which AI observability companies raised money recently?
Raindrop announced a CRV-led Series A in September 2026, bringing its total funding to $50M. Braintrust raised an $80M Series B led by ICONIQ in February 2026 at a reported $800M valuation. Fiddler AI raised a $30M Series C led by RPS Ventures in January 2026, bringing its total to $100M. Arize AI raised a $70M Series C led by Adams Street Partners in February 2025.
How does LLM observability differ from traditional application monitoring?
Traditional APM tracks service latency, errors and infrastructure. LLM observability also captures model inputs, outputs and agent steps, and judges the output itself (accuracy, hallucinations, safety). Datadog now offers both: its Agent Observability product links agent traces to its existing application and infrastructure data.
How do AI observability platforms differ from AI agent infrastructure tools?
Agent infrastructure tools such as frameworks, memory, sandboxes and orchestration are used to build and run agents. Observability platforms such as Raindrop, HoneyHive and Braintrust watch, test and debug those agents once they are deployed. Some vendors span both, but the companies here are those whose main product is observability and evaluation.
More AI company directories
- AI Agent Infrastructure CompaniesLangChain, CrewAI, Browserbase and more
- AI Search CompaniesPerplexity, Glean, Exa and more
- Neocloud CompaniesCoreWeave, Nebius, Crusoe and more
- AI Inference CompaniesCerebras Systems, Fireworks AI, Groq and more
- Voice AI CompaniesElevenLabs, Deepgram, PolyAI and more
- AI Agent CompaniesSierra, Harvey, Decagon and more
About this directory: OpenCurious curates this list for researchers, founders, operators and buyers. Company descriptions and headquarters are based on publicly reported information as of and change often; figures a company reports about itself are its own claims. This is an editorial selection, not a paid ranking: no company paid to be listed, and companies appear in no particular order. Company names and logos are trademarks of their respective owners and are used only to identify each company. See our editorial policy. To suggest a company or a correction, email hello@opencurious.com.
The OpenCurious Newsletter
Stay curious. Get the best of the open web.
New directories, researcher tools, and open-source finds — straight to your inbox. No spam, ever.
- New directories
- Researcher tools
- Open-source finds
Occasional emails from OpenCurious. Unsubscribe anytime. Privacy Policy