Amazon CloudWatch Omni Innovates Observability for AI Systems
For decades, the observability sector has centered on a fundamental question: Is the system running? However, the emergence of agentic artificial intelligence challenges this paradigm. An AI agent may fulfill performance targets and operate error-free but could still provide inaccurate responses, utilize the wrong tools, or access outdated information. While traditional metrics may indicate a healthy system, the ultimate measure of success for a business is whether it delivered the right outcomes. This gap is precisely what Amazon Web Services Inc. (AWS) seeks to address with its new offering, Amazon CloudWatch Omni, which became generally available last week.
AWS frames Omni as the next evolution of its CloudWatch service. This app-centric solution operates independently of the AWS Management Console, leverages OpenTelemetry, and incorporates AI capabilities. By shifting focus from “Is it running?” to “Why did my agent respond this way?”, AWS aims to tackle the complexities most enterprises face when deploying agentic AI at scale. According to an IDC projection, over 1 billion agents will be deployed by 2029, making manual oversight of such systems impractical for operational teams dealing with inherent non-deterministic behaviors.
IT and business leaders should consider several key aspects regarding the deployment of Amazon CloudWatch Omni:
Evaluation Takes Precedence Over Monitoring
The core innovation of Omni lies not in its dashboards, but in its robust evaluation engine. Omni meticulously captures every interaction and includes 17 built-in evaluators that assess coherence, helpfulness, and routing correctness, among other critical metrics. This functionality allows teams to compare different prompt iterations, generate test datasets from live traffic, and automatically identify regressions. Additionally, evaluators can operate continually against real-time traffic, flagging quality drifts similarly to how CPU spikes are monitored. While traditional latency and error rates offer limited insight, scoring metrics provide a clearer picture of answer accuracy.
Sony has emerged as an early adopter of this technology. Masahiro Oba, Senior General Manager of the AI Acceleration Division at Sony, stated that their enterprise-wide agentic AI platform now supports numerous proof-of-concept and production workloads. With Amazon CloudWatch Omni, he can track a single trace straight to evaluation, AI analysis, and dataset creation. This centralized approach alleviates challenges that arise from having numerous teams develop agents independently, each with distinct criteria for what a successful outcome looks like.
Independent Operation Extends Beyond the Console
Omni provides developers with native extensions for popular development tools such as Visual Studio Code, Cursor, and Kiro, allowing for local agent tracing without the necessity of an AWS account. Meanwhile, operators benefit from a standalone web interface equipped with single sign-on functionality via existing identity providers like Okta and Microsoft Entra ID. Both user groups share a unified data layer, ensuring that the trace a developer investigates matches the one an operator reviews.
AWS recognizes that its Management Console was originally designed for infrastructure administrators rather than site reliability engineers, AI engineers, and application owners, who now bear operational responsibility. By integrating into developers’ environments, where advanced AI coding tools like Claude Code and Codex can facilitate instrumentation, quality control is emphasized at the stage where issues are easiest to resolve.
A Unified Data Layer Enhances Investigative Processes
While numerous startups can track large language model calls, the distinguishing feature of Omni lies in its ability to consolidate agent traces, application telemetry, and infrastructure signals within a single CloudWatch data repository. As a result, a troubleshooting process can begin with an agent receiving an erroneous tool result, transition to an application program interface error due to capacity limitations, and conclude with a depleted database connection pool.
In contrast, many organizations currently require multiple tools and teams to piece together these insights, leading to inefficiencies. With the AWS DevOps Agent enabled by default during investigation sessions, the platform correlates signals and maintains a comprehensive investigation history. Capital One, a strategic design partner, highlighted the significance of developing a cohesive, AI-powered observability solution capable of offering topology-aware intelligence and natural-language querying across all telemetry.
Openness and Standardization Mitigate Vendor Lock-In
Omni operates based on OpenInference and the AWS Distro for OpenTelemetry, allowing agents to function on AWS as well as other cloud platforms. The solution supports various technologies, including LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, and third-party evaluators such as DeepEval. Furthermore, integration with Azure is currently supported, with plans for increased multi-cloud functionality on the horizon.
This level of openness is crucial in an increasingly competitive landscape, where companies like Datadog, Dynatrace, New Relic, and Splunk are expanding their own agent observability capabilities. Although OpenTelemetry facilitates data portability, the intelligence layer—including topology, evaluators, and investigation history—remains unique to AWS. Organizations significantly invested in AWS will likely view Omni as the favored option, while those utilizing established platforms like Datadog or Splunk across multiple clouds may prefer to use Omni for agent development and evaluation purposes.
Cost Structure Encourages Adoption, Monitoring is Key
The IDE extension for Amazon CloudWatch Omni is free, with costs incurred based on telemetry data sent and stored. While dashboards and alerts are provided at no additional charge, customers are allowed to query up to five times their monthly ingestion volume. Eligible accounts can access a 30-day trial period and $1,000 in OpenTelemetry ingestion credits.
However, it’s important to monitor costs, as agents generate numerous spans from prompts, model calls, tool interactions, and sub-agent transitions. With hundreds of workloads actively engaged in continuous evaluation, telemetry ingestion costs could easily surpass the budgets allocated for AI initiatives. Additionally, the pricing structure for the DevOps Agent is separate.
Implications for Decision-Makers
Amazon CloudWatch Omni presents a robust platform for agent observability, aiming to bridge the trust gap that hinders widespread implementation of agentic AI. IT leaders are advised to:
- Establish Clear Definitions: Ensure that your business leaders have documented criteria for what constitutes a correct, compliant, and helpful response for each AI agent.
- Standardize Instrumentation Across Agents: Implement OpenTelemetry for every pilot project, enabling flexibility for future platform decisions.
- Model Telemetry Costs Early: Develop policies governing sampling, retention, and evaluation frequency prior to deploying agents.
- Select an Appropriate System of Record: If AWS serves as your primary cloud provider, Omni is a favorable choice; however, those with multi-cloud setups should evaluate its use for agent evaluation without displacing existing observability systems.
- Incorporate Investigation History into Governance: Ensure the captured investigation records contribute to AI risk management and compliance processes, particularly for regulated sectors.
As the industry evolves to understand and observe distributed AI systems, the focus must shift from merely monitoring system functionality to scrutinizing decision-making processes. AWS is positioning itself as the leader in providing this dimension of observability, encouraging companies to treat evaluation as a critical operational practice rather than merely an added feature.
Zeus Kerravala is a principal analyst at ZK Research, a division of Kerravala Consulting. This article was originally published by SiliconANGLE.

