Dom Nguyen
Senior Product Marketing Manager
Managing cloud, AI, and SaaS costs means answering a steady stream of questions from finance, leadership, and engineering teams. What changed? Which team owns the spend? Was an increase expected? Are we still on track against the budget? When each answer requires moving between dashboards, filtering cost data by team or service, or manually correlating billing data with observability data, it can slow down investigations while costs continue to rise.
The Cloud Cost skill in Bits Chat brings Cloud Cost Management analysis into a conversational workflow. You can ask cost questions in plain language, investigate cost anomalies, and get answers grounded in both cost and observability data within minutes. In this post, we’ll show how you can:
Ask Bits Chat anything about your costs
The Cloud Cost skill gives FinOps practitioners and engineers one place to ask cost questions without needing to know which dashboard to open or how to construct a query. You can get started with Bits Chat in the Datadog navigation bar, or you can start an investigation directly from a cost anomaly. From there, Bits Chat analyzes the relevant cost data and responds with a concise summary, then asks what you want to explore next.
Because the Cloud Cost skill works across many cost categories, including cloud, SaaS, AI, custom, and Datadog, teams can use the same workflow for any cost question. FinOps teams can use it to track budgets, identify cost owners, and investigate anomalies. Engineers can use it to understand the cost impact of their own services and workloads, and identify optimization opportunities. For example, you can ask Bits to investigate why Anthropic costs increased last week for team:web-store
, identify which teams are driving the highest OpenAI spend this month, or show total AI cost by provider for the last 30 days.
Bits Chat can investigate cost monitor alerts, cost anomalies, and cost changes; identify teams, services, accounts, regions, or resources that are driving spend; and compare actual or forecasted spend against budgets. It can also correlate cost changes with observability metrics such as CPU, memory, request volume, and storage size. This gives teams the technical context they need to understand whether a cost change came from a usage increase, a configuration change, or another source.
Spot Anthropic cost spikes before they grow
AI usage can become expensive quickly, especially when costs are spread across providers, models, teams, users, and services. AI Costs in Cloud Cost Management gives FinOps and engineering teams a unified location for analyzing AI spend across providers such as OpenAI, Anthropic, Amazon Bedrock, Google Gemini, Vertex AI, GitHub Copilot, and Cursor. Datadog normalizes this spend so teams can understand which providers, models, users, and API keys are contributing to cost changes.
This level of granularity helps teams detect AI cost anomalies before they cause larger budget issues. For example, Datadog can surface an unexpected Anthropic cost spike that might otherwise go unnoticed until the next billing review. But identifying a spike is only half the problem—understanding what caused it is another. Without a connected cost investigation workflow, a team would need to pull billing exports, cross-reference logs, and search for the service or team that caused the change.
With the Cloud Cost skill, the team can start the investigation in one click. When the team clicks “Investigate” directly from the anomaly, Bits automatically pulls in the relevant context and kicks off a root cause analysis. This lets the team move from detection to investigation without navigating to a new page or having to do any manual work.
Find the root cause of a cost change
Once an investigation starts, Bits Chat performs root cause analysis by using both cost data and observability data from across Datadog. For a cost anomaly investigation, the initial analysis typically includes a daily cost chart for the baseline and investigation periods so you can see exactly when a change started. It also summarizes the total dollar change, percentage change, and projected annual impact when applicable.
Bits Chat also provides rate-versus-usage context, which helps teams determine whether spend increased because of a pricing change or because a team consumed more of a resource. For AI spend, this distinction is especially important. An Anthropic spike might come from more requests, larger prompts, higher output token usage, a model change, or a combination of these factors. By correlating cost with observability metrics, Bits can help connect the billing change to the system behavior behind it.
In the preceding example screenshot, Bits identifies that the Anthropic cost spike is driven by two teams: Customer Support (57%) and AI Platform (37%). Together, these teams account for the full $2,895.49 increase—a 188.15% jump—with a projected annual impact of $62,167.87. Instead of knowing only that Anthropic costs went up, the engineer can see exactly which teams are driving it and by how much, which makes it immediately clear who to loop in.
From there, the team can keep drilling into the investigation. Bits can find the top services, accounts, regions, resources, or tags driving the change. It can also compare actual and forecasted spend against related budgets, identify optimization opportunities, and create a Datadog Notebook that captures the investigation for the owning team. This helps preserve the full context, so the team does not need to reconstruct the analysis later from Slack threads or separate reports.
Get started with the Cloud Cost skill in Bits Chat
The Cloud Cost skill in Bits Chat helps FinOps and engineering teams investigate cost changes, budget risks, and AI spend in the same place where they already monitor their systems. By grounding answers in both cost and observability data, Bits Chat can help teams move from a broad cost question to an actionable explanation in minutes.
To get started, read the Cloud Cost skill documentation and make sure you’ve configured Cloud Cost Management for the cost sources you want to analyze. Users also need Bits Chat access and Cloud Cost Management permissions for the data they ask about. If you want Bits to create or update investigation records, you can also grant the relevant Notebooks permissions. Learn more in the Bits Chat documentation and AI Costs in Cloud Cost Management documentation.
If you don’t already have a Datadog account, you can sign up for a 14-day free trial to start investigating your cloud costs with Datadog.
Facts Only
* Cloud Cost skill brings Cloud Cost Management analysis into a conversational workflow.
* Users can ask cost questions in plain language to investigate costs.
* The skill analyzes relevant cost data and responds with a concise summary.
* The skill investigates cost monitor alerts, anomalies, and changes.
* It identifies teams, services, accounts, regions, or resources driving spend.
* The skill correlates cost changes with observability metrics such as CPU, memory, request volume, and storage size.
* AI Costs in Cloud Cost Management unifies spending across providers like OpenAI, Anthropic, Amazon Bedrock, Google Gemini, Vertex AI, GitHub Copilot, and Cursor.
* Investigations can identify drivers such as teams, models, users, and API keys contributing to cost changes.
* Root cause analysis uses both cost data and observability data from Datadog.
* An example investigation identified that an Anthropic cost spike was driven by the Customer Support (57%) and AI Platform (37%) teams, accounting for a $2,895.49 increase.
* The skill allows users to compare actual or forecasted spend against budgets and identify optimization opportunities.
Executive Summary
The Cloud Cost skill in Bits Chat integrates cost and observability data to provide a conversational workflow for managing cloud, SaaS, AI, custom, and Datadog spending. It allows users to ask plain-language questions about costs, investigate anomalies, and correlate billing data with metrics like CPU usage or request volume. This capability enables FinOps practitioners and engineers to move beyond manual dashboard navigation by investigating cost drivers, identifying cost owners, and correlating spend changes with system behavior.
The skill facilitates the detection of AI cost anomalies, such as spikes in Anthropic spending, by normalizing costs across various providers like OpenAI, Anthropic, and Google Gemini. When an anomaly is detected, the skill initiates a root cause analysis by using both cost data and observability metrics to determine if a change resulted from increased usage or configuration shifts. This process can identify specific teams, services, accounts, or resources responsible for cost changes, enabling teams to pinpoint root causes instantly.
Ultimately, the feature aims to streamline the complex process of cloud cost management by unifying billing information with operational data, allowing for faster detection and investigation of cost deviations related to consumption and configuration.
Full Take
The value proposition lies in shifting the investigative paradigm from manual correlation of disparate data sources—billing exports and logs—to an integrated, context-aware causal analysis within a single workflow. This integration addresses a critical bottleneck in FinOps and engineering: the latency between cost detection and actionable root cause identification. The ability to correlate financial markers (cost) with operational reality (observability metrics like CPU or request volume) moves analysis beyond mere reporting into genuine systems understanding, which is essential for effective optimization.
The pattern observed is the systemic friction introduced by data silos. When teams must manually stitch together cost changes with log data across multiple platforms, complexity breeds delay and potential error, directly impeding agility in dynamic cloud environments, especially concerning volatile areas like AI spending. The system attempts to solve this by embedding the analytical correlation engine directly into the query layer.
The deeper implication is how operational context affects financial accountability. By linking spend directly to performance metrics, the framework enforces a model where cost variance is understood not as an abstract number but as a direct consequence of system behavior or configuration decisions made by specific entities (teams/services). The risk lies in the assumption that all relevant observability data is correctly and uniformly mapped to billing events across all services; if the correlation layer introduces even subtle bias, the "root cause" identified may be misleading, inadvertently shifting accountability without fully resolving the underlying architectural or policy issues.
Bridge questions: How robust is the normalization mechanism when dealing with multi-cloud provider pricing structures that change frequently? What are the documented guardrails against attributing cost variance solely to usage metrics versus infrastructure configuration shifts? Does this integrated workflow inherently shift organizational responsibility, or does it simply provide a better tool for existing allocation structures?
Sentinel — Human
The text reads like expert-driven product marketing or technical documentation focused on solving complex data correlation issues in cloud cost management, displaying high human-authored domain knowledge.
