Field GuideLast reviewed October 2026
Observability: metrics, logs and traces
How companies see what their software does in production: collecting metrics, logs and traces, why what you index sets the bill, what OpenTelemetry changed, why noisy alerts let incidents through, and what's real about security and AI using the same data.
The industry on one page
Picture a company that sells scheduling and dispatch software to plumbing and electrical contractors. It runs about 500 hosts on one cloud and writes about 2 TB of logs a day. When a dispatcher drags a job onto a technician's calendar and nothing happens for eight seconds, somebody at the dispatch company needs to know within minutes, find out why and fix it. Observability is everything that makes that possible, and it's one of the larger software bills the company pays.
Your services emit telemetry: metrics (numbers over time, like error rate), logs (lines of text describing what happened), traces (the path of one request across services, made of one span per step) and, increasingly, profiles of which code burns CPU. Collectors pick it up: a vendor's agent on each host, or OpenTelemetry, the open standard for instrumenting code and moving telemetry. They send it to a pipeline that filters, samples (keeps a fraction on purpose) and routes it. The pipeline decides what's kept before the platform stores it, and it's where the bill is set, because everything after it is priced by volume. The platform stores, queries and alerts the on-call team, the engineers whose phones ring, who use dashboards and traces to fix the service. The same data now feeds security and AI products: a SIEM (the security team's log analysis and detection system), cloud security tools, and AI agents that investigate incidents. That reuse is pulling the observability, security and AI markets closer.
What I'd want a new PM in this space to take away:
- The bill: what you index sets the bill. Vendors price on volume units (hosts, gigabytes, events, metric series), and the unit shapes what gets sampled, indexed and kept. A 10% rise in traffic can grow the bill by 40-50%, because cardinality (the number of unique label combinations on a metric), log verbosity, the share of logs indexed and overage rates multiply on top of traffic. Cost control is now a funded category, and security and observability companies are buying it up. More in How the money moves.
- OpenTelemetry: it made collection portable. It graduated in the CNCF (the Linux Foundation's home for cloud software) on May 21, 2026, and every major vendor now accepts its data. Lock-in moved up the stack, to billing units, dashboards, alert rules, query languages and years of history, none of which port. See Interfaces & standards.
- Signal: the hard work is deciding which data and which alerts matter. Alert fatigue tops the list of obstacles to faster incident response in Grafana Labs' practitioner survey, and in Splunk's 2025 survey 73% of respondents said they'd had outages caused by alerts that were ignored or suppressed. Google's own guidance aims for one alert per real incident; few teams get close.
- Convergence: it's only partly real. Security vendors are buying pipelines and observability companies, and observability vendors sell security on the same meters. But Datadog's whole security business is about a seventh the size of the SIEM businesses at Palo Alto Networks or CrowdStrike, and the new AI layer is separating from the data store rather than locking customers into it. More in Observability, security and AI: what's real.
- Rules: they're thinner than in payments or identity, and I'd say so plainly to anyone new. Nobody licenses observability. The rules bite on the data inside telemetry (personal, card and health data), on retention floors for some buyers, and through vendor SLAs that cap liability at service credits.
Monitoring in the compliance sense, watching transactions for money laundering, is a different job with the same word; it's in Identity and trust. Tracing and evaluating AI agents is in AI agents: orchestration platforms. Cloud security (CNAPP) gets its own guide later in this series; here I only cover where it meets observability.
The main players
These are the names behind the diagram's roles, layer by layer, in no particular order.
Collection and open source
- What they do
- Instrument code, collect and store telemetry without a licence fee
- Main players
- OpenTelemetry (CNCF), Prometheus, Grafana's open-source stack (Loki, Mimir, Tempo, Alloy), Fluent Bit, ClickHouse (ClickStack)
- What they control
- The formats and attribute names everyone else reads; the free alternative at every renewal
Pipelines
- What they do
- Filter, sample, redact and route telemetry before it's stored
- Main players
- Cribl, Chronosphere (Palo Alto Networks), Datadog Observability Pipelines, Bindplane (Dynatrace), Onum (CrowdStrike), Observo AI (SentinelOne)
- What they control
- What reaches each backend, and therefore most of the bill
Observability platforms, incumbents
- What they do
- Store, query, alert, investigate; sell many products on one agent
- Main players
- Datadog, Dynatrace, Splunk (Cisco), New Relic, Elastic
- What they control
- Pricing units, retained history, dashboards and alert rules
Observability platforms, open-source-led and newer
- What they do
- The same jobs, built on open formats or cheaper storage
- Main players
- Grafana Labs, Honeycomb, Coralogix, Dash0, groundcover
- What they control
- Price pressure; portability as a selling point
Cloud providers' tools
- What they do
- Default monitoring for each cloud, billed on the cloud invoice
- Main players
- AWS (CloudWatch, X-Ray), Microsoft (Azure Monitor), Google Cloud
- What they control
- The default choice, and commit burn-down
Incident management and on-call
- What they do
- Page the right person, run the incident, record the postmortem
- Main players
- PagerDuty, incident.io, FireHydrant (Freshworks), Atlassian (Opsgenie, closing), Datadog On-Call, Grafana IRM
- What they control
- Escalation rules and the incident timeline
Security convergence
- What they do
- Reuse the same telemetry for detection and cloud security
- Main players
- Wiz (Google), Palo Alto Networks, CrowdStrike, Datadog, Cisco (Splunk), SentinelOne
- What they control
- The CISO's budget and detection content
AI SRE
- What they do
- AI agents that investigate incidents and suggest or make fixes
- Main players
- Datadog (Bits AI SRE), Dynatrace (Dynatrace Intelligence, Bluebox), AWS DevOps Agent, Resolve AI, Traversal, Cleric
- What they control
- The hand-off from alert to code change
How they make money, and who's moving:
- Open source earns through the companies that sell managed versions or support. OpenTelemetry graduated with more than 12,000 contributors from more than 2,800 companies. Prometheus 3.0 (November 2024) accepts OpenTelemetry metrics natively, and Grafana donated its eBPF instrumentation to OpenTelemetry in May 2025. Fluent Bit claims more than 15 billion downloads (company figure).
- Pipelines charge per gigabyte processed or in credits, and pitch savings on the downstream bill. Cribl says it passed $300 million of annual recurring revenue (ARR) in 2025, up from $200 million (self-reported). Chronosphere had more than $160 million of ARR, growing triple digits, when Palo Alto agreed to buy it for $3.35 billion; the deal closed on January 29, 2026. Bindplane added $13 million of ARR to Dynatrace by June 2026. CrowdStrike bought Onum (about $290 million, reported) and SentinelOne bought Observo AI (about $225 million) in 2025.
- Incumbent platforms charge per host, gigabyte, million events and metric series, under annual commitments with overage on top, at gross margins of about 80%. Datadog's revenue grew 36% to $1.12 billion in the quarter to June 2026. Dynatrace's ARR grew 17% to $2.14 billion, with log consumption doubling. Elastic grew 15% to about $478 million a quarter and doesn't split observability out. Cisco's Observability line grew 4% to about $1.1 billion in fiscal 2026; most of Splunk's log business sits in its Security line. New Relic has been private since a take-private of about $6.5 billion in November 2023.
- Newer platforms often charge on a different unit to make a point. Grafana Labs says it passed $400 million of ARR and 7,000 customers in September 2025 (self-reported). Honeycomb prices per event, so adding detail costs nothing; its last disclosed round was in 2023. Coralogix raised at $1.6 billion in June 2026, Dash0 at $1 billion in March 2026 and groundcover $100 million in July 2026. ClickHouse raised at $15 billion in January 2026 and bought Langfuse, an LLM tracing tool, and Snowflake bought Observe in February 2026 for about $596 million, according to its filings.
- Cloud providers charge per gigabyte, metric and trace, and none discloses observability revenue. CloudWatch shows up on the AWS invoice and counts toward a customer's existing AWS commitment.
- Incident tools charge per user, usually $15-50 a month. PagerDuty's ARR was about $500 million in mid-2026, roughly flat. incident.io raised $62 million at about $400 million in April 2025, and Freshworks bought FireHydrant in December 2025. Atlassian ends support for Opsgenie on April 5, 2027, which puts its installed base up for grabs.
- Security vendors sell SIEM on volume and cloud security per host. Google closed its $32 billion purchase of Wiz on March 11, 2026. Palo Alto's XSIAM passed $700 million of ARR and CrowdStrike's Next-Gen SIEM $695 million (secondary coverage). Datadog's security products passed $100 million of ARR in 2025.
- AI SRE is priced as metered work: credits, agent-seconds or a monthly fee. Resolve AI raised $125 million at a $1 billion valuation on what TechCrunch's sources put at about $4 million of ARR, and Traversal raised $48 million in June 2025. AWS DevOps Agent became generally available on March 31, 2026, and Dynatrace opened Bluebox to early adopters in July 2026.
As of October 2026. This market moves through acquisitions every quarter, so treat the list as a map to check before relying on it.
Back to the dispatch company. A dispatcher assigns a job, and the request passes through the API gateway, the scheduling service, a Postgres database and a notification service that texts the technician. The OpenTelemetry SDK in each service creates a span and passes a trace ID along in a header, so the spans join into one trace. The Datadog Agent on each host collects spans, logs and metrics and sends them to a Cribl pipeline, which drops debug logs, keeps every trace with an error plus a sample of the rest, and copies the full log stream to an S3 bucket the company owns. Datadog indexes the errors and a fifth of the other logs. When the scheduling service burns through its error budget too fast, a Datadog monitor pages the on-call engineer through PagerDuty, Bits AI SRE drafts a guess at the cause, and the engineer opens the trace to find a slow query from that morning's deploy. The same logs reach Palo Alto's SIEM for the security team. Six or seven companies touch one slow drag-and-drop, and the dispatch company pays most of them by volume.
How telemetry becomes a fix, step by step
Here's the slow job assignment from the moment the code records it to the fix. Much of the data is sampled out or dropped before it's kept, usually to control cost. And an alert only helps if someone acts on it; noisy alerts get ignored, and that's where many incidents slip through.
- Emitted: the scheduling service records a latency metric, a log line and a span. Spans come from an SDK the team calls in code, from auto-instrumentation that hooks into common frameworks, or from eBPF, which watches network calls from the Linux kernel without code changes. Real-user monitoring (RUM) is emitted from users' browsers and phones. Breaks: a proxy or message queue drops the trace header and the trace splits in two; a host with a bad clock produces negative durations.
- Kept: a collector on each host batches the data and adds host and Kubernetes details, then a pipeline filters, samples, redacts and routes it. Head sampling decides cheaply at the start of a trace. Tail sampling waits until the trace is complete and keeps the errors and slow ones, but it holds whole traces in memory, and OpenTelemetry's docs say it can take "dozens or even hundreds" of servers at scale. The platform indexes some data (fast to search, expensive) and stores the rest cheaply. Breaks: head sampling at 1% also drops 99% of error traces; scaling a tail-sampling tier can split traces silently; a collector buffer fills while the backend is slow, and data is lost during the very incident you need it for.
- Alerted: a rule fires on a threshold, an anomaly or an SLO (service level objective, an internal target such as 99.9% of requests succeeding over 30 days). SLO alerts fire on burn rate, how fast the error budget (the 0.1% you're allowed to fail) is being spent. Breaks: Google's SRE Workbook shows that a plain alert on a 0.1% error rate over ten minutes could fire up to 144 times a day while the service still meets its SLO, which trains people to ignore the pager.
- Investigated: someone looks, usually going from a dashboard to a trace to the logs, then to recent deploys and feature-flag changes. Google budgets about six hours of engineering per incident, postmortem included. This is the step AI SRE products target. Breaks: the trace that would show the cause was sampled out; the logs must be pulled back from the archive; the observability vendor itself is down.
- Resolved: the team mitigates (roll back, fail over, turn off a flag), fixes properly later and writes a blameless postmortem with action items. Breaks: the action items never close, and the incident comes back a quarter later.
Two exits sit off that path. Dropped is data sampled out or filtered because keeping it was too costly; it can't be recovered. Ignored is an alert nobody acted on, usually because the same rule had fired falsely many times. In Splunk's 2025 survey, 52% of respondents said they struggle with high volumes of false alerts.
Metrics
- Question it answers
- Is it healthy? How much, how fast?
- What you pay for
- Active series (each unique label combination), data points
- High-cardinality labels
- Expensive: each combination is a new series
- Typical tools
- Prometheus, Mimir, Datadog, CloudWatch
Logs
- Question it answers
- What exactly happened, here, at this time?
- What you pay for
- Gigabytes ingested, events indexed, days retained
- High-cardinality labels
- Cheap to write, expensive to index
- Typical tools
- Elastic, Splunk, Loki, Datadog
Traces
- Question it answers
- Where in the request did the time or error go?
- What you pay for
- Spans or gigabytes ingested and indexed
- High-cardinality labels
- Cheap: adding a customer ID to a span costs little
- Typical tools
- Jaeger, Tempo, Datadog APM, Honeycomb, X-Ray
Profiles
- Question it answers
- Which lines of code burn CPU or memory?
- What you pay for
- Profiled hosts, gigabytes
- High-cardinality labels
- Not applicable
- Typical tools
- Pyroscope, Datadog, Elastic; OpenTelemetry profiles in alpha
Real-user monitoring
- Question it answers
- What did real users experience?
- What you pay for
- Sessions, cents to a few dollars per thousand
- High-cardinality labels
- Session-level
- Typical tools
- Datadog RUM, New Relic Browser, CloudWatch RUM
The cardinality column is worth remembering. A customer ID on a metric multiplies the series by the number of customers, while on a span it adds almost nothing, so per-customer questions belong in traces.
Collecting the data
Vendor agent (Datadog Agent, Dynatrace OneAgent)
- Setup effort
- Lowest; finds services on its own
- What you get
- Deep, curated integrations
- Portability
- Low
- Who maintains it
- The vendor
OpenTelemetry SDK, called in code
- Setup effort
- Highest
- What you get
- Exactly what you code
- Portability
- High
- Who maintains it
- Your app teams
OpenTelemetry auto-instrumentation
- Setup effort
- Low to medium
- What you get
- Framework-level spans
- Portability
- High
- Who maintains it
- The community and you
eBPF, no code changes
- Setup effort
- Low
- What you get
- Network-level calls (HTTP, gRPC, SQL), little business context
- Portability
- Medium to high
- Who maintains it
- A vendor or an OpenTelemetry group; pre-1.0
Agentless cloud integration
- Setup effort
- Lowest
- What you get
- The cloud's own metrics, 2-10 minutes late
- Portability
- Not applicable
- Who maintains it
- The vendor and the cloud; API calls billed to you
A vendor's OpenTelemetry distribution (Datadog DDOT, Elastic EDOT, Splunk, Grafana Alloy)
- Setup effort
- Low
- What you get
- OpenTelemetry plus vendor extras
- Portability
- Configuration ports, support doesn't
- Who maintains it
- The vendor
Most large companies end up with a mix: in Grafana's 2026 survey, 65% of respondents were investing in both Prometheus and OpenTelemetry.
Which observability setup? Five questions
- How much will you send? Estimate hosts, gigabytes of logs a day, metric series and spans, then price each against the vendor's units. The same telemetry on the same vendor can cost three times more depending on how much you index, so model the indexing share before the contract.
- Who will run it? Self-hosted open source has no licence fee, but at 500 hosts it usually takes two to five platform engineers, who cost more than the software saves until you're very large.
- How portable must you be? If you might switch vendors or run two, instrument with OpenTelemetry and keep a full copy of logs in storage you own. Dashboards, alert rules and query history won't move either way, so keep them as code.
- Who else wants the data? If the security team needs the same logs for its SIEM, a pipeline that routes one stream to both is cheaper than two collectors.
- Which rules apply? Card data, health data, US federal customers and EU residents each change retention, region and certification needs. Region is usually fixed when the account is created, so decide it first.
My defaults: for a B2B software company of a few hundred hosts, instrument new code with OpenTelemetry, use one commercial platform for breadth, and put a pipeline in front of it once logs pass about a terabyte a day. Index errors and the logs you search weekly; send the rest to cheap storage you own. Alert on SLO burn rates for user-facing services and review every page that wasn't actionable. And read the overage rates and the custom-metric definition before the headline price.
The primitives
01
Entity & identity
What is the unit of record, and how do we know it is the same one?
In telemetry the unit of identity is the resource: the service name, its version, the environment, the host, the container, the Kubernetes pod, the cloud account and region. OpenTelemetry standardizes these attribute names, and a trace adds a trace ID and span IDs that follow a request across services.
Vendors bill on entities too, and they don't agree on which entity counts:
| Billing entity | Who uses it | Rough list price (Oct 2026) |
|---|---|---|
| Host | Datadog, Splunk Observability | $15-75 per host a month, by product tier |
| Memory per host-hour | Dynatrace Full-Stack | About $58 a month for an 8 GiB host |
| User | New Relic, Grafana, incident tools | $8-349 per user a month, by role |
| Event | Honeycomb | From $150 a month for 50 million events |
| Active metric series | Grafana Cloud | About $6.50 per 1,000 series |
| Custom metric | Datadog | $5 per 100 a month |
| Gigabyte | Everyone, for logs | Cents to dollars, depending on indexing |
Containers, serverless functions and autoscaling make "how many hosts" a moving number, so vendors bill on percentiles of hourly counts or "host-equivalents". The identity most often missing is ownership. Which team owns checkout-api? That lives in a service catalog if anywhere, and an alert without an owner pages whoever happens to be on call.
"Agent" has a double meaning, often on one product page: the software on each host that collects telemetry, and an AI agent. I'd write "host agent" and "AI agent" every time. Telemetry also carries other people's identities, such as user IDs, IP addresses and emails in URLs, which is where the privacy rules come in.
02
State & lifecycle
What states exist, and what moves an entity between them?
A piece of telemetry moves through states: emitted, collected, processed, ingested, indexed or stored, tiered down, deleted. Each change of state is a pricing event, and regulators have started naming the states: PCI's card-data rules require logs to be "immediately available" for three months, and the US federal memo M-26-14 distinguishes "actively searchable" from "retrievable".
| Tier | What it means | Rough cost |
|---|---|---|
| Hot, indexed | Searchable in seconds, used by alerts and dashboards | Datadog: $1.06-2.50 per million log events, for 3-30 days |
| Warm or flex | Searchable, slower, priced separately | Datadog Flex: $0.05 per million events stored, plus compute |
| Archive | Object storage the customer often owns | About $0.02-0.03 per GB a month |
| Rehydrated | Archive pulled back into the index for an investigation or audit | A fee per GB, plus minutes to hours |
| Deleted | Retention expired, or an erasure request | Free, but hard in append-only stores |
Rehydration is fading. Since 2025, Datadog's Archive Search, Elastic's frozen tier and lakehouse engines let you query old data where it sits.
An alert goes from inactive to pending to firing, then acknowledged and resolved. An incident is detected, triaged, mitigated, resolved, reviewed and closed. Standards have states too: OpenTelemetry's traces, metrics and logs are stable, profiles reached alpha in March 2026, and the conventions for AI and LLM telemetry are still marked "Development".
The contract has states as well: free tier, annual commitment, monthly usage, overage, true-up, renewal. Seat reductions at Datadog take effect only at renewal, and overage history tends to become next year's commitment, so commitments ratchet up.
03
System of record & ledger
Who owns the truth, and how do systems reconcile?
There's no single system of record.
| Record | System of record |
|---|---|
| What happened in production (what was indexed) | The observability backend, for as long as it keeps it |
| The full log history | Often an object-storage bucket the customer owns |
| The incident timeline | The incident tool |
| Who owns each service | A service catalog, if one exists |
| What you'll be billed | The vendor's usage meter |
| Security and audit evidence | The SIEM or the log archive |
The billing ledger is the one I'd watch. The vendor's meter decides the bill, and customers rarely have an independent count. Datadog bills custom metrics on the monthly average of hourly counts, and hosts on the 99th percentile of hourly counts, so the top 1% of hours is forgiven and a week-long spike is billed. A pipeline gives the customer its own meter of what was sent, which helps in a billing dispute.
For security and compliance teams, logs are the record: PCI requires a year of audit logs, SOC 2 auditors want evidence, and a breach investigation depends on what was kept. For the reliability side the same logs are disposable signal. The two retention philosophies collide whenever one platform serves both teams.
04
Rules & policy
What logic decides outcomes, and who can change it?
Most rules in observability are ones the customer writes: what to sample, redact and index, for how long, when to alert and who gets paged. They live in collector and alerting configuration, and they decide both the bill and what an investigator will find.
SLOs matter most for signal. Google's SRE Workbook recommends paging on burn rate over two windows, so a short spike pages nobody and a real outage pages fast:
Page
- Long window
- 1 hour
- Short window
- 5 minutes
- Burn rate
- 14.4x
- Budget spent when it fires
- 2%
Page
- Long window
- 6 hours
- Short window
- 30 minutes
- Burn rate
- 6x
- Budget spent when it fires
- 5%
Ticket
- Long window
- 3 days
- Short window
- 6 hours
- Burn rate
- 1x
- Budget spent when it fires
- 10%
That's for a 99.9% target over 30 days. At a burn rate of 1,000, the whole month's budget is gone in about 43 minutes. The math breaks for low-traffic services: at ten requests an hour, one failure spends about 14% of the monthly budget, so teams add synthetic traffic, group small services or lower the target. An error-budget policy closes the loop: while the SLO is met, releases continue, and when the budget is spent, they pause.
Google's on-call guidance sets a bar few teams meet: every page actionable, at most two incidents in a 12-hour shift, and on-call capped at a quarter of an engineer's time. Vendors add their own rules (cardinality limits, rate limits, index quotas), which are product decisions about how hard to stop a customer's mistake.
05
Effective dating
Which version of the rule applied at that moment?
Everything in observability is measured over a window. Alerts evaluate over minutes, SLOs over a rolling 28 or 30 days, burn rates over 5 minutes to 3 days, and billing over hours, months and contract years. Retention clocks usually start when data arrives, so late mobile or batch data can land in the wrong window.
The meaning of a field has a date. OpenTelemetry's semantic conventions are versioned, and renames break things quietly: database spans renamed db.system to db.system.name in version 1.33.0 in 2025, and dashboards keyed on the old name stopped matching.
Dates I'd keep on a calendar as of October 2026:
| Change | Effective | Status (Oct 2026) |
|---|---|---|
| PCI DSS 4.0's future-dated requirements, including automated log review | March 31, 2025 | Live |
| Atlassian Opsgenie end of sale | June 4, 2025 | Done; support ends April 5, 2027 |
| California's cybersecurity audit rules, covering audit-log management | January 1, 2026 | Live; first certifications from April 1, 2028 for the largest businesses |
| OpenTelemetry graduates in the CNCF | May 21, 2026 | Done |
| OMB M-26-14 replaces M-21-31 for US federal logging | May 22, 2026 | Live |
| Datadog's AI SRE moves to AI credits | Around June 2026 | Live |
| FedRAMP 20x rules launched | June 25, 2026 | Mandatory January 1, 2027 |
| OpenTelemetry Governance Committee election results | October 30, 2026 | Upcoming |
| EU Data Act bans switching and egress fees | January 12, 2027 | Upcoming |
| HIPAA Security Rule update | Proposed January 2025 | Final rule targeted around July 2027; uncertain |
Take the dispatch company's commitment renewing in March. Overage from a cardinality spike in January sits in the usage history the vendor brings to the renewal, so a mistake fixed in a week can still raise next year's commitment.
06
Interfaces & standards
What format and protocol do counterparties speak?
The collection side is now mostly standard:
- OTLP, OpenTelemetry's protocol for sending traces, metrics and logs, is stable and accepted by every major vendor.
- W3C Trace Context is the
traceparentheader that carries a trace ID between services. - Semantic conventions name attributes the same way everywhere (HTTP stable since 2023, databases since 2025, AI still in development).
- Prometheus exposition format and PromQL are the de facto standard for metrics and the closest thing to a common query language; many backends accept PromQL.
The storage and query side has no standard. Logs and traces are queried in each vendor's language: LogQL and TraceQL at Grafana, SPL at Splunk, KQL at Microsoft, ES|QL at Elastic, DQL at Dynatrace, Datadog's own syntax, SQL on ClickHouse. Dashboards and alert rules are written in these languages, which is why they don't port.
Every major vendor now ships its own OpenTelemetry distribution, and the distributions recreate some lock-in. Splunk gives official support only for its own and "best-effort" support upstream. Datadog's wasn't FedRAMP or FIPS compliant when the research was done, which steers government buyers back to the proprietary agent. Tail sampling is "often tied to a specific vendor", in OpenTelemetry's own words. Datadog's FY2025 annual report doesn't mention OpenTelemetry at all; it sells "a single agent" with over 1,000 integrations. Its price list charges $0.50 per GB for spans ingested over OTLP against $0.10 for its own APM ingestion, and application metrics sent through OpenTelemetry generally count as custom metrics.
Vendors are also giving their collection technology away, which is what you do once collection stops being where you make money: Elastic donated its eBPF profiler and Grafana its Beyla. For AI, two newer interfaces matter. MCP lets AI agents query observability tools, and Datadog said MCP tool calls grew more than 22-fold from late 2025 to mid-2026 (company figure). GitHub issues and pull requests are the hand-off between observability AI and coding agents.
07
Networks & counterparties
Who sits between us and the outcome, and what do they want?
| Party | What they control | What stops if they fail |
|---|---|---|
| App teams | What gets instrumented, log levels, labels | Telemetry quality; the bill, through cardinality and verbosity |
| Platform or observability team | Collectors, pipelines, budgets, vendor choice | Collection, sometimes for everyone |
| Observability vendor | Storage, query, alerting, pricing units | Visibility during incidents |
| Pipeline vendor | Routing, sampling, redaction | Data to every destination at once |
| Cloud provider | Compute, network charges, default tools, marketplace billing | Everything running on it |
| Incident tool | Paging, escalation, incident records | Waking the right person |
| Security team and SIEM | Detection rules on the same logs | Security monitoring |
| AI agent vendors | Read access to telemetry, sometimes write access to code | Automated investigation |
Concentration runs both ways. Datadog's largest customer, widely reported to be OpenAI, renewed at nine figures in 2026 but is cutting usage from July, and Datadog built that into its guidance. Coinbase's roughly $65 million a year Datadog bill was restructured in 2023 after Coinbase built a parallel open-source stack. The biggest customers can credibly leave, and the threat works best at renewal.
The pipeline is the new switching point. Whoever owns it decides which backend gets which data, which is why security and observability vendors have both bought pipelines since 2025.
I'd write down early which single failure blinds us during an incident. For many teams it's the collector tier: in OpenTelemetry's 2026 survey of collector users, 23% didn't monitor their collectors at all.
08
Regulatory layering
Jurisdiction × activity × entity type: is it a license or a certification?
Observability has no licensing regime. Nobody needs permission to collect or store telemetry, and no regulator supervises observability vendors as such. Rules reach this industry through four other doors.
Data inside telemetry
- What it covers
- GDPR and CCPA for personal data, HIPAA for health data, PCI DSS for card data
- Binds
- The customer as controller; the vendor as processor by contract
- Enforced by
- Data protection authorities, the FTC and states, card networks through acquirers
The buyer
- What it covers
- FedRAMP for US federal agencies, OMB logging memos, SEC disclosure for listed companies
- Binds
- Vendors selling to those buyers; the buyers themselves
- Enforced by
- Agencies, auditors, the SEC
The market
- What it covers
- EU Data Act switching rules, competition review of deals
- Binds
- Vendors
- Enforced by
- The EU, competition authorities
Attestation
- What it covers
- SOC 2, ISO 27001
- Binds
- Vendors, as a sales requirement
- Enforced by
- Customers' procurement teams
GDPR and CCPA don't mention logs; they apply because logs carry IP addresses, user IDs, emails and tokens. SOC 2 is a contract artifact with no legal force, and in practice it's the report every vendor is asked for first.
The test I'd use on any requirement: is it about the data, the buyer or the vendor? Data rules follow the data into every tool, buyer rules only matter if you sell to that buyer, and vendor rules mostly live in contracts.
09
Exceptions & reversals
What goes wrong, and how is it undone?
| Exception | The way back | Clock |
|---|---|---|
| Data needed from the archive | Rehydrate into the index, or query in place | Minutes to hours, plus fees |
| Sampled or dropped data | None; it's gone | Never |
| Cardinality or log spike on the bill | Ask for a one-time credit; vendors sometimes grant one | A negotiation |
| Bad deploy | Roll back, the most common mitigation | Minutes |
| Alert flapping or false pages | Tune the rule or move to SLO burn rates | Weeks |
| Postmortem actions not done | Track them; Google says an unreviewed postmortem "might as well never have existed" | Next incident |
| Legal hold or security investigation | Freeze deletion | Until released |
| Forced tool migration | Move off Opsgenie (support ends April 5, 2027) or Grafana OnCall open source (archived March 24, 2026) | Months |
Rules reverse too. The US federal logging memo M-21-31 was rescinded on May 22, 2026, a federal court struck down part of HHS's guidance on tracking technologies in 2024, and the HIPAA Security Rule update has stalled. Product reversals follow public pressure: after the Storm-0558 intrusion in 2023, which agencies could only detect with logs from a premium tier, Microsoft gave standard customers more than 30 log types at no extra cost.
10
Liability allocation
When it fails, who pays?
The customer carries most of the risk in observability, and the contracts say so.
| Failure | Who pays first | How |
|---|---|---|
| Bill shock from a spike | The customer | Usage is billable by contract; credits are a goodwill gesture |
| Vendor outage | The customer, in lost visibility; the vendor, in credits and lost usage revenue | SLA service credits, the "sole and exclusive remedy" |
| Missed alert leads to an outage | The customer and its users | Nothing in the vendor's contract covers it |
| Personal data logged | The customer as controller | GDPR fines land on the controller, not the logging tool |
| Privileged host agent crashes machines | Customers and insurers, then the vendor if negligence is proven | Liability caps, tested in court |
| AI agent's wrong suggestion | The customer, through the human who approved it | Vendors' terms don't accept liability for remediations |
The SLAs are short. Grafana Cloud promises 99.5% of requests succeed each month and credits 10% to 100% of the affected service's monthly fee as availability drops; it also covers alert evaluation, so a missed alert is contractually defined there. New Relic targets 99.8%, with termination as the remedy if it misses badly two months running, and I found no public Datadog SLA to compare.
Datadog's own outage on March 8, 2023 shows the vendor side. A security update auto-applied across its nodes knocked out networking, and the company said it lost about $5 million of revenue, because under usage pricing it couldn't bill for data it couldn't ingest. Customers were blind for most of a day and got credits at best.
CrowdStrike's July 2024 outage is the extreme case for any privileged host agent, which observability agents are becoming as they take on security work. One estimate put direct losses to Fortune 500 companies at $5.4 billion, and CrowdStrike says its liability cap limits its exposure to "single-digit millions". A court let Delta's negligence claims proceed in May 2025.
What's different here
How the money moves
The customer pays almost everyone, mostly by volume. The app teams who decide what to log rarely see the bill; the platform team owns it.
| Flow | Who pays whom | Rough amount (Oct 2026, list) |
|---|---|---|
| Infrastructure and APM | Customer → platform | $15-75 per host a month, or by memory or user |
| Log ingest | Customer → platform | $0.10-0.60 per GB |
| Log indexing | Customer → platform | About $1-2.50 per million events at Datadog, for 3-30 days |
| Metrics | Customer → platform | $5 per 100 custom metrics (Datadog); about $6.50 per 1,000 series (Grafana); 5-30 cents per metric (CloudWatch) |
| Users and on-call seats | Customer → platform or incident tool | $15-50 per responder; up to about $350 for a full New Relic user |
| Pipeline | Customer → pipeline vendor | Roughly $0.10-0.20 per GB processed |
| Archive storage and network | Customer → cloud provider | About $0.02-0.03 per GB a month to store, plus charges for moving data between zones |
| Security on the same data | Customer → platform | Datadog: $5 per million events analyzed by its SIEM; $10-25 per host for cloud security |
| AI investigation | Customer → platform or cloud | A few dollars per investigation; AWS DevOps Agent about $30 per agent-hour |
Enterprise contracts are discounted from list, often a lot, and nobody publishes by how much. Overage is the opposite: above the commitment, Datadog's on-demand rates run 20-50% higher than its annual rates.
Gartner's research, as summarized by others, puts its median client's spend with a single observability vendor above $800,000 a year. At the top end, OpenAI was reported in 2025 to spend about $170 million a year on Datadog.
The vendors' growth model is the customer's bill seen from the other side. Datadog, Dynatrace and Elastic report net retention between about 110% and the low 120s, meaning existing customers spend that much more each year, and in Grafana's 2026 survey most respondents expecting higher spend blamed broader adoption more than price rises.
A year of observability for 500 hosts and 2 TB of logs a day
Back to the dispatch company: 500 hosts with APM on all of them, about 2 TB of logs a day (roughly 60 billion log events a month at an average of 1 KB each), half a million to 1.5 million metric series, and 10-20 TB of sampled traces a month. Searchable logs are kept 30 days, metrics 13 months, and a log archive for a year. Every figure is a list-price range for a year, before discounts.
| Option | Per year (list) | What drives it |
|---|---|---|
| Datadog, logs indexed with discipline (10-20% indexed, the rest in cheap storage) | About $0.6-1.05 million | Hosts, custom metrics, indexed share |
| Datadog, everything indexed for 15-30 days | About $1.6-2.6 million | Log indexing alone is $1.2-1.8 million |
| New Relic | About $0.5-0.7 million | Gigabytes and full users; no per-host fee |
| Dynatrace | About $0.55-0.9 million | Memory per host-hour; larger nodes cost more |
| Grafana Cloud (managed open source) | About $0.45-0.8 million | Gigabytes and active series |
| Self-hosted open source on a cloud | About $0.55-2 million | Two to five engineers at $200,000-300,000 each |
| AWS CloudWatch | About $0.6-1 million | Gigabytes and per-metric fees; counts toward the AWS commitment |
| Add a pipeline | Plus $70,000-150,000 | Gigabytes processed |
The striking gap is between the two Datadog rows: same telemetry, same vendor, about three times the bill, and the only difference is how much gets indexed. Under ingest-only pricing (New Relic, Grafana) savings from dropping data scale with bytes, so a pipeline matters less there.
As I read the numbers, self-hosting doesn't save money at this size once people are counted. It wins at Airbnb's scale, where moving metrics to OpenTelemetry and open-source storage at more than 100 million samples a second reportedly cut cost by about an order of magnitude. CloudWatch looks expensive on paper, and Gartner's clients complain about its cost, but teams pick it anyway because it draws down a commitment they've already signed.
Why the bill grew 47% when traffic grew 10%
Now the next quarter. Traffic grows 10%, and the disciplined Datadog setup starts at about $66,000 a month.
Hosts +10% (autoscaling follows traffic)
- Monthly cost before
- About $27,000
- After
- About $29,700
- Why
- The linear part
A team adds a customer_id or pod tag to three busy metrics
- Monthly cost before
- About $10,000 (250,000 custom metrics)
- After
- About $22,500 (500,000)
- Why
- Series multiply by every tag value
New services and retries add 25% more log events, and the indexed share creeps from 15% to 20%
- Monthly cost before
- About $15,300 (9 billion indexed)
- After
- About $25,500 (15 billion)
- Why
- Verbosity and indexing defaults
Usage passes the commitment, so the excess is billed at on-demand rates
- Monthly cost before
- After
- About $5,100 more
- Why
- Overage runs 50% above the committed rate for indexed logs
Total
- Monthly cost before
- About $66,000
- After
- About $97,000
- Why
- About +47% on +10% traffic
None of those drivers shows up on a traffic dashboard. A smaller example in the research, 200 hosts and 1.5 TB of logs a day, comes out at +39% on the same mechanics, so I'd tell a CFO "40-50%" and explain the four causes.
The fix is mostly in the pipeline and the tag allow-list: drop debug logs before ingest, index less, keep the customer tag ingested but unindexed, or move per-customer detail into traces. In the smaller example those moves bring the bill back to about 6% above baseline. The pipeline itself cost about 15% of the bill there. At index-heavy prices it pays for itself many times over if it cuts what gets indexed; cutting ingest alone barely covers its cost.
What the pipelines are worth
Cost control is a funded category, and it's being bought: Palo Alto, Dynatrace, CrowdStrike and SentinelOne all bought pipelines in 2025-2026 (see The main players). The platforms ship their own cost tools too, such as Grafana's Adaptive Metrics and Adaptive Logs, which Grafana says cut series by up to 80% and logs by up to 50%. Cribl's secondary-market value of about $3 billion in September 2026 sits below its 2024 round. The buyers seem to treat pipelines as a feature of a SIEM or a platform, and I'm not sure cost control survives as a standalone category.
Who holds the power
Power in observability follows whoever controls the meter, the history and the pipeline.
- Platform vendors hold the meter and the unit definitions, years of history, dashboards and alert rules written in their query languages, multi-product bundles (58% of Datadog's customers used four or more products by mid-2026) and an agent on every host. Each makes a partial exit harder.
- Customers gain power through OpenTelemetry, a pipeline they control, an archive in their own bucket, PromQL-compatible backends and the renewal date. In Grafana's survey, 37% of OpenTelemetry adopters cite freedom to switch vendors. The power shows up mostly at renewal and mostly for large buyers who can afford to run two stacks.
- Cloud providers hold the default, the invoice and the commitment. Their tools are one click away, their charges count toward commitments customers already signed, and AWS gives DevOps Agent credits worth 30-100% of a customer's support spend, which specialists can't match.
- Security vendors hold the CISO's budget and the detection content, and they're buying pipelines to control what reaches any SIEM.
- OpenTelemetry's maintainers, many of them vendor employees, decide what attributes are called. A rename costs everyone downstream.
- Coding-agent vendors are a new channel. Whoever feeds production context into Claude, Copilot, Cursor or AWS's Kiro gets embedded in how code is written.
- Regulators hold little direct power. They set retention floors for some buyers and data rules for what's inside the logs.
How the rules work
There's no observability regulator and no licence, so most of this section is about data rules and buyer rules that happen to land on telemetry. I'd still take them seriously, because the fines come from what's in the logs.
Personal data in logs. Telemetry carries IP addresses, user IDs, emails in URLs, auth tokens in headers and now LLM prompts. The largest GDPR fine I found about data sitting in internal systems is Meta's €91 million from Ireland's regulator in September 2024, for passwords stored in plaintext, partly for failing to notify and document the breach. France's CNIL recommends keeping access logs six months to a year, minimizing personal data in them, and storing them separately and append-only. Deleting one person's data from append-only stores is hard, so the common practice is to redact in the pipeline and keep retention short. The customer is the controller and carries the liability; the vendor is a processor.
Diagnostic files are the leak people forget. In 2023 an attacker reached support files for 134 Okta customers; some were browser recordings (HAR files) holding session tokens, and five customers' sessions were hijacked. Those files never went through a pipeline. Session replay in RUM tools has its own exposure: since a court struck down part of HHS's tracking guidance in 2024, claims over replay and tracking pixels have moved to state wiretap lawsuits against website operators.
Retention floors.
PCI DSS v4.0.1, 10.5.1
- Who it covers
- Systems in the card-data environment
- Retention
- 12 months, the last 3 immediately available
- Status (Oct 2026)
- In force; automated log review required since March 31, 2025
OMB M-26-14
- Who it covers
- US federal civilian agencies
- Retention
- 6 months actively searchable, 12 months retrievable
- Status (Oct 2026)
- In force since May 22, 2026
OMB M-21-31
- Who it covers
- US federal agencies
- Retention
- 12 months hot plus 18 cold
- Status (Oct 2026)
- Rescinded May 22, 2026
CNIL recommendation
- Who it covers
- France, under GDPR
- Retention
- 6-12 months for access logs
- Status (Oct 2026)
- Guidance
HIPAA audit controls
- Who it covers
- Covered entities and their vendors
- Retention
- No period for logs; policies kept 6 years
- Status (Oct 2026)
- Update proposed, stalled
California cybersecurity audits
- Who it covers
- Large businesses under CCPA
- Retention
- Audits cover audit-log management
- Status (Oct 2026)
- Effective January 1, 2026; first certifications 2028
The federal change is the clearest example of rules following the bill. M-26-14 cut the floor from 30 months to 12, and the memo itself says keeping "vast quantities of logging data without clear utility" was neither operationally feasible nor cost-effective. Its top maturity level grades signal: logs should generate actionable alerts covering at least 95% of baseline requirements. PCI's 12-and-3 maps neatly onto vendors' hot and archive tiers.
Buyer rules. FedRAMP decides who can sell to US federal agencies. Datadog for Government and Grafana's federal cloud say they're authorized at High, Dynatrace holds Moderate and said in July 2026 it would pursue High, and FedRAMP's new 20x process becomes mandatory on January 1, 2027. Listed US companies must file an 8-K within four business days of deciding a cyber incident is material, under SEC Item 1.05; trade groups want it rescinded, but I found no proposal as of October 2026. Telemetry is the evidence for scoping, and "we can't determine the scope" is an answer regulators read badly.
Residency and switching. Regions are fixed at setup. Datadog's sites are separate ("you cannot share data across sites"), and New Relic says the customer is "solely responsible" for choosing its region. Moving is a re-implementation. From January 12, 2027, the EU Data Act bars cloud and SaaS providers, observability vendors included, from charging EU customers switching or data-egress fees when they move.
SLAs and contracts. Most of the customer's protection sits here, and it's thin; see Liability allocation. EU financial firms have their own operational resilience rules (DORA), which I haven't covered.
What mistakes cost
- Bill shock: a tag or a debug flag can add tens of thousands of dollars a month within days, and overage history raises the next commitment.
- Ignored alerts: users find the outage first. In Splunk's UK cut of its 2025 survey, 54% said false alerts hurt morale.
- Cost-cutting blind spots: dropped logs and sampled traces lengthen the next investigation, and rehydrating the archive mid-incident costs time you don't have.
- Secrets in logs: Meta's €91 million fine was about passwords in internal systems, and Okta's 2023 breach ran through troubleshooting files.
- Logs you didn't keep: after Storm-0558, the US Cyber Safety Review Board found Microsoft had "no evidence or logs" for parts of its explanation.
- Vendor and agent failures: Datadog's 2023 outage and CrowdStrike's 2024 update, both in Liability allocation, left customers carrying most of the cost.
- The wrong region: moving an account to another region is a rebuild, with dashboards and alerts recreated by hand.
Observability, security and AI: what's real
My view as of October 2026: the convergence is real at two layers, the pipeline and the agent on the host, and weak at the revenue line, where security buyers still mostly buy from security vendors. AI investigation is improving but solves a minority of incidents on public benchmarks, and the AI layer is separating from the data store. If I were at an observability vendor, I'd treat the pipeline and the context (who owns what, what changed, how services connect) as what's worth defending, more than the raw data.
What's real in the security convergence
- Buying the data path: Palo Alto closed Chronosphere on January 29, 2026 and calls it "our next-generation observability platform"; its annual report now names Datadog, Dynatrace and Elastic as competitors. CrowdStrike and SentinelOne bought telemetry pipelines in 2025. Going the other way, Cribl bought detection-engineering and AI security-operations assets in mid-2026.
- Same meters: Datadog prices its SIEM at $5 per million events analyzed and cloud security at $10-25 per host, on the same agent and bill. On a company with the dispatch company's host count and about half its log volume, adding SIEM, cloud security and a sensitive-data scanner would run roughly 20-35% on top of the log bill. One of Datadog's mid-2026 wins was a bank adopting its SIEM for how it handles personal data.
- The host agent: Wiz, built agentless, now ships a runtime sensor on the workload, while Datadog sells workload security from its observability agent. Google closed Wiz for $32 billion on March 11, 2026, so a cloud provider now owns a cloud-security leader.
What's not there yet
- Scale: Datadog's security products passed $100 million of ARR in 2025. By their 2026 figures, Palo Alto's XSIAM and CrowdStrike's Next-Gen SIEM are each around $700 million, roughly seven times larger even allowing for the year's gap, and Datadog's security business is a few percent of its revenue by my estimate.
- Growth from owning both: Cisco owns Splunk on both sides, and in its fiscal 2026 Security grew 2% and Observability 4%.
- Different buyers: reliability teams measure time to recover, SLOs and cost per gigabyte; security teams measure detection coverage, how long an attacker goes unnoticed, and audit evidence. Security vendors stay ahead on detection rules and threat intelligence, and observability vendors win on consolidation and data handling.
What's real in AI SRE
AI SRE (site reliability engineering) products promise to turn an alert into a diagnosis, and sometimes a fix.
| Product | What it does | How it's priced (Oct 2026) |
|---|---|---|
| Datadog Bits AI SRE | Investigates alerts; marketed as "fully autonomous" detection, investigation and remediation since mid-2026 | AI credits, about $1.30 each; one listing estimates about 6.5 credits an investigation |
| Dynatrace Bluebox | Gives coding agents production context; opens GitHub issues with the trace, the offending lines and the blast radius | Free with 5 GiB of ingest; Pro $75 a month; Enterprise custom |
| AWS DevOps Agent | Investigates incidents across AWS, Azure and on-premises apps | $0.0083 per agent-second, about $4 for an eight-minute investigation, offset by support credits |
| Azure SRE Agent | Investigates incidents on Azure | Agent units plus tokens |
| PagerDuty SRE agent | Triage inside the incident tool; a "fully autonomous responder" in early access from late 2026 | Not public |
| Resolve AI, Traversal, Cleric | Startups that sit on top of existing tools | Not public |
The benchmarks are sobering. On OpenRCA, a public root-cause benchmark presented in 2025, the best model solved about 11% of 335 failure cases, and ITBench put AI agents at 11-14% of its SRE scenarios. A 2026 follow-up, OpenRCA 2.0, found about 21% exact root causes across 11 frontier models. Scores are rising from one paper to the next, but I'd treat any claim of autonomous root-cause analysis as running ahead of the evidence.
Dynatrace shows the gap in its own messaging. Its CTO has talked about taking humans out of the loop while also saying nobody has moved from AI-assisted to autonomous operations yet, and Bluebox's documentation says it "never pushes a fix on its own": it opens an issue, a coding agent proposes the change, and a person merges it. That's a policy statement more than a technical guarantee, and the liability sits with whoever approved the change.
Practitioners want to see the reasoning: in Grafana's 2026 survey, 95% said AI must show how it reached its answer, and the top barrier was having to feed it context by hand.
Where the AI layer is heading
- Away from the store: Datadog separated its AI security analyst from its SIEM "so customers can use it with other SIEMs", AWS DevOps Agent investigates apps outside AWS, and the startups store little telemetry. If the AI works on anyone's data, the store is less of a lock.
- Into coding agents: Bluebox, Datadog's Bits Code and AWS's Kiro feed production data into the tools that write code. The bottleneck moves to reviewing changes before production, the theme of Becoming AI native.
- AI workloads as telemetry: Dynatrace says more than 1,000 customers monitor AI workloads with it, and LLM prompts make traces bigger. OpenTelemetry's AI conventions are still in development, so vendors fill the gap with their own SDKs; agent tracing and evaluation are in AI agents: orchestration platforms.
- Metered pricing: credits, agent-seconds and tokens are priced like telemetry, and Datadog's move from per-investigation bundles to credits in mid-2026 cut the list price per investigation by about three-quarters by my arithmetic.
Questions to ask about an AI investigation agent
- On your last hundred real incidents, how often was the agent's initial answer the root cause, and who checked?
- What data does it read, through which interface, and does it copy telemetry into tickets or chat outside our retention policy?
- Can it change anything in production, and what's the approval step and the audit trail when it does?
- Does it work on data stored with another vendor, or only yours?
- How is it priced per investigation at our volume, including the query charges it triggers?
What usually goes wrong
| Symptom | Likely cause | First thing to check |
|---|---|---|
| Bill up 40-50% on 10% more traffic | A new high-cardinality tag, debug logs left on, indexed share creeping up, overage at on-demand rates | Usage by product, top custom metrics by tag, indexed share |
| Self-hosted Prometheus runs out of memory | Cardinality explosion from a label like user ID or raw URL path | Active series by label |
| Traces arrive in pieces | A proxy or queue drops the trace header | Whether traceparent survives each hop |
| Error traces missing during an incident | Head sampling dropped them | The sampling policy; tail sampling for errors |
| Pager fires dozens of times a day and gets ignored | Threshold alerts on noisy metrics | Alerts per real incident; SLO burn-rate alerts |
| Customers report the outage first | No alert, a misconfigured one, or an ignored one | The detection section of the postmortem |
| Data gaps with no errors anywhere | Collector buffers full, or collectors not monitored | The collectors' own metrics |
| Double bill or duplicate spans during migration | Vendor agent and OpenTelemetry both instrumenting the same app | Which one instruments each service |
| Emails or tokens found in logs | No redaction before storage | Redaction rules in the collector or pipeline |
| Auditor asks for last year's card-system logs | Only the hot index was kept | Retention tiers against PCI's 12 months |
Words that mean something else here
| Term | What you'd assume | What it means here |
|---|---|---|
| Agent | An AI agent | Usually the collector software on a host; now also an AI agent, often on the same product page |
| Event | Something that happened | Honeycomb: a wide record (a span is one). Datadog: a log line or a change notification. Security: a SIEM record |
| Cardinality | A row count | The number of unique values of a label, and so the number of metric series |
| Custom metric | A metric you designed | At Datadog, any metric not from one of its integrations, counted per unique name and tag combination |
| Host | A server you own | A billing unit: a VM, a Kubernetes node, or a "host-equivalent" for containers |
| Ingest vs index | The same thing | Ingest is receiving data, priced per GB; indexing makes it searchable for a period, priced per event |
| Sampling | Taking a sample to inspect | Dropping data on purpose, by a rule |
| Pipeline | A CI/CD pipeline | A telemetry router that filters, samples and routes |
| Monitor | Watching something | Datadog's word for an alert rule. In compliance, transaction monitoring (Identity and trust) |
| SLO, SLA, SLI | Interchangeable | SLI is the measurement, SLO the internal target, SLA the contract with credits |
| Error budget | Money | The failure you're allowed (1 − the SLO) over a window |
| Burn rate | Startup cash burn | How fast the error budget is being spent; 1x is exactly on budget |
| Incident | One thing | A reliability incident, a security incident, or an SEC "cybersecurity incident" with a materiality test |
| Autonomous | Acts on its own | Usually "investigates without being asked"; few products merge and deploy without a person |
| Credit | A refund | An SLA service credit, or a prepaid unit of AI usage |
What surprised me
Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.
"Observability is a tool you buy." Coming from payments, I pictured a product with a price. It turned out to be a set of choices about what to keep, and the same telemetry on the same vendor could cost three times as much depending on how much we indexed.
"Open standards end lock-in." OpenTelemetry made the collection portable. The dashboards, alert rules, query language and two years of history stayed exactly where they were.
"More data, fewer outages." The incidents that hurt were the ones where an alert fired and nobody looked, because it had fired falsely all month.
"Security and observability are merging." The deals say so. The revenue says the security vendors' SIEMs are still several times bigger than the observability vendors' security lines.
"This is a regulated industry like the others." There's no licence and no supervisor. The rules that bit were about what we'd written into the logs.
Sources
Undated entries were read on October 8, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator.
Standards, open source and research
- CNCF, OpenTelemetry's graduation (May 2026)
- OpenTelemetry: specification status; sampling; collector follow-up survey (2026); eBPF instrumentation goals for 2026; Governance Committee elections (2026, search result)
- ClickHouse, OpenTelemetry semantic conventions; Greptime, GenAI semantic conventions (May 2026, search result)
- Prometheus, Prometheus 3.0 (Nov 2024)
- Grafana Labs, Beyla donated to OpenTelemetry (May 2025); Elastic, profiling agent accepted by OpenTelemetry
- Chronosphere, Fluent Bit v4 (Apr 2025)
- Google SRE: Alerting on SLOs; Monitoring Distributed Systems; Being On-Call; Postmortem Culture
- AI root-cause benchmarks: OpenRCA (ICLR 2025, search result); OpenRCA 2.0 (2026, search result); ITBench (2025, search result); ITBench leaderboard (search result)
Rules and regulators
- OMB, M-26-14 on agency logging (May 2026); cloud.gov, M-21-31 compliance (search result)
- Data Protection Commission, €91 million fine on Meta (Sep 2024, search result)
- CNIL, logging recommendation (Oct 2021, search result)
- PCI DSS v4 requirement 10.5.1, control summary (search result)
- Hunton: California's cybersecurity audit regulations (2025, search result); HHS tracking guidance ruling (Jun 2024, search result)
- Baird Holm, HHS withdraws its appeal (2024, search result)
- Fierce Healthcare, HIPAA Security Rule pushed to July 2027 (search result)
- Latham & Watkins, EU Data Act switching requirements (Aug 2025, search result)
- SEC, petition comments on cyber disclosure, File 4-856 (2026, search result); Rane, the future of the SEC's cybersecurity rules (search result)
- FedRAMP: A-LIGN, FedRAMP 20x (2026, search result); Datadog, FedRAMP High (search result); Grafana Labs, FedRAMP High (Apr 2025, search result); Dynatrace, intent to pursue FedRAMP High (Jul 2026, search result)
Company filings and results
- Datadog: FY2025 10-K (Feb 2026); Q2 2026 results (Aug 2026); Q2 2026 earnings call transcript (Aug 2026); Q3 2025 call transcript (search result)
- Dynatrace, Q1 FY27 results (Aug 2026); MarketBeat, Dynatrace Q1 call highlights (Aug 2026, search result)
- Elastic, Q1 FY27 results (Aug 2026, search result)
- Cisco: FY2026 results (Aug 2026); 10-Q for the quarter to April 2026
- PagerDuty, Q2 FY27 results (2026, search result)
- Palo Alto Networks: FY2026 10-K (Sep 2026); Chronosphere acquisition completed (Jan 2026); Chronosphere deal announcement (Nov 2025, search result); Zacks, XSIAM ARR (search result)
- CrowdStrike: Q2 FY27 results (Aug 2026, search result); Eastern Progress, Next-Gen SIEM momentum (search result); GovInfoSecurity, CrowdStrike buys Onum (Aug 2025, search result)
- SecurityWeek, SentinelOne to acquire Observo AI (Sep 2025, search result)
- Google, completes acquisition of Wiz (Mar 2026, search result); Wiz, runtime sensor (search result)
- Snowflake: Observe acquisition closed (2026); FY2026 10-K (search result)
Private companies and deals
- Cribl: $200 million ARR (Jan 2025); Axios Pro, $300 million ARR and secondary offering (Feb 2026); StockAnalysis profile (Sep 2026, search result); SaaSRise, Cribl buys Radiant Security's AI SOC technology (Aug 2026, search result)
- Grafana Labs, $400 million ARR and 7,000 customers (Sep 2025, search result)
- Bloomberg Government, ClickHouse at $15 billion (Jan 2026, search result); ClickHouse, ClickStack in ClickHouse Cloud
- SecurityWeek, Coralogix raises $200 million at $1.6 billion (Jun 2026)
- The Next Web, Dash0's $110 million Series B (Mar 2026)
- SiliconANGLE, groundcover raises $100 million (Jul 2026)
- BigDATAwire, New Relic to be acquired by Francisco Partners and TPG (2023, search result)
- TechCrunch, incident.io raises $62 million (Apr 2025, search result)
- Freshworks, acquisition of FireHydrant (Dec 2025, search result)
- Atlassian, Opsgenie licensing; Grafana Labs, OnCall open source in maintenance mode (Mar 2025) and OnCall OSS docs
- TechCrunch, Resolve AI confirms $125 million raise (Feb 2026); BuiltIn NYC, Traversal raises $48 million (Jun 2025); Business Wire, Traversal strategic investment (Mar 2026, search result); Business Wire, Cleric launch (Dec 2025)
Pricing, product and legal pages
- Datadog: price list; custom metrics billing; billing; DDOT collector; Flex Frozen, Archive Search and CloudPrem (Jun 2025); sites (search result); March 2023 incident
- Dynatrace: pricing; Bluebox on Dynatrace Hub
- Bluebox: product and pricing
- Grafana Labs: pricing; Grafana Cloud SLA
- New Relic: pricing; EU region (search result); service level commitment (search result)
- Splunk: Observability pricing; OpenTelemetry Collector distribution
- Honeycomb, pricing; incident.io, pricing; PagerDuty, incident management pricing (search result)
- AWS: CloudWatch pricing; DevOps Agent pricing; DevOps Agent generally available (Mar 2026, search result); Kiro, DevOps Agent and Bluebox (search result)
- Microsoft, Azure SRE Agent pricing
Surveys, press and analysis
- Grafana Labs: Observability Survey 2025; 2026 survey release (Mar 2026); open standards in the 2026 survey
- Splunk, State of Observability 2025 (Oct 2025); IT Brief UK, Splunk's UK findings
- Catchpoint, summary of Gartner's observability spend research; AWSInsider, Gartner on CloudWatch costs (Jul 2025)
- The Pragmatic Engineer: inside the Datadog outage (2023); Datadog's $65 million customer (May 2023, search result); The Pulse #148, OpenAI's Datadog spend (Oct 2025, search result)
- SaaSRise, Datadog's biggest AI client cuts usage (2026, search result)
- InfoQ, Airbnb's OpenTelemetry metrics pipeline (Apr 2026, search result)
- Okta, support system breach root cause (Nov 2023, search result)
- Dark Reading, Microsoft offers free logging (2023, search result); TechTarget, Cyber Safety Review Board on Microsoft (Mar 2024, search result)
- CrowdStrike outage: Axios, Fortune 500 impact (Jul 2024, search result); The Register, Delta's lawsuit allowed to proceed (May 2025, search result)
- TechTarget, Dynatrace executives on Bluebox (2026)
- Datadog AI credits change, pricing tracker (Jun 2026, search result)
Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.