The Platform PM

Field GuideLast reviewed October 2026

Observability: metrics, logs and traces

How companies see what their software does in production: collecting metrics, logs and traces, why what you index sets the bill, what OpenTelemetry changed, why noisy alerts let incidents through, and what's real about security and AI using the same data.

The industry on one page

The parties. Your services emit telemetry, collectors send it on, and a pipeline decides what to keep before the platform stores it, alerts the on-call team and helps them fix the service. The same data now feeds security and AI products, which is pulling those markets closer.

Picture a company that sells scheduling and dispatch software to plumbing and electrical contractors. It runs about 500 hosts on one cloud and writes about 2 TB of logs a day. When a dispatcher drags a job onto a technician's calendar and nothing happens for eight seconds, somebody at the dispatch company needs to know within minutes, find out why and fix it. Observability is everything that makes that possible, and it's one of the larger software bills the company pays.

Your services emit telemetry: metrics (numbers over time, like error rate), logs (lines of text describing what happened), traces (the path of one request across services, made of one span per step) and, increasingly, profiles of which code burns CPU. Collectors pick it up: a vendor's agent on each host, or OpenTelemetry, the open standard for instrumenting code and moving telemetry. They send it to a pipeline that filters, samples (keeps a fraction on purpose) and routes it. The pipeline decides what's kept before the platform stores it, and it's where the bill is set, because everything after it is priced by volume. The platform stores, queries and alerts the on-call team, the engineers whose phones ring, who use dashboards and traces to fix the service. The same data now feeds security and AI products: a SIEM (the security team's log analysis and detection system), cloud security tools, and AI agents that investigate incidents. That reuse is pulling the observability, security and AI markets closer.

What I'd want a new PM in this space to take away:

  1. The bill: what you index sets the bill. Vendors price on volume units (hosts, gigabytes, events, metric series), and the unit shapes what gets sampled, indexed and kept. A 10% rise in traffic can grow the bill by 40-50%, because cardinality (the number of unique label combinations on a metric), log verbosity, the share of logs indexed and overage rates multiply on top of traffic. Cost control is now a funded category, and security and observability companies are buying it up. More in How the money moves.
  2. OpenTelemetry: it made collection portable. It graduated in the CNCF (the Linux Foundation's home for cloud software) on May 21, 2026, and every major vendor now accepts its data. Lock-in moved up the stack, to billing units, dashboards, alert rules, query languages and years of history, none of which port. See Interfaces & standards.
  3. Signal: the hard work is deciding which data and which alerts matter. Alert fatigue tops the list of obstacles to faster incident response in Grafana Labs' practitioner survey, and in Splunk's 2025 survey 73% of respondents said they'd had outages caused by alerts that were ignored or suppressed. Google's own guidance aims for one alert per real incident; few teams get close.
  4. Convergence: it's only partly real. Security vendors are buying pipelines and observability companies, and observability vendors sell security on the same meters. But Datadog's whole security business is about a seventh the size of the SIEM businesses at Palo Alto Networks or CrowdStrike, and the new AI layer is separating from the data store rather than locking customers into it. More in Observability, security and AI: what's real.
  5. Rules: they're thinner than in payments or identity, and I'd say so plainly to anyone new. Nobody licenses observability. The rules bite on the data inside telemetry (personal, card and health data), on retention floors for some buyers, and through vendor SLAs that cap liability at service credits.

Monitoring in the compliance sense, watching transactions for money laundering, is a different job with the same word; it's in Identity and trust. Tracing and evaluating AI agents is in AI agents: orchestration platforms. Cloud security (CNAPP) gets its own guide later in this series; here I only cover where it meets observability.

The main players

These are the names behind the diagram's roles, layer by layer, in no particular order.

Collection and open source

What they do
Instrument code, collect and store telemetry without a licence fee
Main players
OpenTelemetry (CNCF), Prometheus, Grafana's open-source stack (Loki, Mimir, Tempo, Alloy), Fluent Bit, ClickHouse (ClickStack)
What they control
The formats and attribute names everyone else reads; the free alternative at every renewal

Pipelines

What they do
Filter, sample, redact and route telemetry before it's stored
Main players
Cribl, Chronosphere (Palo Alto Networks), Datadog Observability Pipelines, Bindplane (Dynatrace), Onum (CrowdStrike), Observo AI (SentinelOne)
What they control
What reaches each backend, and therefore most of the bill

Observability platforms, incumbents

What they do
Store, query, alert, investigate; sell many products on one agent
Main players
Datadog, Dynatrace, Splunk (Cisco), New Relic, Elastic
What they control
Pricing units, retained history, dashboards and alert rules

Observability platforms, open-source-led and newer

What they do
The same jobs, built on open formats or cheaper storage
Main players
Grafana Labs, Honeycomb, Coralogix, Dash0, groundcover
What they control
Price pressure; portability as a selling point

Cloud providers' tools

What they do
Default monitoring for each cloud, billed on the cloud invoice
Main players
AWS (CloudWatch, X-Ray), Microsoft (Azure Monitor), Google Cloud
What they control
The default choice, and commit burn-down

Incident management and on-call

What they do
Page the right person, run the incident, record the postmortem
Main players
PagerDuty, incident.io, FireHydrant (Freshworks), Atlassian (Opsgenie, closing), Datadog On-Call, Grafana IRM
What they control
Escalation rules and the incident timeline

Security convergence

What they do
Reuse the same telemetry for detection and cloud security
Main players
Wiz (Google), Palo Alto Networks, CrowdStrike, Datadog, Cisco (Splunk), SentinelOne
What they control
The CISO's budget and detection content

AI SRE

What they do
AI agents that investigate incidents and suggest or make fixes
Main players
Datadog (Bits AI SRE), Dynatrace (Dynatrace Intelligence, Bluebox), AWS DevOps Agent, Resolve AI, Traversal, Cleric
What they control
The hand-off from alert to code change

How they make money, and who's moving:

  • Open source earns through the companies that sell managed versions or support. OpenTelemetry graduated with more than 12,000 contributors from more than 2,800 companies. Prometheus 3.0 (November 2024) accepts OpenTelemetry metrics natively, and Grafana donated its eBPF instrumentation to OpenTelemetry in May 2025. Fluent Bit claims more than 15 billion downloads (company figure).
  • Pipelines charge per gigabyte processed or in credits, and pitch savings on the downstream bill. Cribl says it passed $300 million of annual recurring revenue (ARR) in 2025, up from $200 million (self-reported). Chronosphere had more than $160 million of ARR, growing triple digits, when Palo Alto agreed to buy it for $3.35 billion; the deal closed on January 29, 2026. Bindplane added $13 million of ARR to Dynatrace by June 2026. CrowdStrike bought Onum (about $290 million, reported) and SentinelOne bought Observo AI (about $225 million) in 2025.
  • Incumbent platforms charge per host, gigabyte, million events and metric series, under annual commitments with overage on top, at gross margins of about 80%. Datadog's revenue grew 36% to $1.12 billion in the quarter to June 2026. Dynatrace's ARR grew 17% to $2.14 billion, with log consumption doubling. Elastic grew 15% to about $478 million a quarter and doesn't split observability out. Cisco's Observability line grew 4% to about $1.1 billion in fiscal 2026; most of Splunk's log business sits in its Security line. New Relic has been private since a take-private of about $6.5 billion in November 2023.
  • Newer platforms often charge on a different unit to make a point. Grafana Labs says it passed $400 million of ARR and 7,000 customers in September 2025 (self-reported). Honeycomb prices per event, so adding detail costs nothing; its last disclosed round was in 2023. Coralogix raised at $1.6 billion in June 2026, Dash0 at $1 billion in March 2026 and groundcover $100 million in July 2026. ClickHouse raised at $15 billion in January 2026 and bought Langfuse, an LLM tracing tool, and Snowflake bought Observe in February 2026 for about $596 million, according to its filings.
  • Cloud providers charge per gigabyte, metric and trace, and none discloses observability revenue. CloudWatch shows up on the AWS invoice and counts toward a customer's existing AWS commitment.
  • Incident tools charge per user, usually $15-50 a month. PagerDuty's ARR was about $500 million in mid-2026, roughly flat. incident.io raised $62 million at about $400 million in April 2025, and Freshworks bought FireHydrant in December 2025. Atlassian ends support for Opsgenie on April 5, 2027, which puts its installed base up for grabs.
  • Security vendors sell SIEM on volume and cloud security per host. Google closed its $32 billion purchase of Wiz on March 11, 2026. Palo Alto's XSIAM passed $700 million of ARR and CrowdStrike's Next-Gen SIEM $695 million (secondary coverage). Datadog's security products passed $100 million of ARR in 2025.
  • AI SRE is priced as metered work: credits, agent-seconds or a monthly fee. Resolve AI raised $125 million at a $1 billion valuation on what TechCrunch's sources put at about $4 million of ARR, and Traversal raised $48 million in June 2025. AWS DevOps Agent became generally available on March 31, 2026, and Dynatrace opened Bluebox to early adopters in July 2026.

As of October 2026. This market moves through acquisitions every quarter, so treat the list as a map to check before relying on it.

Back to the dispatch company. A dispatcher assigns a job, and the request passes through the API gateway, the scheduling service, a Postgres database and a notification service that texts the technician. The OpenTelemetry SDK in each service creates a span and passes a trace ID along in a header, so the spans join into one trace. The Datadog Agent on each host collects spans, logs and metrics and sends them to a Cribl pipeline, which drops debug logs, keeps every trace with an error plus a sample of the rest, and copies the full log stream to an S3 bucket the company owns. Datadog indexes the errors and a fifth of the other logs. When the scheduling service burns through its error budget too fast, a Datadog monitor pages the on-call engineer through PagerDuty, Bits AI SRE drafts a guess at the cause, and the engineer opens the trace to find a slow query from that morning's deploy. The same logs reach Palo Alto's SIEM for the security team. Six or seven companies touch one slow drag-and-drop, and the dispatch company pays most of them by volume.

How telemetry becomes a fix, step by step

From telemetry to a fix. Much of the data is sampled out or dropped before it's kept, usually to control cost. An alert only helps if someone acts on it; noisy alerts get ignored, and that's where many incidents slip through.

Here's the slow job assignment from the moment the code records it to the fix. Much of the data is sampled out or dropped before it's kept, usually to control cost. And an alert only helps if someone acts on it; noisy alerts get ignored, and that's where many incidents slip through.

  1. Emitted: the scheduling service records a latency metric, a log line and a span. Spans come from an SDK the team calls in code, from auto-instrumentation that hooks into common frameworks, or from eBPF, which watches network calls from the Linux kernel without code changes. Real-user monitoring (RUM) is emitted from users' browsers and phones. Breaks: a proxy or message queue drops the trace header and the trace splits in two; a host with a bad clock produces negative durations.
  2. Kept: a collector on each host batches the data and adds host and Kubernetes details, then a pipeline filters, samples, redacts and routes it. Head sampling decides cheaply at the start of a trace. Tail sampling waits until the trace is complete and keeps the errors and slow ones, but it holds whole traces in memory, and OpenTelemetry's docs say it can take "dozens or even hundreds" of servers at scale. The platform indexes some data (fast to search, expensive) and stores the rest cheaply. Breaks: head sampling at 1% also drops 99% of error traces; scaling a tail-sampling tier can split traces silently; a collector buffer fills while the backend is slow, and data is lost during the very incident you need it for.
  3. Alerted: a rule fires on a threshold, an anomaly or an SLO (service level objective, an internal target such as 99.9% of requests succeeding over 30 days). SLO alerts fire on burn rate, how fast the error budget (the 0.1% you're allowed to fail) is being spent. Breaks: Google's SRE Workbook shows that a plain alert on a 0.1% error rate over ten minutes could fire up to 144 times a day while the service still meets its SLO, which trains people to ignore the pager.
  4. Investigated: someone looks, usually going from a dashboard to a trace to the logs, then to recent deploys and feature-flag changes. Google budgets about six hours of engineering per incident, postmortem included. This is the step AI SRE products target. Breaks: the trace that would show the cause was sampled out; the logs must be pulled back from the archive; the observability vendor itself is down.
  5. Resolved: the team mitigates (roll back, fail over, turn off a flag), fixes properly later and writes a blameless postmortem with action items. Breaks: the action items never close, and the incident comes back a quarter later.

Two exits sit off that path. Dropped is data sampled out or filtered because keeping it was too costly; it can't be recovered. Ignored is an alert nobody acted on, usually because the same rule had fired falsely many times. In Splunk's 2025 survey, 52% of respondents said they struggle with high volumes of false alerts.

Metrics

Question it answers
Is it healthy? How much, how fast?
What you pay for
Active series (each unique label combination), data points
High-cardinality labels
Expensive: each combination is a new series
Typical tools
Prometheus, Mimir, Datadog, CloudWatch

Logs

Question it answers
What exactly happened, here, at this time?
What you pay for
Gigabytes ingested, events indexed, days retained
High-cardinality labels
Cheap to write, expensive to index
Typical tools
Elastic, Splunk, Loki, Datadog

Traces

Question it answers
Where in the request did the time or error go?
What you pay for
Spans or gigabytes ingested and indexed
High-cardinality labels
Cheap: adding a customer ID to a span costs little
Typical tools
Jaeger, Tempo, Datadog APM, Honeycomb, X-Ray

Profiles

Question it answers
Which lines of code burn CPU or memory?
What you pay for
Profiled hosts, gigabytes
High-cardinality labels
Not applicable
Typical tools
Pyroscope, Datadog, Elastic; OpenTelemetry profiles in alpha

Real-user monitoring

Question it answers
What did real users experience?
What you pay for
Sessions, cents to a few dollars per thousand
High-cardinality labels
Session-level
Typical tools
Datadog RUM, New Relic Browser, CloudWatch RUM

The cardinality column is worth remembering. A customer ID on a metric multiplies the series by the number of customers, while on a span it adds almost nothing, so per-customer questions belong in traces.

Collecting the data

Vendor agent (Datadog Agent, Dynatrace OneAgent)

Setup effort
Lowest; finds services on its own
What you get
Deep, curated integrations
Portability
Low
Who maintains it
The vendor

OpenTelemetry SDK, called in code

Setup effort
Highest
What you get
Exactly what you code
Portability
High
Who maintains it
Your app teams

OpenTelemetry auto-instrumentation

Setup effort
Low to medium
What you get
Framework-level spans
Portability
High
Who maintains it
The community and you

eBPF, no code changes

Setup effort
Low
What you get
Network-level calls (HTTP, gRPC, SQL), little business context
Portability
Medium to high
Who maintains it
A vendor or an OpenTelemetry group; pre-1.0

Agentless cloud integration

Setup effort
Lowest
What you get
The cloud's own metrics, 2-10 minutes late
Portability
Not applicable
Who maintains it
The vendor and the cloud; API calls billed to you

A vendor's OpenTelemetry distribution (Datadog DDOT, Elastic EDOT, Splunk, Grafana Alloy)

Setup effort
Low
What you get
OpenTelemetry plus vendor extras
Portability
Configuration ports, support doesn't
Who maintains it
The vendor

Most large companies end up with a mix: in Grafana's 2026 survey, 65% of respondents were investing in both Prometheus and OpenTelemetry.

Which observability setup? Five questions

  1. How much will you send? Estimate hosts, gigabytes of logs a day, metric series and spans, then price each against the vendor's units. The same telemetry on the same vendor can cost three times more depending on how much you index, so model the indexing share before the contract.
  2. Who will run it? Self-hosted open source has no licence fee, but at 500 hosts it usually takes two to five platform engineers, who cost more than the software saves until you're very large.
  3. How portable must you be? If you might switch vendors or run two, instrument with OpenTelemetry and keep a full copy of logs in storage you own. Dashboards, alert rules and query history won't move either way, so keep them as code.
  4. Who else wants the data? If the security team needs the same logs for its SIEM, a pipeline that routes one stream to both is cheaper than two collectors.
  5. Which rules apply? Card data, health data, US federal customers and EU residents each change retention, region and certification needs. Region is usually fixed when the account is created, so decide it first.

My defaults: for a B2B software company of a few hundred hosts, instrument new code with OpenTelemetry, use one commercial platform for breadth, and put a pipeline in front of it once logs pass about a terabyte a day. Index errors and the logs you search weekly; send the rest to cheap storage you own. Alert on SLO burn rates for user-facing services and review every page that wasn't actionable. And read the overage rates and the custom-metric definition before the headline price.

The primitives

01

Entity & identity

What is the unit of record, and how do we know it is the same one?

In telemetry the unit of identity is the resource: the service name, its version, the environment, the host, the container, the Kubernetes pod, the cloud account and region. OpenTelemetry standardizes these attribute names, and a trace adds a trace ID and span IDs that follow a request across services.

Vendors bill on entities too, and they don't agree on which entity counts:

Billing entityWho uses itRough list price (Oct 2026)
HostDatadog, Splunk Observability$15-75 per host a month, by product tier
Memory per host-hourDynatrace Full-StackAbout $58 a month for an 8 GiB host
UserNew Relic, Grafana, incident tools$8-349 per user a month, by role
EventHoneycombFrom $150 a month for 50 million events
Active metric seriesGrafana CloudAbout $6.50 per 1,000 series
Custom metricDatadog$5 per 100 a month
GigabyteEveryone, for logsCents to dollars, depending on indexing

Containers, serverless functions and autoscaling make "how many hosts" a moving number, so vendors bill on percentiles of hourly counts or "host-equivalents". The identity most often missing is ownership. Which team owns checkout-api? That lives in a service catalog if anywhere, and an alert without an owner pages whoever happens to be on call.

"Agent" has a double meaning, often on one product page: the software on each host that collects telemetry, and an AI agent. I'd write "host agent" and "AI agent" every time. Telemetry also carries other people's identities, such as user IDs, IP addresses and emails in URLs, which is where the privacy rules come in.

More on Entity & identity →

02

State & lifecycle

What states exist, and what moves an entity between them?

A piece of telemetry moves through states: emitted, collected, processed, ingested, indexed or stored, tiered down, deleted. Each change of state is a pricing event, and regulators have started naming the states: PCI's card-data rules require logs to be "immediately available" for three months, and the US federal memo M-26-14 distinguishes "actively searchable" from "retrievable".

TierWhat it meansRough cost
Hot, indexedSearchable in seconds, used by alerts and dashboardsDatadog: $1.06-2.50 per million log events, for 3-30 days
Warm or flexSearchable, slower, priced separatelyDatadog Flex: $0.05 per million events stored, plus compute
ArchiveObject storage the customer often ownsAbout $0.02-0.03 per GB a month
RehydratedArchive pulled back into the index for an investigation or auditA fee per GB, plus minutes to hours
DeletedRetention expired, or an erasure requestFree, but hard in append-only stores

Rehydration is fading. Since 2025, Datadog's Archive Search, Elastic's frozen tier and lakehouse engines let you query old data where it sits.

An alert goes from inactive to pending to firing, then acknowledged and resolved. An incident is detected, triaged, mitigated, resolved, reviewed and closed. Standards have states too: OpenTelemetry's traces, metrics and logs are stable, profiles reached alpha in March 2026, and the conventions for AI and LLM telemetry are still marked "Development".

The contract has states as well: free tier, annual commitment, monthly usage, overage, true-up, renewal. Seat reductions at Datadog take effect only at renewal, and overage history tends to become next year's commitment, so commitments ratchet up.

More on State & lifecycle →

03

System of record & ledger

Who owns the truth, and how do systems reconcile?

There's no single system of record.

RecordSystem of record
What happened in production (what was indexed)The observability backend, for as long as it keeps it
The full log historyOften an object-storage bucket the customer owns
The incident timelineThe incident tool
Who owns each serviceA service catalog, if one exists
What you'll be billedThe vendor's usage meter
Security and audit evidenceThe SIEM or the log archive

The billing ledger is the one I'd watch. The vendor's meter decides the bill, and customers rarely have an independent count. Datadog bills custom metrics on the monthly average of hourly counts, and hosts on the 99th percentile of hourly counts, so the top 1% of hours is forgiven and a week-long spike is billed. A pipeline gives the customer its own meter of what was sent, which helps in a billing dispute.

For security and compliance teams, logs are the record: PCI requires a year of audit logs, SOC 2 auditors want evidence, and a breach investigation depends on what was kept. For the reliability side the same logs are disposable signal. The two retention philosophies collide whenever one platform serves both teams.

More on System of record & ledger →

04

Rules & policy

What logic decides outcomes, and who can change it?

Most rules in observability are ones the customer writes: what to sample, redact and index, for how long, when to alert and who gets paged. They live in collector and alerting configuration, and they decide both the bill and what an investigator will find.

SLOs matter most for signal. Google's SRE Workbook recommends paging on burn rate over two windows, so a short spike pages nobody and a real outage pages fast:

Page

Long window
1 hour
Short window
5 minutes
Burn rate
14.4x
Budget spent when it fires
2%

Page

Long window
6 hours
Short window
30 minutes
Burn rate
6x
Budget spent when it fires
5%

Ticket

Long window
3 days
Short window
6 hours
Burn rate
1x
Budget spent when it fires
10%

That's for a 99.9% target over 30 days. At a burn rate of 1,000, the whole month's budget is gone in about 43 minutes. The math breaks for low-traffic services: at ten requests an hour, one failure spends about 14% of the monthly budget, so teams add synthetic traffic, group small services or lower the target. An error-budget policy closes the loop: while the SLO is met, releases continue, and when the budget is spent, they pause.

Google's on-call guidance sets a bar few teams meet: every page actionable, at most two incidents in a 12-hour shift, and on-call capped at a quarter of an engineer's time. Vendors add their own rules (cardinality limits, rate limits, index quotas), which are product decisions about how hard to stop a customer's mistake.

More on Rules & policy →

05

Effective dating

Which version of the rule applied at that moment?

Everything in observability is measured over a window. Alerts evaluate over minutes, SLOs over a rolling 28 or 30 days, burn rates over 5 minutes to 3 days, and billing over hours, months and contract years. Retention clocks usually start when data arrives, so late mobile or batch data can land in the wrong window.

The meaning of a field has a date. OpenTelemetry's semantic conventions are versioned, and renames break things quietly: database spans renamed db.system to db.system.name in version 1.33.0 in 2025, and dashboards keyed on the old name stopped matching.

Dates I'd keep on a calendar as of October 2026:

ChangeEffectiveStatus (Oct 2026)
PCI DSS 4.0's future-dated requirements, including automated log reviewMarch 31, 2025Live
Atlassian Opsgenie end of saleJune 4, 2025Done; support ends April 5, 2027
California's cybersecurity audit rules, covering audit-log managementJanuary 1, 2026Live; first certifications from April 1, 2028 for the largest businesses
OpenTelemetry graduates in the CNCFMay 21, 2026Done
OMB M-26-14 replaces M-21-31 for US federal loggingMay 22, 2026Live
Datadog's AI SRE moves to AI creditsAround June 2026Live
FedRAMP 20x rules launchedJune 25, 2026Mandatory January 1, 2027
OpenTelemetry Governance Committee election resultsOctober 30, 2026Upcoming
EU Data Act bans switching and egress feesJanuary 12, 2027Upcoming
HIPAA Security Rule updateProposed January 2025Final rule targeted around July 2027; uncertain

Take the dispatch company's commitment renewing in March. Overage from a cardinality spike in January sits in the usage history the vendor brings to the renewal, so a mistake fixed in a week can still raise next year's commitment.

More on Effective dating →

06

Interfaces & standards

What format and protocol do counterparties speak?

The collection side is now mostly standard:

  • OTLP, OpenTelemetry's protocol for sending traces, metrics and logs, is stable and accepted by every major vendor.
  • W3C Trace Context is the traceparent header that carries a trace ID between services.
  • Semantic conventions name attributes the same way everywhere (HTTP stable since 2023, databases since 2025, AI still in development).
  • Prometheus exposition format and PromQL are the de facto standard for metrics and the closest thing to a common query language; many backends accept PromQL.

The storage and query side has no standard. Logs and traces are queried in each vendor's language: LogQL and TraceQL at Grafana, SPL at Splunk, KQL at Microsoft, ES|QL at Elastic, DQL at Dynatrace, Datadog's own syntax, SQL on ClickHouse. Dashboards and alert rules are written in these languages, which is why they don't port.

Every major vendor now ships its own OpenTelemetry distribution, and the distributions recreate some lock-in. Splunk gives official support only for its own and "best-effort" support upstream. Datadog's wasn't FedRAMP or FIPS compliant when the research was done, which steers government buyers back to the proprietary agent. Tail sampling is "often tied to a specific vendor", in OpenTelemetry's own words. Datadog's FY2025 annual report doesn't mention OpenTelemetry at all; it sells "a single agent" with over 1,000 integrations. Its price list charges $0.50 per GB for spans ingested over OTLP against $0.10 for its own APM ingestion, and application metrics sent through OpenTelemetry generally count as custom metrics.

Vendors are also giving their collection technology away, which is what you do once collection stops being where you make money: Elastic donated its eBPF profiler and Grafana its Beyla. For AI, two newer interfaces matter. MCP lets AI agents query observability tools, and Datadog said MCP tool calls grew more than 22-fold from late 2025 to mid-2026 (company figure). GitHub issues and pull requests are the hand-off between observability AI and coding agents.

More on Interfaces & standards →

07

Networks & counterparties

Who sits between us and the outcome, and what do they want?

PartyWhat they controlWhat stops if they fail
App teamsWhat gets instrumented, log levels, labelsTelemetry quality; the bill, through cardinality and verbosity
Platform or observability teamCollectors, pipelines, budgets, vendor choiceCollection, sometimes for everyone
Observability vendorStorage, query, alerting, pricing unitsVisibility during incidents
Pipeline vendorRouting, sampling, redactionData to every destination at once
Cloud providerCompute, network charges, default tools, marketplace billingEverything running on it
Incident toolPaging, escalation, incident recordsWaking the right person
Security team and SIEMDetection rules on the same logsSecurity monitoring
AI agent vendorsRead access to telemetry, sometimes write access to codeAutomated investigation

Concentration runs both ways. Datadog's largest customer, widely reported to be OpenAI, renewed at nine figures in 2026 but is cutting usage from July, and Datadog built that into its guidance. Coinbase's roughly $65 million a year Datadog bill was restructured in 2023 after Coinbase built a parallel open-source stack. The biggest customers can credibly leave, and the threat works best at renewal.

The pipeline is the new switching point. Whoever owns it decides which backend gets which data, which is why security and observability vendors have both bought pipelines since 2025.

I'd write down early which single failure blinds us during an incident. For many teams it's the collector tier: in OpenTelemetry's 2026 survey of collector users, 23% didn't monitor their collectors at all.

More on Networks & counterparties →

08

Regulatory layering

Jurisdiction × activity × entity type: is it a license or a certification?

Observability has no licensing regime. Nobody needs permission to collect or store telemetry, and no regulator supervises observability vendors as such. Rules reach this industry through four other doors.

Data inside telemetry

What it covers
GDPR and CCPA for personal data, HIPAA for health data, PCI DSS for card data
Binds
The customer as controller; the vendor as processor by contract
Enforced by
Data protection authorities, the FTC and states, card networks through acquirers

The buyer

What it covers
FedRAMP for US federal agencies, OMB logging memos, SEC disclosure for listed companies
Binds
Vendors selling to those buyers; the buyers themselves
Enforced by
Agencies, auditors, the SEC

The market

What it covers
EU Data Act switching rules, competition review of deals
Binds
Vendors
Enforced by
The EU, competition authorities

Attestation

What it covers
SOC 2, ISO 27001
Binds
Vendors, as a sales requirement
Enforced by
Customers' procurement teams

GDPR and CCPA don't mention logs; they apply because logs carry IP addresses, user IDs, emails and tokens. SOC 2 is a contract artifact with no legal force, and in practice it's the report every vendor is asked for first.

The test I'd use on any requirement: is it about the data, the buyer or the vendor? Data rules follow the data into every tool, buyer rules only matter if you sell to that buyer, and vendor rules mostly live in contracts.

More on Regulatory layering →

09

Exceptions & reversals

What goes wrong, and how is it undone?

ExceptionThe way backClock
Data needed from the archiveRehydrate into the index, or query in placeMinutes to hours, plus fees
Sampled or dropped dataNone; it's goneNever
Cardinality or log spike on the billAsk for a one-time credit; vendors sometimes grant oneA negotiation
Bad deployRoll back, the most common mitigationMinutes
Alert flapping or false pagesTune the rule or move to SLO burn ratesWeeks
Postmortem actions not doneTrack them; Google says an unreviewed postmortem "might as well never have existed"Next incident
Legal hold or security investigationFreeze deletionUntil released
Forced tool migrationMove off Opsgenie (support ends April 5, 2027) or Grafana OnCall open source (archived March 24, 2026)Months

Rules reverse too. The US federal logging memo M-21-31 was rescinded on May 22, 2026, a federal court struck down part of HHS's guidance on tracking technologies in 2024, and the HIPAA Security Rule update has stalled. Product reversals follow public pressure: after the Storm-0558 intrusion in 2023, which agencies could only detect with logs from a premium tier, Microsoft gave standard customers more than 30 log types at no extra cost.

More on Exceptions & reversals →

10

Liability allocation

When it fails, who pays?

The customer carries most of the risk in observability, and the contracts say so.

FailureWho pays firstHow
Bill shock from a spikeThe customerUsage is billable by contract; credits are a goodwill gesture
Vendor outageThe customer, in lost visibility; the vendor, in credits and lost usage revenueSLA service credits, the "sole and exclusive remedy"
Missed alert leads to an outageThe customer and its usersNothing in the vendor's contract covers it
Personal data loggedThe customer as controllerGDPR fines land on the controller, not the logging tool
Privileged host agent crashes machinesCustomers and insurers, then the vendor if negligence is provenLiability caps, tested in court
AI agent's wrong suggestionThe customer, through the human who approved itVendors' terms don't accept liability for remediations

The SLAs are short. Grafana Cloud promises 99.5% of requests succeed each month and credits 10% to 100% of the affected service's monthly fee as availability drops; it also covers alert evaluation, so a missed alert is contractually defined there. New Relic targets 99.8%, with termination as the remedy if it misses badly two months running, and I found no public Datadog SLA to compare.

Datadog's own outage on March 8, 2023 shows the vendor side. A security update auto-applied across its nodes knocked out networking, and the company said it lost about $5 million of revenue, because under usage pricing it couldn't bill for data it couldn't ingest. Customers were blind for most of a day and got credits at best.

CrowdStrike's July 2024 outage is the extreme case for any privileged host agent, which observability agents are becoming as they take on security work. One estimate put direct losses to Fortune 500 companies at $5.4 billion, and CrowdStrike says its liability cap limits its exposure to "single-digit millions". A court let Delta's negligence claims proceed in May 2025.

More on Liability allocation →

What's different here

How the money moves

The customer pays almost everyone, mostly by volume. The app teams who decide what to log rarely see the bill; the platform team owns it.

FlowWho pays whomRough amount (Oct 2026, list)
Infrastructure and APMCustomer → platform$15-75 per host a month, or by memory or user
Log ingestCustomer → platform$0.10-0.60 per GB
Log indexingCustomer → platformAbout $1-2.50 per million events at Datadog, for 3-30 days
MetricsCustomer → platform$5 per 100 custom metrics (Datadog); about $6.50 per 1,000 series (Grafana); 5-30 cents per metric (CloudWatch)
Users and on-call seatsCustomer → platform or incident tool$15-50 per responder; up to about $350 for a full New Relic user
PipelineCustomer → pipeline vendorRoughly $0.10-0.20 per GB processed
Archive storage and networkCustomer → cloud providerAbout $0.02-0.03 per GB a month to store, plus charges for moving data between zones
Security on the same dataCustomer → platformDatadog: $5 per million events analyzed by its SIEM; $10-25 per host for cloud security
AI investigationCustomer → platform or cloudA few dollars per investigation; AWS DevOps Agent about $30 per agent-hour

Enterprise contracts are discounted from list, often a lot, and nobody publishes by how much. Overage is the opposite: above the commitment, Datadog's on-demand rates run 20-50% higher than its annual rates.

Gartner's research, as summarized by others, puts its median client's spend with a single observability vendor above $800,000 a year. At the top end, OpenAI was reported in 2025 to spend about $170 million a year on Datadog.

The vendors' growth model is the customer's bill seen from the other side. Datadog, Dynatrace and Elastic report net retention between about 110% and the low 120s, meaning existing customers spend that much more each year, and in Grafana's 2026 survey most respondents expecting higher spend blamed broader adoption more than price rises.

A year of observability for 500 hosts and 2 TB of logs a day

Back to the dispatch company: 500 hosts with APM on all of them, about 2 TB of logs a day (roughly 60 billion log events a month at an average of 1 KB each), half a million to 1.5 million metric series, and 10-20 TB of sampled traces a month. Searchable logs are kept 30 days, metrics 13 months, and a log archive for a year. Every figure is a list-price range for a year, before discounts.

OptionPer year (list)What drives it
Datadog, logs indexed with discipline (10-20% indexed, the rest in cheap storage)About $0.6-1.05 millionHosts, custom metrics, indexed share
Datadog, everything indexed for 15-30 daysAbout $1.6-2.6 millionLog indexing alone is $1.2-1.8 million
New RelicAbout $0.5-0.7 millionGigabytes and full users; no per-host fee
DynatraceAbout $0.55-0.9 millionMemory per host-hour; larger nodes cost more
Grafana Cloud (managed open source)About $0.45-0.8 millionGigabytes and active series
Self-hosted open source on a cloudAbout $0.55-2 millionTwo to five engineers at $200,000-300,000 each
AWS CloudWatchAbout $0.6-1 millionGigabytes and per-metric fees; counts toward the AWS commitment
Add a pipelinePlus $70,000-150,000Gigabytes processed

The striking gap is between the two Datadog rows: same telemetry, same vendor, about three times the bill, and the only difference is how much gets indexed. Under ingest-only pricing (New Relic, Grafana) savings from dropping data scale with bytes, so a pipeline matters less there.

As I read the numbers, self-hosting doesn't save money at this size once people are counted. It wins at Airbnb's scale, where moving metrics to OpenTelemetry and open-source storage at more than 100 million samples a second reportedly cut cost by about an order of magnitude. CloudWatch looks expensive on paper, and Gartner's clients complain about its cost, but teams pick it anyway because it draws down a commitment they've already signed.

Why the bill grew 47% when traffic grew 10%

Now the next quarter. Traffic grows 10%, and the disciplined Datadog setup starts at about $66,000 a month.

Hosts +10% (autoscaling follows traffic)

Monthly cost before
About $27,000
After
About $29,700
Why
The linear part

A team adds a customer_id or pod tag to three busy metrics

Monthly cost before
About $10,000 (250,000 custom metrics)
After
About $22,500 (500,000)
Why
Series multiply by every tag value

New services and retries add 25% more log events, and the indexed share creeps from 15% to 20%

Monthly cost before
About $15,300 (9 billion indexed)
After
About $25,500 (15 billion)
Why
Verbosity and indexing defaults

Usage passes the commitment, so the excess is billed at on-demand rates

Monthly cost before
After
About $5,100 more
Why
Overage runs 50% above the committed rate for indexed logs

Total

Monthly cost before
About $66,000
After
About $97,000
Why
About +47% on +10% traffic

None of those drivers shows up on a traffic dashboard. A smaller example in the research, 200 hosts and 1.5 TB of logs a day, comes out at +39% on the same mechanics, so I'd tell a CFO "40-50%" and explain the four causes.

The fix is mostly in the pipeline and the tag allow-list: drop debug logs before ingest, index less, keep the customer tag ingested but unindexed, or move per-customer detail into traces. In the smaller example those moves bring the bill back to about 6% above baseline. The pipeline itself cost about 15% of the bill there. At index-heavy prices it pays for itself many times over if it cuts what gets indexed; cutting ingest alone barely covers its cost.

What the pipelines are worth

Cost control is a funded category, and it's being bought: Palo Alto, Dynatrace, CrowdStrike and SentinelOne all bought pipelines in 2025-2026 (see The main players). The platforms ship their own cost tools too, such as Grafana's Adaptive Metrics and Adaptive Logs, which Grafana says cut series by up to 80% and logs by up to 50%. Cribl's secondary-market value of about $3 billion in September 2026 sits below its 2024 round. The buyers seem to treat pipelines as a feature of a SIEM or a platform, and I'm not sure cost control survives as a standalone category.

Who holds the power

Power in observability follows whoever controls the meter, the history and the pipeline.

  • Platform vendors hold the meter and the unit definitions, years of history, dashboards and alert rules written in their query languages, multi-product bundles (58% of Datadog's customers used four or more products by mid-2026) and an agent on every host. Each makes a partial exit harder.
  • Customers gain power through OpenTelemetry, a pipeline they control, an archive in their own bucket, PromQL-compatible backends and the renewal date. In Grafana's survey, 37% of OpenTelemetry adopters cite freedom to switch vendors. The power shows up mostly at renewal and mostly for large buyers who can afford to run two stacks.
  • Cloud providers hold the default, the invoice and the commitment. Their tools are one click away, their charges count toward commitments customers already signed, and AWS gives DevOps Agent credits worth 30-100% of a customer's support spend, which specialists can't match.
  • Security vendors hold the CISO's budget and the detection content, and they're buying pipelines to control what reaches any SIEM.
  • OpenTelemetry's maintainers, many of them vendor employees, decide what attributes are called. A rename costs everyone downstream.
  • Coding-agent vendors are a new channel. Whoever feeds production context into Claude, Copilot, Cursor or AWS's Kiro gets embedded in how code is written.
  • Regulators hold little direct power. They set retention floors for some buyers and data rules for what's inside the logs.

How the rules work

There's no observability regulator and no licence, so most of this section is about data rules and buyer rules that happen to land on telemetry. I'd still take them seriously, because the fines come from what's in the logs.

Personal data in logs. Telemetry carries IP addresses, user IDs, emails in URLs, auth tokens in headers and now LLM prompts. The largest GDPR fine I found about data sitting in internal systems is Meta's €91 million from Ireland's regulator in September 2024, for passwords stored in plaintext, partly for failing to notify and document the breach. France's CNIL recommends keeping access logs six months to a year, minimizing personal data in them, and storing them separately and append-only. Deleting one person's data from append-only stores is hard, so the common practice is to redact in the pipeline and keep retention short. The customer is the controller and carries the liability; the vendor is a processor.

Diagnostic files are the leak people forget. In 2023 an attacker reached support files for 134 Okta customers; some were browser recordings (HAR files) holding session tokens, and five customers' sessions were hijacked. Those files never went through a pipeline. Session replay in RUM tools has its own exposure: since a court struck down part of HHS's tracking guidance in 2024, claims over replay and tracking pixels have moved to state wiretap lawsuits against website operators.

Retention floors.

PCI DSS v4.0.1, 10.5.1

Who it covers
Systems in the card-data environment
Retention
12 months, the last 3 immediately available
Status (Oct 2026)
In force; automated log review required since March 31, 2025

OMB M-26-14

Who it covers
US federal civilian agencies
Retention
6 months actively searchable, 12 months retrievable
Status (Oct 2026)
In force since May 22, 2026

OMB M-21-31

Who it covers
US federal agencies
Retention
12 months hot plus 18 cold
Status (Oct 2026)
Rescinded May 22, 2026

CNIL recommendation

Who it covers
France, under GDPR
Retention
6-12 months for access logs
Status (Oct 2026)
Guidance

HIPAA audit controls

Who it covers
Covered entities and their vendors
Retention
No period for logs; policies kept 6 years
Status (Oct 2026)
Update proposed, stalled

California cybersecurity audits

Who it covers
Large businesses under CCPA
Retention
Audits cover audit-log management
Status (Oct 2026)
Effective January 1, 2026; first certifications 2028

The federal change is the clearest example of rules following the bill. M-26-14 cut the floor from 30 months to 12, and the memo itself says keeping "vast quantities of logging data without clear utility" was neither operationally feasible nor cost-effective. Its top maturity level grades signal: logs should generate actionable alerts covering at least 95% of baseline requirements. PCI's 12-and-3 maps neatly onto vendors' hot and archive tiers.

Buyer rules. FedRAMP decides who can sell to US federal agencies. Datadog for Government and Grafana's federal cloud say they're authorized at High, Dynatrace holds Moderate and said in July 2026 it would pursue High, and FedRAMP's new 20x process becomes mandatory on January 1, 2027. Listed US companies must file an 8-K within four business days of deciding a cyber incident is material, under SEC Item 1.05; trade groups want it rescinded, but I found no proposal as of October 2026. Telemetry is the evidence for scoping, and "we can't determine the scope" is an answer regulators read badly.

Residency and switching. Regions are fixed at setup. Datadog's sites are separate ("you cannot share data across sites"), and New Relic says the customer is "solely responsible" for choosing its region. Moving is a re-implementation. From January 12, 2027, the EU Data Act bars cloud and SaaS providers, observability vendors included, from charging EU customers switching or data-egress fees when they move.

SLAs and contracts. Most of the customer's protection sits here, and it's thin; see Liability allocation. EU financial firms have their own operational resilience rules (DORA), which I haven't covered.

What mistakes cost

  • Bill shock: a tag or a debug flag can add tens of thousands of dollars a month within days, and overage history raises the next commitment.
  • Ignored alerts: users find the outage first. In Splunk's UK cut of its 2025 survey, 54% said false alerts hurt morale.
  • Cost-cutting blind spots: dropped logs and sampled traces lengthen the next investigation, and rehydrating the archive mid-incident costs time you don't have.
  • Secrets in logs: Meta's €91 million fine was about passwords in internal systems, and Okta's 2023 breach ran through troubleshooting files.
  • Logs you didn't keep: after Storm-0558, the US Cyber Safety Review Board found Microsoft had "no evidence or logs" for parts of its explanation.
  • Vendor and agent failures: Datadog's 2023 outage and CrowdStrike's 2024 update, both in Liability allocation, left customers carrying most of the cost.
  • The wrong region: moving an account to another region is a rebuild, with dashboards and alerts recreated by hand.

Observability, security and AI: what's real

My view as of October 2026: the convergence is real at two layers, the pipeline and the agent on the host, and weak at the revenue line, where security buyers still mostly buy from security vendors. AI investigation is improving but solves a minority of incidents on public benchmarks, and the AI layer is separating from the data store. If I were at an observability vendor, I'd treat the pipeline and the context (who owns what, what changed, how services connect) as what's worth defending, more than the raw data.

What's real in the security convergence

  • Buying the data path: Palo Alto closed Chronosphere on January 29, 2026 and calls it "our next-generation observability platform"; its annual report now names Datadog, Dynatrace and Elastic as competitors. CrowdStrike and SentinelOne bought telemetry pipelines in 2025. Going the other way, Cribl bought detection-engineering and AI security-operations assets in mid-2026.
  • Same meters: Datadog prices its SIEM at $5 per million events analyzed and cloud security at $10-25 per host, on the same agent and bill. On a company with the dispatch company's host count and about half its log volume, adding SIEM, cloud security and a sensitive-data scanner would run roughly 20-35% on top of the log bill. One of Datadog's mid-2026 wins was a bank adopting its SIEM for how it handles personal data.
  • The host agent: Wiz, built agentless, now ships a runtime sensor on the workload, while Datadog sells workload security from its observability agent. Google closed Wiz for $32 billion on March 11, 2026, so a cloud provider now owns a cloud-security leader.

What's not there yet

  • Scale: Datadog's security products passed $100 million of ARR in 2025. By their 2026 figures, Palo Alto's XSIAM and CrowdStrike's Next-Gen SIEM are each around $700 million, roughly seven times larger even allowing for the year's gap, and Datadog's security business is a few percent of its revenue by my estimate.
  • Growth from owning both: Cisco owns Splunk on both sides, and in its fiscal 2026 Security grew 2% and Observability 4%.
  • Different buyers: reliability teams measure time to recover, SLOs and cost per gigabyte; security teams measure detection coverage, how long an attacker goes unnoticed, and audit evidence. Security vendors stay ahead on detection rules and threat intelligence, and observability vendors win on consolidation and data handling.

What's real in AI SRE

AI SRE (site reliability engineering) products promise to turn an alert into a diagnosis, and sometimes a fix.

ProductWhat it doesHow it's priced (Oct 2026)
Datadog Bits AI SREInvestigates alerts; marketed as "fully autonomous" detection, investigation and remediation since mid-2026AI credits, about $1.30 each; one listing estimates about 6.5 credits an investigation
Dynatrace BlueboxGives coding agents production context; opens GitHub issues with the trace, the offending lines and the blast radiusFree with 5 GiB of ingest; Pro $75 a month; Enterprise custom
AWS DevOps AgentInvestigates incidents across AWS, Azure and on-premises apps$0.0083 per agent-second, about $4 for an eight-minute investigation, offset by support credits
Azure SRE AgentInvestigates incidents on AzureAgent units plus tokens
PagerDuty SRE agentTriage inside the incident tool; a "fully autonomous responder" in early access from late 2026Not public
Resolve AI, Traversal, ClericStartups that sit on top of existing toolsNot public

The benchmarks are sobering. On OpenRCA, a public root-cause benchmark presented in 2025, the best model solved about 11% of 335 failure cases, and ITBench put AI agents at 11-14% of its SRE scenarios. A 2026 follow-up, OpenRCA 2.0, found about 21% exact root causes across 11 frontier models. Scores are rising from one paper to the next, but I'd treat any claim of autonomous root-cause analysis as running ahead of the evidence.

Dynatrace shows the gap in its own messaging. Its CTO has talked about taking humans out of the loop while also saying nobody has moved from AI-assisted to autonomous operations yet, and Bluebox's documentation says it "never pushes a fix on its own": it opens an issue, a coding agent proposes the change, and a person merges it. That's a policy statement more than a technical guarantee, and the liability sits with whoever approved the change.

Practitioners want to see the reasoning: in Grafana's 2026 survey, 95% said AI must show how it reached its answer, and the top barrier was having to feed it context by hand.

Where the AI layer is heading

  • Away from the store: Datadog separated its AI security analyst from its SIEM "so customers can use it with other SIEMs", AWS DevOps Agent investigates apps outside AWS, and the startups store little telemetry. If the AI works on anyone's data, the store is less of a lock.
  • Into coding agents: Bluebox, Datadog's Bits Code and AWS's Kiro feed production data into the tools that write code. The bottleneck moves to reviewing changes before production, the theme of Becoming AI native.
  • AI workloads as telemetry: Dynatrace says more than 1,000 customers monitor AI workloads with it, and LLM prompts make traces bigger. OpenTelemetry's AI conventions are still in development, so vendors fill the gap with their own SDKs; agent tracing and evaluation are in AI agents: orchestration platforms.
  • Metered pricing: credits, agent-seconds and tokens are priced like telemetry, and Datadog's move from per-investigation bundles to credits in mid-2026 cut the list price per investigation by about three-quarters by my arithmetic.

Questions to ask about an AI investigation agent

  1. On your last hundred real incidents, how often was the agent's initial answer the root cause, and who checked?
  2. What data does it read, through which interface, and does it copy telemetry into tickets or chat outside our retention policy?
  3. Can it change anything in production, and what's the approval step and the audit trail when it does?
  4. Does it work on data stored with another vendor, or only yours?
  5. How is it priced per investigation at our volume, including the query charges it triggers?

What usually goes wrong

SymptomLikely causeFirst thing to check
Bill up 40-50% on 10% more trafficA new high-cardinality tag, debug logs left on, indexed share creeping up, overage at on-demand ratesUsage by product, top custom metrics by tag, indexed share
Self-hosted Prometheus runs out of memoryCardinality explosion from a label like user ID or raw URL pathActive series by label
Traces arrive in piecesA proxy or queue drops the trace headerWhether traceparent survives each hop
Error traces missing during an incidentHead sampling dropped themThe sampling policy; tail sampling for errors
Pager fires dozens of times a day and gets ignoredThreshold alerts on noisy metricsAlerts per real incident; SLO burn-rate alerts
Customers report the outage firstNo alert, a misconfigured one, or an ignored oneThe detection section of the postmortem
Data gaps with no errors anywhereCollector buffers full, or collectors not monitoredThe collectors' own metrics
Double bill or duplicate spans during migrationVendor agent and OpenTelemetry both instrumenting the same appWhich one instruments each service
Emails or tokens found in logsNo redaction before storageRedaction rules in the collector or pipeline
Auditor asks for last year's card-system logsOnly the hot index was keptRetention tiers against PCI's 12 months

Words that mean something else here

TermWhat you'd assumeWhat it means here
AgentAn AI agentUsually the collector software on a host; now also an AI agent, often on the same product page
EventSomething that happenedHoneycomb: a wide record (a span is one). Datadog: a log line or a change notification. Security: a SIEM record
CardinalityA row countThe number of unique values of a label, and so the number of metric series
Custom metricA metric you designedAt Datadog, any metric not from one of its integrations, counted per unique name and tag combination
HostA server you ownA billing unit: a VM, a Kubernetes node, or a "host-equivalent" for containers
Ingest vs indexThe same thingIngest is receiving data, priced per GB; indexing makes it searchable for a period, priced per event
SamplingTaking a sample to inspectDropping data on purpose, by a rule
PipelineA CI/CD pipelineA telemetry router that filters, samples and routes
MonitorWatching somethingDatadog's word for an alert rule. In compliance, transaction monitoring (Identity and trust)
SLO, SLA, SLIInterchangeableSLI is the measurement, SLO the internal target, SLA the contract with credits
Error budgetMoneyThe failure you're allowed (1 − the SLO) over a window
Burn rateStartup cash burnHow fast the error budget is being spent; 1x is exactly on budget
IncidentOne thingA reliability incident, a security incident, or an SEC "cybersecurity incident" with a materiality test
AutonomousActs on its ownUsually "investigates without being asked"; few products merge and deploy without a person
CreditA refundAn SLA service credit, or a prepaid unit of AI usage

What surprised me

Placeholders in your voice, drafted from the research and the earlier guides. Rewrite each with your own moment.

"Observability is a tool you buy." Coming from payments, I pictured a product with a price. It turned out to be a set of choices about what to keep, and the same telemetry on the same vendor could cost three times as much depending on how much we indexed.

"Open standards end lock-in." OpenTelemetry made the collection portable. The dashboards, alert rules, query language and two years of history stayed exactly where they were.

"More data, fewer outages." The incidents that hurt were the ones where an alert fired and nobody looked, because it had fired falsely all month.

"Security and observability are merging." The deals say so. The revenue says the security vendors' SIEMs are still several times bigger than the observability vendors' security lines.

"This is a regulated industry like the others." There's no licence and no supervisor. The rules that bit were about what we'd written into the logs.

Sources

Undated entries were read on October 8, 2026; "search result" means seen only as a search snippet or summary. Company figures are self-reported unless they come from a filing or a regulator.

Standards, open source and research

Rules and regulators

Company filings and results

Private companies and deals

Pricing, product and legal pages

Surveys, press and analysis

Field Guides are learning notes, not legal or compliance advice. Rules and fees change; check the cited primary sources before you act on anything here.