Cloud Native

Atlassian’s Faster Incident Pipeline Exposes the Coverage Trap

Atlassian has published an unusually candid account of rebuilding the automated incident-detection platform behind more than ten cloud products. The headline result is technically impressive: a pipeline based on Apache Kafka, Apache Flink on Kubernetes, and OpenTelemetry cut event-to-metric latency from more than 40 seconds to under 10. The more important result is less flattering and more useful. Even after the rebuild, the system detected only 30.4% of all major incidents measured over a nine-month period.

That gap does not invalidate the architecture. It reveals the distinction that platform teams often blur when evaluating observability: detection recall measures performance inside the territory a system can see, while coverage measures how much territory has been instrumented at all. Atlassian’s pipeline became faster, cheaper, and more reliable within its defined scope. Yet the mix of incidents expanded beyond that scope, and hard failures sometimes removed the very telemetry on which detection depended.

The lesson for cloud-native operators is that a sophisticated streaming architecture can improve the speed and quality of a signal, but it cannot compensate for a missing signal. Teams planning a similar system should treat telemetry coverage, signal independence, and detector self-monitoring as first-class reliability objectives—not as cleanup work after the data path is fast.

A rebuild driven by latency, isolation, and cost

Atlassian’s first-generation system aggregated client-side operational events through a shared cloud queue, an in-memory cache, and roughly 90 virtual machines. It provided valuable domain knowledge, but it had three structural problems. Event-to-metric latency exceeded 40 seconds even under favorable conditions. A shared queue exposed the detection system to noisy-neighbor failures. Costs grew as products and user experiences were onboarded, rising from roughly $120,000 to $230,000 annually.

The replacement starts by filtering upstream. A subscription filter generated from configuration selects relevant events from the company’s Kafka-based event bus and places them in a dedicated topic with seven days of retention. A single Flink 1.20 job, managed by the Flink Kubernetes Operator, parses and transforms events, enriches them with tenant context, produces telemetry, deduplicates users, and builds minute-level impact aggregates.

Those aggregates feed several destinations. OpenTelemetry carries metrics to a Prometheus-compatible time-series database. Idempotent records land in a key-value store for impact analysis, while checkpoint-committed Parquet files support longer-term reporting. A separate decision engine converts detector alerts into anomalies, estimates impact, applies severity rules, suppresses short-lived blips, opens incidents, and pages responders.

This design reflects a sound cloud-native principle: assign different delivery guarantees according to the value of the data. Impact records use deterministic keys so Kafka replays overwrite rather than double-count. Parquet output is committed on checkpoints. Raw metrics tolerate at-least-once behavior because detectors use pre-aggregated state instead. The system avoids paying an exactly-once tax everywhere merely because correctness matters somewhere.

The architecture delivered real operational gains

The rebuild was not an exercise in adopting fashionable projects. Several implementation choices directly addressed known failure modes. Filtering at the bus reduced the volume reaching the job. Async tenant enrichment uses caching, a concurrency bulkhead, and a circuit breaker so a slow dependency cannot stall the entire stream. HyperLogLog sketches make distinct-user counts mergeable across minutes, tenants, regions, and products without moving individual user identifiers through the analytical path.

State is stored in RocksDB, checkpoints go to object storage every 30 seconds, and upgrades restore from savepoints. Atlassian also found that stable Flink operator identifiers behave like an API: changing topology carelessly can discard state. Autoscaling required similarly deliberate treatment. Early settings caused about eight job restarts per day; tuning target utilization, scaling boundaries, parallelism limits, and the scale-down interval reduced that figure to zero.

The system also separates operators according to their resource profile. Parsing and filtering were merged where doing so eliminated redundant deserialization, while CPU-heavy transformations were split from I/O-bound work so the two stages could scale independently. That is a more durable heuristic than maximizing or minimizing the number of operators: merge stages that make the same pass over data, and split stages that have different scaling requirements.

Economically, pre-aggregation produced a dramatic result. The monthly cost of the impact dashboard fell from more than $20,000 to about $1,000 because queries stopped scanning logs and instead read prepared state. Operationally, median detection time for major incidents with available timing data ranged from one to eight minutes. In one case, automation opened an incident nine minutes before a human did.

Recall improved while overall coverage remained constrained

The uncomfortable numbers make this case study more valuable than a conventional migration story. In-scope recall rose from about 60% in March 2025 to 86% in June 2026 and 80% in July, before falling to 64% in August. On three core products, recall reached 100% in three separate months. Those improvements correlated with concrete changes including blip suppression, early-warning tickets, the Flink cutover, 4xx detection, and volume-drop detection.

But among 263 major incidents recorded from November 2025 through July 2026, only 117 affected instrumented experiences. The platform detected 80 of those, producing 68.4% recall within scope but just 30.4% coverage across all major incidents. The share of incidents that fell within instrumented scope declined as failures spread into newer products and causes unrelated to conventional service reliability.

This is the coverage trap. A team can optimize the detector against visible incidents and report a rising recall rate while a growing fraction of real failures remains invisible. A mature reliability program therefore needs at least two separate service-level measures: one for the percentage of eligible incidents detected, and another for the percentage of consequential incidents represented by eligible telemetry. Combining the two rewards local optimization and hides portfolio risk.

Precision also needs segmentation. Atlassian reports about 90% precision on covered scope for core products, but lower figures when low-severity early warnings are included. Some rejected tickets were genuine signals below the incident threshold, while others reflected process problems such as duplicates or missing rejection reasons. One burst of false volume-drop incidents came from metrics-ingestion lag being interpreted as zero traffic. Adding a 120-second minimum delay traded detection speed for better precision.

That trade is instructive. “Fastest possible” is not the correct objective for every detector. The useful target is the shortest delay that preserves enough confidence for the action being automated. A dashboard can tolerate speculative signals; paging and incident creation impose organizational costs. Detection latency should therefore be budgeted by response consequence rather than applied as a universal threshold.

Silence is the failure mode streaming systems cannot infer away

Atlassian’s client-side events describe whether a user action begins, succeeds, fails, or is abandoned. That makes them close to the customer experience during partial degradation. It also creates a fundamental blind spot: when a database shard is completely unavailable and a page never loads, the client emits no event. A detector watching failure rates can interpret the absence of failures as health.

The same dependency problem applies to the detection pipeline itself. A stall in the upstream event path can starve the detector, leaving the compute healthy while its evidence disappears. In a single-region design, a regional outage may preserve the active-active decision engine while taking away the feed it needs to decide anything.

Volume-drop detection helps but cannot be the sole answer. Expected traffic varies with geography, product usage, weekends, and holidays. Backend ingestion delays can resemble outages. Independent signals are essential: synthetic transactions, edge-level 5xx rates, telemetry freshness indicators, and checks that continuously inject known events through the complete incident workflow. Atlassian runs an end-to-end synthetic check every 15 minutes, testing alert ingestion, impact analysis, ticket creation, and closure.

OpenTelemetry supports the portability of this self-observation layer across the Java stream job and Go decision engine. Its metrics model also permits spatial and temporal reaggregation, which is useful for controlling cardinality and cost. But standardizing transport does not decide which signals should exist. The architecture still needs an explicit dependency map showing which detector can survive the loss of each monitored component.

What platform teams should copy—and what they should measure

Teams do not need Atlassian’s event volume to adopt the most transferable parts of the design. The priorities are architectural boundaries and measurement discipline:

  • Filter before expensive processing. Keep subscription policy as code and reduce the stream near its source.
  • Design replay-safe storage keys early. Recovery is far simpler when reprocessing produces deterministic writes.
  • Match guarantees to each sink. Use idempotence and checkpoint commits where duplication is dangerous; tolerate weaker semantics where brief distortion is acceptable.
  • Treat Flink state identity as production configuration. Stable operator identifiers, savepoints, checkpoint tests, and rescaling policy belong in deployment reviews.
  • Measure recall, precision, and coverage separately. Add reason codes for rejected incidents so tuning data reflects reality.
  • Test the watcher end to end. A green pod or successful HTTP request does not prove that an incident ticket can be created.
  • Add signals that remain present during absence. Pair client telemetry with synthetics, infrastructure indicators, freshness alarms, and independent regional paths.

The next phase of Atlassian’s work follows directly from its findings: move detection closer to the OpenTelemetry-fed store and Flink job, add active-active regional processing, introduce roll-ups and caching for impact queries, and establish telemetry-coverage objectives per product. Those changes are more consequential than shaving another few seconds from the hot path because they expand the system’s trustworthy operating envelope.

The larger cloud-native lesson

Kubernetes, Flink, Kafka, and OpenTelemetry gave Atlassian useful primitives for isolation, stateful stream processing, replay, standardized telemetry, and portable operation. The projects solved difficult infrastructure problems. They did not decide whether the chosen events represented the full failure surface, whether an absent event meant success, or whether the organization was measuring detector quality instead of customer coverage.

That boundary is the central takeaway. Cloud-native architecture can make detection faster and cheaper, but reliability comes from pairing that machinery with independent evidence and honest denominators. The best number in this case study is not “under 10 seconds.” It is the admission that 68.4% in-scope recall translated into 30.4% of all major incidents. Platform teams that publish and act on both numbers will build better detection systems than teams that optimize only the metric their pipeline can already see.

Sources