DevOps
90 Day OpenTelemetry Application Observability Plan for Engineers
4 October 2026

Application observability is the discipline of inferring a system’s internal state from the external signals it emits, namely metrics, logs and traces, typically instrumented through OpenTelemetry. Where monitoring tells us that something has failed, observability tells us why, and well-instrumented pipelines routinely cut mean time to resolution from tens of minutes to a handful.
TL;DR:
- Using tail sampling ensures error and high-latency traces are retained, improving incident diagnosis without excessive data costs.
- Centralized, multi-tenant backends with pre-aggregation keep telemetry costs manageable while maintaining cross-team visibility.
- Automated root-cause analysis reduces manual triage time from 15-20 minutes to 2-3 minutes by correlating signals using dependency graphs.
- Scaling observability requires careful label hygiene and pre-aggregation to prevent exponential growth in series cardinality and storage costs.
- Implementing a phased approach with initial focus on critical services and defining clear policies helps ensure sustainable, effective observability adoption.
Table of Contents
- What application observability actually means
- Observability versus monitoring: what each actually does
- Choosing a sampling strategy without losing diagnostic power
- Architecture patterns: collectors, allocators and multi-tenant backends
- How automated root-cause analysis cuts triage time
- Keeping observability affordable as telemetry volume grows
- A 90-day plan for adopting application observability
- Who stands behind this guidance
- Instrumenting for observability without breaking context
- Beyond OpenTelemetry: the wider tool landscape
- Security and privacy in an observability pipeline
- Why an open, engineering-first approach pays off
- How we support observability adoption at Vicedomini Softworks
- FAQ
- Sources
What application observability actually means
Application observability rests on the premise that we can reconstruct what a system is doing internally by examining what it emits externally, without having to predict every failure mode in advance. This distinguishes it from static health checks: a system is observable when the telemetry it produces is rich enough to answer questions we have not yet thought to ask. The concept depends on what the industry has settled into as three pillars, each serving a distinct diagnostic function.
Metrics give us aggregated, numerical measures over time, the substrate for service level objectives and key performance indicators such as latency percentiles and error rates. Logs capture discrete events with contextual detail, the narrative record of what happened at a specific moment in a specific component. Traces follow a request across every service it touches, exposing the dependency chain and the latency contributed by each hop.
A fourth category has become central to practical implementations: span-derived metrics, which convert individual trace spans into aggregated, dashboardable time series. A typical example is extracting latency histograms from database call spans to flag slow SQL queries before they degrade a customer-facing endpoint, turning a trace that would otherwise sit unexamined into a metric that triggers an alert.
-
Metrics quantify system health against defined thresholds and SLOs.
-
Logs supply the granular, timestamped context engineers need during investigation.
-
Traces reconstruct the path and latency breakdown of individual requests across services.
-
Span-derived metrics convert trace data into aggregate signals suitable for baseline alerting.
One practitioner workflow uses the spanmetrics connector to export latency histograms from traces into a Prometheus-compatible store, then applies anomaly detection to flag deviations from baseline. This lets teams prioritise optimisation work by traffic-weighted impact rather than by instinct.
Observability versus monitoring: what each actually does
Confusion between these two terms persists because they address overlapping concerns with different intent. Monitoring is reactive: it watches predefined thresholds, runs synthetic checks against known endpoints, and raises an alert when a metric crosses a line we decided in advance mattered. It answers “is something wrong”, and it answers it well, provided we anticipated the failure mode.
Observability addresses what monitoring cannot: the unknown unknowns, the failure patterns nobody wrote a threshold for. According to AWS’s comparison of the two disciplines, observability is a system property that enables root-cause analysis in dynamic, distributed environments, rather than a single tool or dashboard. It achieves this by correlating metrics, logs and traces so an engineer can follow a thread from symptom to cause without first guessing which component is responsible.
The two are not competitors. Monitoring is, in practice, a prerequisite for observability: a mature pipeline still relies on thresholds to decide when a human should look, then hands that human the correlated telemetry needed to diagnose the cause quickly.
- Monitoring detects known failure modes against predefined thresholds.
- Observability supports diagnosis of failure modes nobody anticipated.
- Synthetic checks and alerting remain the trigger; correlated telemetry is the investigation layer.
- Teams that treat observability as “monitoring with more dashboards” tend to under-instrument traces and logs, the signals diagnosis actually depends on.
Workflows that integrate both typically route monitoring alerts directly into an observability backend, so the alert that fires is the entry point into a trace, not a dead end that requires a separate investigation tool.
Choosing a sampling strategy without losing diagnostic power
Collecting every trace from every request is rarely affordable at scale, so sampling policy becomes one of the first architectural decisions a team makes, and getting it wrong either inflates costs or blinds diagnosis exactly when it matters.
- Head sampling decides whether to keep a trace at the moment it starts, typically using a fixed percentage. It is computationally cheap and simple to operate, and it suits high-volume, largely healthy traffic where a representative slice is enough to establish baselines.
- Tail sampling defers the decision until the trace completes, allowing rules to retain traces that ended in an error or exceeded a latency threshold regardless of overall sampling rate. The OpenTelemetry documentation on sampling recommends tail sampling specifically to guarantee that error traces are not discarded by chance, since low-frequency anomalies are exactly what a fixed percentage head sample tends to miss.
- Hybrid strategies combine both: a modest head sample for baseline visibility, paired with tail rules that always capture errors and outliers, with the remaining unsampled data routed to cheaper cold storage rather than discarded outright.
The trade-off is operational, not just financial. Tail sampling requires buffering complete traces before a keep or drop decision, which demands more collector memory and compute than head sampling, and the rules themselves need ongoing maintenance as service behaviour and traffic patterns shift. Teams that set tail-sampling rules once and never revisit them tend to find coverage drifting away from what current incidents actually need.
Pro Tip: Start with head sampling at a low, affordable rate, add tail rules that always retain errors and high-latency traces, and revisit both percentages quarterly as traffic and failure patterns change.
Architecture patterns: collectors, allocators and multi-tenant backends
How we place the OpenTelemetry Collector, and how we route data once it leaves an application, determines whether an observability pipeline stays affordable as it scales or collapses under its own telemetry volume.
Three collector deployment patterns dominate in practice. A daemonset collector runs once per node, intercepting all traffic from pods on that machine, which minimises duplicated network hops but concentrates load. A sidecar collector runs alongside each application container, offering tighter isolation at the cost of duplicated resource overhead per pod. A central, gateway-tier collector receives data forwarded from lighter per-node agents, which centralises processing like tail sampling and pre-aggregation but introduces a network hop and a single point requiring careful capacity planning.

A target allocator solves a problem that emerges once Prometheus-style scraping is involved at any scale: without it, multiple collector replicas can scrape the same targets redundantly, inflating both load and cost. The allocator assigns scrape targets across collector instances so each node’s metrics are collected exactly once, which matters increasingly as fleets grow past a handful of nodes.
On the storage side, teams choose between centralised, multi-tenant backends with strict tenant isolation, and per-cluster stacks replicated across environments. Centralised backends simplify cross-team correlation and reduce operational surface area, but demand careful tenant isolation so one noisy team cannot degrade another’s query performance. Per-cluster stacks avoid that risk at the cost of duplicated infrastructure and harder cross-cluster correlation during incidents that span environments.
- Daemonset collectors suit most Kubernetes fleets as a default, balancing overhead and isolation.
- Target allocators prevent duplicated scraping once more than one collector replica exists.
- Centralised, multi-tenant backends reduce operational surface area but need enforced isolation.
- Pre-aggregation at the tenant-minute level caps series cardinality before it reaches the storage tier.
That last pattern, aggregating metrics by tenant and minute before they hit the time series database, is what keeps cardinality manageable as the number of services and labels grows, a concern we return to in the cost section below.
How automated root-cause analysis cuts triage time
Correlated telemetry only becomes useful once something ties the signals together: timestamp, service topology and the shape of the anomaly itself. Automated root-cause analysis works by correlating metrics, logs and traces across these dimensions, using the service dependency graph as a prior that narrows the search space from “anywhere in the system” to “somewhere downstream of this known-affected node”.
CNCF’s documentation on multi-signal correlation frames the goal precisely: observability is about context, not volume, and producing span-derived metrics alongside dependency graphs is what makes a mountain of traces actually actionable during an incident rather than merely archived.
- Correlation combines signal type, timestamp proximity and topological distance to rank candidate causes.
- Dependency graphs act as a causal prior, so an anomaly in a downstream service is weighted against known upstream failures first.
- Curated runbooks and exclusion rules focus automated agents on relevant evidence instead of exhaustive, unguided search.
- A feedback loop, where responders mark a hypothesis accepted or rejected, tunes inference weights for future incidents.
Practitioners applying automated multi-signal correlation report manual triage time for Kubernetes alerts dropping from 15 to 20 minutes down to roughly 2 to 3 minutes, with around 40% of standard incidents resolved automatically through pattern matching. That scale of reduction depends on the dependency graph being accurate and the runbooks being specific. Automated RCA engines should emit a narrative that cites trace identifiers and the evidence behind each claim, so an engineer reviewing the output can verify the reasoning rather than simply trusting a conclusion.
Keeping observability affordable as telemetry volume grows
Cardinality, the number of unique label combinations a metric can take, is the single most common reason observability budgets spiral. Every new label dimension multiplies the number of distinct time series a backend must store and query, and unbounded labels such as raw user identifiers or request paths can turn a modest metric into millions of series within days.
Label hygiene is the first defence: bound cardinality by using templated route names instead of raw URLs, and reserve high-cardinality identifiers for traces and logs rather than metrics, where they belong. Pre-aggregation is the second: rolling metrics up by tenant and by minute before storage, rather than storing every raw data point, caps series growth at the source rather than fighting it downstream.
On infrastructure sizing, CNCF guidance on cost-effective OpenTelemetry platforms recommends provisioning collectors with at least 4GB of memory, version-locking operator and collector components to avoid compatibility drift, and applying per-tenant pre-aggregation specifically to keep series counts under control at scale.
- Template route and resource names in metric labels; keep raw identifiers in traces and logs only.
- Pre-aggregate by tenant and time window before data reaches the storage tier.
- Size collector nodes with a minimum of 4GB of memory and pin component versions.
- Define a telemetry-coverage SLO, the share of production traffic actually instrumented, so coverage gaps are visible rather than assumed away.
Pro Tip: Treat telemetry coverage as its own SLO with its own dashboard; a pipeline that silently drops 30% of traces is a blind spot nobody notices until the next incident.
One organisation reported a 72% reduction in observability platform costs while moving from 5% sampled traces to near full coverage, simply by standardising on OpenTelemetry and decoupling telemetry generation from backend storage. That combination, cheaper ingestion and better coverage at once, is unusual outside a genuine architectural change.
A 90-day plan for adopting application observability
A phased rollout avoids the common failure mode of instrumenting everything shallowly and nothing deeply. Phase one should cover the handful of services that sit on the critical path, with latency and error traces and a small set of SLOs, before any broader rollout begins.
- Weeks 1 to 4: instrument the top five to ten customer-facing services with OpenTelemetry traces, basic metrics and structured logs.
- Weeks 4 to 8: define SLOs for latency and error rate on those services, and build dashboards that surface breaches against them.
- Weeks 8 to 12: introduce tail sampling for error retention, pre-aggregation for cost control, and a feedback loop where incident responders tag which traces actually helped diagnosis.
| Milestone | Target by day 90 |
|---|---|
| Instrumentation coverage | Critical-path services fully traced |
| Dashboards | SLO breach visibility for each instrumented service |
| Sampling policy | Hybrid head and tail rules in place |
| RCA feedback loop | Responders routinely tagging hypothesis outcomes |
Operational guardrails matter as much as the rollout itself: agree sampling rules, retention windows and a cost ceiling before volume grows, not after a bill arrives that forces a rushed cut to coverage.
Who stands behind this guidance
This guide draws on engineering-first practice from our team, an Italian software engineering company serving organisations across EMEA and North America. Our work spans custom software, cloud-native infrastructure and production observability, built on modern technologies, with every engagement backed by peer-reviewed code and production monitoring.
Instrumenting for observability without breaking context
The hardest part of instrumenting an application is rarely writing the first span: it is keeping context intact as a request crosses process, language and network boundaries. Context propagation, passing trace and span identifiers along with every call, fails silently when a message queue, a background job or a third-party SDK does not forward the right headers, and the result is a trace that simply stops, with no error to point at the gap.
Code-level visibility suffers most in older codebases, where manual instrumentation competes with developer time for features, and partial coverage leaves exactly the paths most likely to fail under load uninstrumented. OpenTelemetry’s auto-instrumentation libraries close much of that gap for common frameworks, but they rarely cover custom internal libraries or legacy integration layers without manual spans added around them.
Good practice treats instrumentation as part of the definition of done for new code, not a retrofit project. Asynchronous boundaries, queues, event buses, scheduled jobs, deserve explicit attention, since context propagation across them usually requires deliberately passing trace headers through message metadata rather than relying on automatic propagation designed for synchronous calls. Naming conventions for spans and attributes matter too: inconsistent naming across teams makes correlation and dashboarding far harder later, so a shared semantic convention, ideally the one OpenTelemetry itself publishes, is worth agreeing early rather than reconciling after the fact.
Beyond OpenTelemetry: the wider tool landscape
OpenTelemetry has become the dominant instrumentation standard, but the backends and platforms that consume its data vary considerably in approach. Prometheus and Prometheus-compatible, multi-tenant stores such as Mimir handle metrics at scale, typically paired with Tempo or an equivalent backend for traces and Loki or an equivalent for logs, forming what practitioners often call an LGTM-style stack built entirely on open standards.
Cloud-vendor platforms take a more integrated route. Microsoft’s Application Insights, for instance, now supports OpenTelemetry ingestion directly and layers application maps, transaction search, live metrics and profiler integration on top, illustrating how vendor platforms are adapting to OTel conventions rather than requiring proprietary instrumentation. That pattern, an open collection standard feeding into either open-source or vendor-managed storage and analysis, is becoming the norm rather than the exception.
The practical choice is less about which single tool wins and more about where we want to own operational complexity. Open-source backends give full control over retention, cardinality limits and query performance, at the cost of running and scaling that infrastructure ourselves. Managed platforms remove that operational burden but typically charge for the data volume and retention we configure, so the sampling and pre-aggregation decisions covered earlier in this guide matter regardless of which backend ultimately stores the data.
Security and privacy in an observability pipeline
Telemetry pipelines carry more sensitive information than most teams initially assume. Traces and logs frequently capture request parameters, headers and database query fragments, any of which can include personal data, authentication tokens or internal identifiers that should never leave an organisation’s control unredacted.
Redaction belongs at the instrumentation or collector layer, not as an afterthought applied at the storage tier. The OpenTelemetry Collector supports processors that can strip or hash sensitive attributes before export, and applying these rules consistently across every service is far more reliable than trusting individual developers to avoid logging sensitive fields.
Access control over the observability backend itself deserves the same scrutiny as access to production data, since a trace can reveal as much about system internals and customer behaviour as the database it describes. Multi-tenant backends need enforced isolation so that one team’s dashboards and queries cannot surface another team’s telemetry, a concern that becomes sharper once observability data spans multiple business units or client environments.
Retention policy is a privacy decision as much as a cost one: keeping high-cardinality, potentially sensitive traces for months by default extends the window during which a breach of the observability platform would expose historical data, so retention periods are worth setting deliberately rather than leaving at a backend’s default.
Why an open, engineering-first approach pays off
The strongest argument for building on OpenTelemetry rather than a proprietary agent is not cost alone, though the cost difference is real. It is that vendor-agnostic pipelines let engineering teams change their storage backend without re-instrumenting every service, which keeps the feedback loop between writing code and understanding its production behaviour short rather than hostage to a renewal cycle.
Engineering-first teams, ones that own their instrumentation decisions rather than inheriting them from a locked-in platform, tend to maintain better telemetry coverage over time, simply because changing a backend no longer means starting instrumentation over. That portability is the real defence against vendor lock-in, and it is worth treating as a architectural requirement, not a nice-to-have.
— Pepe F.
How we support observability adoption at Vicedomini Softworks
Building a sound observability pipeline is an architecture problem before it is a tooling problem, and that is precisely where we work. Our CTO advisory and Fractional CTO engagements give engineering teams direct access to senior architects who have designed telemetry pipelines, sampling policies and cardinality controls like the ones described above.

- An initial assessment maps your current telemetry coverage against the 90-day plan outlined earlier, with written recommendations.
- Architecture and implementation work follows directly from that assessment, through our services page, with no account-manager layer between you and the engineers doing the work.
- Current prices for CTO Advisory services are available on the pricing page on our website.
If your telemetry pipeline needs an architecture review or a hands-on build, get in touch with our team to scope an initial assessment.
FAQ
What is application observability (osservabilità applicativa)?
Application observability is the capacity to infer a system’s internal state from the metrics, logs and traces it emits externally, which lets engineers diagnose failures they did not anticipate. It differs from simple health checks because it supports open-ended investigation rather than only confirming predefined thresholds, as AWS explains in its comparison of the two disciplines.
What are observable benefits in practice?
The practical benefits include faster root-cause analysis, better visibility into dependency chains during incidents, and the ability to prioritise performance work by measured impact rather than guesswork. Teams that correlate signals through dependency graphs have seen manual triage time for Kubernetes alerts fall from 15 to 20 minutes to roughly 2 to 3 minutes.
What is observability as a discipline?
Observability is a system property, not a single product, built on the idea that external telemetry, metrics, logs and traces, should be rich enough to answer unplanned diagnostic questions. OpenTelemetry has become the standard way to generate that telemetry in a vendor-neutral format.
How does monitoring relate to observability?
Monitoring and observability are complementary rather than competing: monitoring watches predefined thresholds and raises the alert, while observability supplies the correlated telemetry needed to diagnose why that alert fired. In practice, mature pipelines route monitoring alerts directly into observability tooling so the alert becomes the entry point into an investigation rather than a separate step.