DevOps
SLO vs SLA: 99.95% SLO Protects a 99.9% SLA for Engineers
7 October 2026

An SLA is the customer-facing contract with consequences. An SLO is an internal reliability target measured by SLIs, and it should sit comfortably stricter than whatever the SLA promises. When an SLA is breached, the result is usually a contractual remedy such as a service credit. When an SLO is missed, the result is operational: an error-budget policy kicks in, releases pause, and engineers redirect toward reliability work.
TL;DR:
- A 99.9% availability SLO allows about 43 minutes of downtime during a month of 30 days; document the remaining budget for engineers responding to incidents.
- Track latency at the 95th percentile, such as a target below 300 milliseconds, because averages can hide slow responses experienced by customers.
- When the error budget runs low, slow risky releases; when it is exhausted, freeze nonessential releases and redirect engineers toward reliability fixes.
- If an SLA reports performance per customer, monitor each customer’s traffic separately; aggregate measurements can conceal a local service failure.
- Use a rolling window of 28 days for internal alerts; calendar month reporting aligns more naturally with customer SLA cycles.
Table of Contents
- SLI, SLO and SLA: the three terms defined
- Who owns SLOs and SLAs, and how they connect
- Measuring SLOs: SLIs, windows and error budgets
- Availability and latency SLOs in practice
- A practical framework for setting SLOs and linking them to SLAs
- What an SLO miss policy and an SLA breach each demand
- Best practices and the pitfalls that undermine them
- How an engineering-first delivery model supports SLO practice
- Why error budgets work better as a shared language than a technical rule
- Turning SLO discipline into a working system
- FAQ
- Sources
SLI, SLO and SLA: the three terms defined
Confusion between these terms persists because they describe a chain, not three unrelated ideas. The SRE handbook treats the Service Level Indicator as the measurable signal, the Service Level Objective as the internal target set against that signal, and the Service Level Agreement as the external promise that may carry financial consequences.
- SLI: a quantitative measurement, such as request latency, error rate or availability percentage.
- SLO: the internal target for that SLI over a defined window, for example “99.9% of requests succeed over 28 days.”
- SLA: the contractual commitment, often set looser than the SLO, which may specify remedies such as service credits when breached.
A numeric example clarifies the gap: an engineering team might set an internal SLO of 99.95% availability while the signed SLA with the customer promises only 99.9%. The difference between those two figures is the safety margin that protects the business from financial exposure every time reliability dips slightly.
Who owns SLOs and SLAs, and how they connect
Ownership splits along functional lines, and that split is deliberate rather than accidental. Engineering and SRE teams typically define SLIs and propose SLOs because they understand what the telemetry can actually support, while product, business and legal functions negotiate and sign the SLA because it carries commercial and legal weight.
- Engineering/SRE: selects SLIs, proposes SLOs, operates the error budget.
- Product and business: translates user expectations into SLO targets worth defending.
- Legal and commercial teams: draft and sign the SLA, including remedies for breach.
The Google Cloud SRE fundamentals post notes that SRE teams usually help shape the SLA rather than author it outright, precisely because the internal target has to stay stricter than the external promise. When an SLA requires per-customer reporting, monitoring that customer’s traffic separately from the aggregate SLI becomes necessary to avoid masking a localised problem.
Measuring SLOs: SLIs, windows and error budgets
Selecting the right SLI matters more than the precision of the measurement itself. A rolling 28-day window smooths out transient spikes and suits internal alerting, whereas a calendar-month window is easier to communicate externally and tends to map more cleanly onto SLA reporting cycles.
The error budget is simply 100% minus the SLO, expressed as the amount of acceptable failure a service can spend before action is required. The error budget policy described in Google’s SRE workbook sets out how teams act once that budget runs low.
- Track budget consumption continuously against the chosen window.
- When the budget runs low, slow the pace of risky releases.
- When the budget is exhausted, freeze non-essential releases and redirect engineering toward reliability fixes.
When error budgets are exhausted, standard SRE practice is to freeze non-essential releases and prioritise reliability work according to Google’s error budget policy, which turns a reliability target into a shared decision rule rather than a purely technical threshold.
Availability and latency SLOs in practice
Abstract targets become usable once translated into plain operational terms. A 99.9% monthly availability SLO permits roughly 43 minutes of downtime across a 30-day month, a figure worth writing into run books so that on-call engineers know exactly how much budget remains after an incident.
- Availability example: 99.9% SLO converts to about 43 minutes of allowable downtime per month.
- Latency example: an SLI defined as “95th percentile response time under 300 milliseconds” gives engineers a concrete target distinct from an average, which hides outliers.
- SLA scope: an SLA may intentionally reference only availability, leaving latency as an internal-only SLO, since customers often care less about the exact percentile than about the service being reachable.
The Datadog documentation on service level objectives recommends starting with a small number of user-facing SLIs, availability and latency being the usual pair, before expanding into freshness, correctness or durability metrics as telemetry matures.
A practical framework for setting SLOs and linking them to SLAs
Choosing the right numbers is less about statistical rigour and more about disciplined sequencing. Skipping a step tends to produce SLOs that are either unrealistically tight or disconnected from what the business has promised customers.
- Pick SLIs that reflect genuine user experience rather than convenient internal counters.
- Choose the lowest SLO that still satisfies users, balancing reliability against delivery speed and cost.
- Decide the measurement window and reporting cadence before negotiating any external commitment.
- Agree SLA language and remedies only after engineering, product and legal have reviewed the proposed targets together.
Pro Tip: Draft the SLO first and let the SLA follow it, never the reverse, so the safety margin is deliberate rather than accidental.
Atlassian’s explainer on SLOs, SLAs and SLIs frames the SLO as the internal mechanism that helps a team actually meet the external expectations written into the SLA, which is why step four always comes last.
What an SLO miss policy and an SLA breach each demand
The operational paths diverge sharply once a target is missed, and conflating them during an incident wastes time that should go toward mitigation.
- Contractual remedy: an SLA breach typically triggers a service credit or another remedy defined in the signed agreement.
- Internal miss policy: an SLO miss triggers a release freeze, a postmortem and a shift of engineering focus toward reliability.
- Detection: instrument the SLI close to the user experience so a miss is caught before it becomes an SLA exposure.
- Ownership and communication: assign a clear owner for the postmortem and a clear owner for any customer-facing communication, since they are rarely the same person.
Early detection is the real defence here: a team that notices SLO erosion days before the SLA threshold is crossed has time to act without ever triggering a contractual remedy.
Best practices and the pitfalls that undermine them
Most SLO programmes fail for the same handful of reasons, and they are avoidable once named plainly.
- Never set the SLO equal to the SLA. The SRE handbook warns that doing so removes the safety margin, turning every minor miss into a contractual crisis.
- Choose SLIs tied to real user experience, not internal counters that look reassuring but miss what customers actually notice.
- Treat the error budget as a shared policy tool between product and engineering, not a purely technical dashboard.
- Resist over-engineering reliability beyond what users need: squeezing out the last fraction of a percentage point often costs far more than it is worth, a trade-off explored in Koritsu’s guide to serverless cost optimisation, which shows how reliability choices ripple directly into infrastructure spend.
Pro Tip: If an SLI cannot be explained to a product manager in one sentence, it is probably measuring the wrong thing.
How an engineering-first delivery model supports SLO practice
Defining the right SLIs and keeping SLOs honest both depend on close, continuous contact between the people writing the code and the people deciding what reliability is worth. We built our delivery model around direct engineer-to-client collaboration specifically because SLI selection, dashboard design and monitoring thresholds are decisions that lose precision when filtered through an account manager.
Every engagement we run includes peer-reviewed code, targeted testing, production monitoring and transparent reporting, which gives the client visibility into exactly how an SLO is tracking rather than a summarised status update. That direct line between engineers and stakeholders is also how latency-sensitive services get the right SLIs from the start, a point echoed in Vadacom’s guide to QoS for VoIP engineers, which shows how real-time services demand metrics chosen for what users actually perceive.

Why error budgets work better as a shared language than a technical rule
The most durable fix for SLO versus SLA confusion is organisational, not technical. Pick one SLI this quarter, set a modest SLO around it, and let the error budget become the sentence product and engineering use to settle release-timing arguments instead of relitigating reliability from scratch each time.
— Pepe F.
Turning SLO discipline into a working system
Getting SLIs, SLOs and SLA language right on paper is one thing. Operating them reliably across dashboards, alerting and release gates over years is another, and it is where many teams stall after the initial framework is agreed.

We offer CTO Advisory, Fractional CTO and CTO Partner engagements for organisations that need governance support to align product and engineering around reliability targets, alongside hands-on delivery through our broader service catalogue.
- Architecture and technology stack work to design observability that supports SLIs.
- Code audit and technical debt reviews to identify issues affecting the error budget.
- Long-term maintenance to keep monitoring and alerting accurate as systems evolve.
Our services overview sets out the full scope if you want to see how an engagement is structured before reaching out.
FAQ
Is an SLO part of an SLA?
An SLO is not part of an SLA. It is an internal target that helps a team meet the wider expectations described in an SLA, as the Atlassian explainer sets out, and the two are typically documented separately.
What is the difference between an SLA and an SOP?
An SLA is an external, often contractual commitment to a customer about service performance, while a Standard Operating Procedure (SOP) is an internal document describing how a team carries out a specific operational task. They serve different audiences and neither replaces the other.
What does SLO mean in ITIL?
Within ITIL practice, an SLO carries the same meaning it does in SRE: a measurable internal target for a service attribute, tracked against an SLI and used to judge whether service delivery meets expectations before any formal SLA threshold is reached.
What does SLO mean in software engineering?
In software engineering, an SLO is the internal reliability target a team sets for a measurable attribute such as availability or latency, tracked through an SLI over a defined time window. Teams use it to decide when to pause feature work and focus on stability, as described in Google’s error budget policy.
Sources
This article was produced with AI assistance and reviewed for accuracy. It is provided for general information only and is not professional advice.