SLA, SLO and SLI definition
SLA, SLO and SLI are related reliability terms. A service level indicator (SLI) is a measured metric, such as the share of successful requests. A service level objective (SLO) is the internal target for that metric, such as 99.9 percent. A service level agreement (SLA) is a contractual promise to customers, with credits or penalties if it is missed.
How SLIs, SLOs and SLAs fit together
Start with the indicator. An SLI measures something users care about, expressed as good events divided by total events: the percentage of API requests that succeed, the percentage served in under 300 ms, or the percentage of uploads processed within a minute. The SLO sets a target for that SLI over a window, for example 99.9 percent of requests succeed over 28 days.
The SLA is a business document. It promises customers a level of service, often lower than the internal SLO, and defines what happens if it is missed, usually service credits. Keeping the SLO stricter than the SLA gives the team warning before a contractual breach. Site reliability engineering practices built this vocabulary into everyday operations.
Examples for a web application
A typical SaaS product might define its reliability like this, with SLIs measured at the load balancer or from synthetic checks rather than estimated from application logs alone, which miss failures that happen before traffic reaches the app.
Notice that the SLA is looser than the SLO, and that each SLI is defined precisely enough that two engineers would calculate the same number from the same data. Vague definitions, such as uptime without saying how it is measured, cause disputes exactly when something goes wrong. An example set:
- Availability SLI: share of HTTP requests that do not return 5xx errors
- Latency SLI: share of requests completed within 300 ms
- SLO: 99.9 percent availability and 95 percent of requests under 300 ms over 28 days
- SLA: 99.5 percent monthly availability, with a service credit below that
- Freshness SLI for data products: share of dashboards updated within 15 minutes of new source data
Error budgets
An SLO of 99.9 percent implies a 0.1 percent error budget: about 43 minutes of full downtime in a 30-day month, or the equivalent in partial failures. The budget turns reliability into a shared, measurable decision. While budget remains, teams ship features freely; when it is nearly spent, they slow releases and prioritize stability work. Alerts based on how fast the budget burns are far less noisy than alerts on every error spike.
Error budgets also end arguments about perfection. Aiming for 100 percent is neither possible nor worth the cost, since users cannot tell 99.99 percent from 100 percent through their own networks and devices, while the engineering effort to chase that last fraction is enormous.
How to set good SLOs
Choose a few SLIs tied to user journeys, such as login, search and checkout, rather than server metrics like CPU. Base targets on current performance and what users actually need, start slightly loose and tighten as you learn. Make sure observability tooling can measure each SLI accurately, and review SLOs quarterly with product owners.
When you sign an SLA with a vendor, check how availability is measured, what is excluded, such as scheduled maintenance, and whether credits are meaningful. Nexzem defines SLOs and monitoring for systems it runs under managed cloud services, so support commitments are measured rather than assumed.