Skip to content

What is High Availability?

DevOps & Reliability, explained by the engineers who build it. Definition, how it works, use cases and common questions.

High Availability definition

High availability (HA) is the ability of a system to keep working with minimal downtime when individual components fail, typically by removing single points of failure through redundancy and automatic failover. It is usually expressed as an uptime percentage, such as 99.9 percent or 99.99 percent, measured over a month or a year.

What the nines of uptime mean

Availability targets are often quoted in nines. 99.9 percent allows roughly 43 minutes of downtime in a 30-day month, 99.95 percent about 22 minutes, 99.99 percent about 4 minutes and 99.999 percent under half a minute. Each extra nine typically costs significantly more, because it demands more redundancy, more automation and faster detection of failures.

The right target depends on the cost of downtime. An internal reporting tool may be fine at 99.5 percent, while a payment API or hospital system needs far more. Availability also compounds: a request that depends on three services each at 99.9 percent can only be about 99.7 percent available, so dependencies matter as much as your own code.

How high availability works

High availability comes from assuming every component will fail and arranging for something else to take over automatically. The core techniques appear at every layer of a modern stack, from DNS down to the database:

  • Redundancy: at least two of everything, from servers to network paths
  • Load balancing: a load balancer spreads traffic and stops sending it to unhealthy instances
  • Multiple availability zones: separate data centers within a cloud region, so a power or network failure in one does not stop service
  • Database replication with automatic failover to a standby
  • Health checks and self-healing, such as Kubernetes restarting failed containers
  • Graceful degradation: switching off non-essential features when a dependency fails

Active-active vs active-passive

In an active-passive setup, a standby waits idle and takes over when the primary fails. It is simpler, but failover takes time and the standby may hide configuration drift until the moment it is needed. In active-active, all nodes serve traffic at once, so a failure only reduces capacity. It recovers faster but needs data that can be written in several places, which brings the consistency trade-offs described by the CAP theorem.

Within a cloud region, active-active across availability zones is now the normal pattern for stateless tiers, while managed databases such as Amazon RDS Multi-AZ or zone-redundant Azure SQL handle failover for the data tier with a short, automatic switchover.

High availability vs disaster recovery

High availability handles routine failures automatically, such as a crashed server, a failed disk or a lost availability zone, usually with seconds of disruption and no data loss. Disaster recovery handles rarer events that take out a whole region, corrupt data or involve ransomware. A system needs both, because HA does not protect against a bad deployment or a deleted table replicated instantly to every copy.

Nexzem designs HA architectures on AWS, Azure and Google Cloud sized to the business case, and tests failover in practice, because a standby that has never been exercised is a hope rather than a plan, and the first real failover is the worst time to discover a missing permission.

High Availability: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What does 99.99 percent uptime mean?

It means the service may be unavailable for about 52 minutes per year, or roughly 4 minutes per month. Achieving it usually requires redundancy across availability zones, automated failover, careful deployments and fast incident response, since a single slow manual recovery can use up the whole year's allowance.

Is high availability the same as fault tolerance?

They are related but not identical. Fault-tolerant systems keep running with no interruption at all when a component fails, often through fully duplicated hardware. Highly available systems accept a brief disruption, such as seconds of failover, in exchange for lower cost and complexity. Most web services aim for high availability.

Does Kubernetes make an application highly available?

It helps by restarting failed containers, spreading replicas across nodes and routing only to healthy pods. But availability still depends on running multiple replicas across zones, a resilient control plane and database, safe deployment strategies and dependencies that are themselves available. Kubernetes is a tool, not a guarantee.

Keep exploring the devops & reliability glossary

Need High Availability in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.