High Availability definition
High availability (HA) is the ability of a system to keep working with minimal downtime when individual components fail, typically by removing single points of failure through redundancy and automatic failover. It is usually expressed as an uptime percentage, such as 99.9 percent or 99.99 percent, measured over a month or a year.
What the nines of uptime mean
Availability targets are often quoted in nines. 99.9 percent allows roughly 43 minutes of downtime in a 30-day month, 99.95 percent about 22 minutes, 99.99 percent about 4 minutes and 99.999 percent under half a minute. Each extra nine typically costs significantly more, because it demands more redundancy, more automation and faster detection of failures.
The right target depends on the cost of downtime. An internal reporting tool may be fine at 99.5 percent, while a payment API or hospital system needs far more. Availability also compounds: a request that depends on three services each at 99.9 percent can only be about 99.7 percent available, so dependencies matter as much as your own code.
How high availability works
High availability comes from assuming every component will fail and arranging for something else to take over automatically. The core techniques appear at every layer of a modern stack, from DNS down to the database:
- Redundancy: at least two of everything, from servers to network paths
- Load balancing: a load balancer spreads traffic and stops sending it to unhealthy instances
- Multiple availability zones: separate data centers within a cloud region, so a power or network failure in one does not stop service
- Database replication with automatic failover to a standby
- Health checks and self-healing, such as Kubernetes restarting failed containers
- Graceful degradation: switching off non-essential features when a dependency fails
Active-active vs active-passive
In an active-passive setup, a standby waits idle and takes over when the primary fails. It is simpler, but failover takes time and the standby may hide configuration drift until the moment it is needed. In active-active, all nodes serve traffic at once, so a failure only reduces capacity. It recovers faster but needs data that can be written in several places, which brings the consistency trade-offs described by the CAP theorem.
Within a cloud region, active-active across availability zones is now the normal pattern for stateless tiers, while managed databases such as Amazon RDS Multi-AZ or zone-redundant Azure SQL handle failover for the data tier with a short, automatic switchover.
High availability vs disaster recovery
High availability handles routine failures automatically, such as a crashed server, a failed disk or a lost availability zone, usually with seconds of disruption and no data loss. Disaster recovery handles rarer events that take out a whole region, corrupt data or involve ransomware. A system needs both, because HA does not protect against a bad deployment or a deleted table replicated instantly to every copy.
Nexzem designs HA architectures on AWS, Azure and Google Cloud sized to the business case, and tests failover in practice, because a standby that has never been exercised is a hope rather than a plan, and the first real failover is the worst time to discover a missing permission.