Auto Scaling definition
Auto scaling is a cloud capability that automatically adds or removes computing resources, such as servers, containers or function instances, based on demand. Policies watch metrics like CPU use, request rate or queue length, scaling out when load rises and scaling in when it falls, so applications stay responsive without paying for idle capacity.
How does auto scaling work?
An auto scaling group defines a template for new instances, such as a machine image or container spec, plus a minimum, maximum and desired count. A monitoring service such as Amazon CloudWatch, Azure Monitor or Google Cloud Monitoring collects metrics, and scaling policies compare them with targets. When load rises, new instances launch, pass health checks and register with the load balancer. When load falls, instances are drained of active connections and removed.
A typical policy for a web API keeps average CPU around 60 percent. If traffic doubles at lunchtime, CPU climbs, the policy adds instances until the average returns to target, and removes them again in the afternoon. The application must be stateless, storing sessions and files in shared services, so any instance can be added or removed at any time without losing user data.
Types of auto scaling
- Horizontal scaling: add or remove instances; the standard approach in the cloud.
- Vertical scaling: give an instance more CPU or memory; limited and often needs a restart.
- Target tracking: keep a metric, such as CPU or requests per instance, near a set value.
- Step scaling: add or remove a set number of instances at thresholds.
- Scheduled scaling: change capacity at known times, such as business hours.
- Predictive scaling: forecast load from history and scale ahead of it.
- Event-driven scaling: scale on queue length or events, including down to zero.
Auto scaling in Kubernetes
Kubernetes scales at two levels. The Horizontal Pod Autoscaler adds or removes pod replicas based on CPU, memory or custom metrics, and the Vertical Pod Autoscaler adjusts their resource requests. When pods no longer fit on existing machines, the Cluster Autoscaler or Karpenter adds nodes, and removes them when they sit underused. KEDA extends this with event-driven scaling from queues, streams and databases, including scaling idle workloads to zero.
Common auto scaling pitfalls
Most scaling problems come from the system around the scaling group, not from the policy itself. Load tests that push past expected peaks reveal these limits before customers do. Each one needs its own fix, and none is solved by a better scaling policy alone.
- Slow startup: instances that take minutes to boot cannot absorb sudden spikes; use pre-built images or warm pools.
- The wrong metric: CPU does not reflect load for I/O-bound or queue-driven services.
- Flapping: aggressive policies that scale out and in repeatedly; add cooldowns.
- Bottlenecks that do not scale, such as a single database or a third-party API rate limit.
- Account quotas and maximum limits that block scaling at the worst moment.
- Runaway costs from bugs, retry storms or attacks without a sensible maximum.
Example: a ticketing flash sale
A ticketing platform knows a popular concert goes on sale at 10:00. A scheduled action raises minimum capacity at 9:45, so instances are warm before the rush. Target tracking then handles the actual surge, a queue smooths checkout requests to protect the payment provider, and the database uses read replicas for seat browsing. After the sale, capacity falls back automatically. Nexzem designs scaling like this with load tests that prove each limit before launch day.