The short answer
Quick answer: Auto-scaling automatically adds or removes computing capacity to match demand. An autoscaler watches a metric, such as average CPU use or requests per instance, and compares it with a target. When the metric rises above the target, it starts more instances (scaling out); when it falls, it removes some (scaling in), within limits you set. A load balancer spreads traffic across whatever instances exist. The benefits are that the service stays responsive during spikes and you stop paying for idle capacity when things are quiet. The main catch is that scaling takes time, so truly sudden spikes also need spare capacity and other defences.
The problem
Traffic is never flat. It follows daily and weekly cycles, jumps during sales or news events, and occasionally surges without warning.
You have two bad options without auto-scaling:
- Provision for the peak. The site copes, but most of the hardware sits idle most of the time, and you pay for all of it.
- Provision for the average. It is cheap, and the site falls over whenever demand rises.
Auto-scaling replaces a fixed guess with a control loop.
Horizontal and vertical
| Horizontal (out / in) | Vertical (up / down) | |
|---|---|---|
| What changes | The number of instances | The size of an instance: more CPU and memory |
| Limit | Practically none | The largest machine available |
| Disruption | None; instances are added behind a load balancer | Often needs a restart |
| Fault tolerance | Improves: more instances | Unchanged: still one machine |
| Requirement | The application must be able to run as many copies | None |
Auto-scaling normally means horizontal scaling. Vertical scaling is still useful for systems that are hard to distribute, such as a single primary database.
The control loop
Every autoscaler runs the same cycle:
- Measure. Collect a metric across the group.
- Compare it with the target.
- Decide how many instances are needed.
- Act. Launch or terminate instances.
- Wait, then repeat.
A load balancer sits in front and distributes requests across healthy instances, adding new ones as they pass health checks. See what a load balancer does.
The calculation
The Kubernetes Horizontal Pod Autoscaler uses a simple proportional formula:
desired replicas = ceil(current replicas × current metric / target metric)
If 4 pods are running at 90% CPU and the target is 60%:
ceil(4 × 90 / 60) = 6
It scales to 6 pods, which should bring the average back to about 60%. The Kubernetes autoscaling documentation describes this and the other autoscalers.
You also set a minimum and maximum number of instances. The minimum guarantees availability and some headroom. The maximum caps your bill and protects downstream systems.
Types of scaling policy
| Policy | How it works | Good for |
|---|---|---|
| Target tracking | "Keep average CPU at 50%"; the system adds and removes capacity to hold it | The sensible default |
| Step scaling | "If CPU is above 70%, add 2; above 90%, add 5" | Fine control over the response |
| Scheduled | "Run 20 instances from 08:00 on weekdays" | Predictable patterns and known events |
| Predictive | Forecast demand from history and scale ahead of it | Regular daily or weekly cycles |
AWS's EC2 Auto Scaling documentation covers these for virtual machines.
Choosing what to measure
The metric must rise and fall in proportion to load, and adding instances must bring it down.
| Metric | Suits |
|---|---|
| CPU utilisation | CPU-bound services; the common default |
| Requests per instance | Web services and APIs |
| Queue length, or the age of the oldest message | Background workers. See how message queues work |
| Concurrent connections | WebSocket and streaming servers |
| A custom business metric | Anything that reflects real work |
CPU is not always right. A service that mostly waits on a database can be overloaded while its CPU is nearly idle. For queue workers, the size of the backlog is the meaningful signal. Event-driven autoscalers such as KEDA scale on these external signals and can scale to zero when there is no work.
Memory is usually a poor trigger: many runtimes do not release memory when load drops, so the metric never comes back down.
All of this depends on good metrics; see logging, metrics and tracing.
Why scaling is not instant
A spike that arrives in seconds can outrun the autoscaler. The delays add up:
- Detection. Metrics are collected and averaged over a period.
- Decision. The policy may require the condition to persist.
- Provisioning. A new virtual machine takes a minute or more; a container takes seconds if a node has room.
- Start-up. The application must start, load configuration and warm its caches.
- Health checks. The load balancer waits until the instance reports ready.
Altogether this is often one to several minutes. Ways to shorten it or cope with it:
- Keep headroom. Set the target well below 100%, for example 50 to 60%, so existing instances can absorb a surge while new ones start.
- Make start-up fast. Use pre-built images and small containers, and do less work at launch.
- Keep warm spare capacity ready to enter service.
- Scale ahead of known events with scheduled or predictive policies.
- Buffer with a queue so work waits instead of failing.
- Shed load with rate limiting. See how rate limiters work.
- Cache at the edge so less traffic reaches your servers. See how CDNs work.
Avoiding flapping
If the autoscaler reacts to every wobble, it adds instances, sees the metric drop, removes them, sees it rise, and so on. This oscillation is called flapping or thrashing.
The remedies:
- Cooldown or stabilisation periods: wait after a scaling action before taking another.
- Scale out quickly, scale in slowly. Being briefly over-provisioned is cheap. Being under-provisioned causes errors.
- Warm-up time: do not count a new instance's metrics until it is fully serving.
Scaling in safely
Removing an instance must not cut off the requests it is handling.
- Connection draining: the load balancer stops sending new requests and lets existing ones finish.
- Graceful shutdown: the application handles the termination signal, completes in-flight work and closes connections.
- Long-running jobs need protection from termination or must be safe to retry.
What the application must do
Auto-scaling assumes instances are interchangeable.
- Stateless servers. Sessions, uploads and caches must live in shared stores, not on an individual instance.
- Fast, self-contained start-up.
- Health endpoints that accurately report readiness.
- Disposable instances. Any one can vanish at any moment.
The layers that scale
In a container platform, scaling happens at more than one level. See what Kubernetes does.
| Autoscaler | Adjusts |
|---|---|
| Horizontal Pod Autoscaler | The number of pods |
| Vertical Pod Autoscaler | Each pod's CPU and memory allocation |
| Cluster autoscaler | The number of nodes (machines) in the cluster |
More pods need somewhere to run. If the nodes are full, the cluster autoscaler adds a node, which takes longer.
Serverless platforms take this to its conclusion: scaling is per request and fully managed. See what serverless really means.
Pitfalls
- The bottleneck moves. Doubling your web servers doubles the load on the database, which may not scale the same way. Check connection limits; see connection pooling.
- Runaway cost. A bug, a retry storm or an attack can scale you to the maximum. Set sensible limits and billing alerts.
- Scaling in response to an attack just pays for the attacker's traffic. Filter it first.
- The wrong metric, so the system scales when it should not, or does not when it should.
- A maximum set too low, discovered during the busiest hour of the year.
- Provider quotas and capacity limits that block new instances.
- Untested behaviour. Run load tests so you know how the system scales before it matters.
Frequently asked questions
What is the difference between horizontal and vertical scaling?
Horizontal scaling changes the number of instances. Vertical scaling changes the size of each instance.
How fast does auto-scaling react?
Typically within one to a few minutes, covering detection, launching instances and start-up. Containers are quicker than virtual machines.
What metric should I scale on?
One that tracks load directly. CPU is a reasonable default for web services; queue length works better for background workers.
Can auto-scaling handle a sudden viral spike?
Partly. It will catch up, but not instantly. Headroom, caching, queues and rate limiting cover the gap.
Conclusion
Auto-scaling is a feedback loop: measure, compare with a target, adjust capacity. It keeps a service responsive and avoids paying for idle machines, provided the application is stateless, the metric reflects real load, and you allow for the minutes it takes new capacity to arrive. Combine it with headroom and a few protective measures, and sudden spikes become manageable.
Related articles
- What Is a Load Balancer and How Does It Decide Where Traffic Goes?
- What Kubernetes Does and Why Companies Use It
- What "Serverless" Really Means (Spoiler: There Are Servers)
- How Logging, Metrics, and Tracing Help You Debug Production
