Rate Limiting definition
Rate limiting is a technique that caps how many requests a client, user or IP address can make to a service within a time window, such as 100 requests per minute. It protects APIs and websites from abuse, brute-force attacks and accidental overload, and keeps capacity fair between customers. Requests over the limit are typically rejected with HTTP 429.
Why rate limiting matters
Every service has finite capacity. Without limits, one misbehaving client, a scraper, a buggy retry loop or a deliberate attack can consume it all and degrade the service for everyone else. Rate limits stop credential stuffing on login pages, slow down card-testing fraud at checkout, protect expensive endpoints such as search or AI generation, and keep a free tier from starving paying customers.
They also protect your costs. An endpoint that calls a paid third-party API or a large language model can generate a surprising bill if a single client loops on it. Per-user limits on such endpoints are cheap insurance, and they make capacity planning predictable because peak load has a known ceiling.
Rate limiting algorithms
The algorithm decides how strictly bursts are handled and how much state the limiter must store. Most API gateways and libraries implement one of these four, usually backed by a shared in-memory store such as Redis so every server sees the same counters.
Token bucket is a common choice for public APIs because it tolerates natural bursts, such as a page loading several resources at once, while still enforcing a long-run average. Whatever the algorithm, key limits on something meaningful, such as API key, user ID, tenant or IP address, depending on who you want to constrain. The four common algorithms are:
- Fixed window: count requests per calendar minute; simple, but allows double bursts at window boundaries
- Sliding window: smooths the boundary problem by weighting the previous window or logging timestamps
- Token bucket: tokens refill at a steady rate and each request spends one, allowing short bursts up to the bucket size
- Leaky bucket: requests drain at a constant rate, smoothing traffic into a steady stream
Responding to limits: HTTP 429 and client behavior
When a client exceeds its limit, return HTTP 429 Too Many Requests with a Retry-After header saying when to try again, and ideally headers showing the limit, remaining requests and reset time. Clear signals let well-behaved clients slow down instead of hammering the service. Our HTTP status codes reference lists 429 alongside related codes.
On the client side, respect Retry-After, retry with exponential backoff and jitter, and never retry non-idempotent operations blindly. Queue work that can wait and batch requests where the API supports it. Many integration failures with Stripe, Shopify, Salesforce or LLM APIs under load turn out to be rate-limit handling bugs rather than provider outages.
Where to enforce limits and how to choose them
Limits can live at several layers: a CDN or web application firewall for IP-based protection against floods, an API gateway for per-key and per-plan quotas, and application code for business rules such as five password attempts per account per hour. Layering matters, because IP limits alone fail against attackers using many addresses, while per-account limits do nothing for anonymous endpoints.
Choose numbers from real traffic: look at normal usage per client, set limits comfortably above it and log near-misses before enforcing. Nexzem configures rate limiting on APIs and login flows during application security work, tuning limits so attackers are slowed down without legitimate users ever noticing.