Skip to main content

The Art of CTO API Rate Limit Planner sizes a token-bucket refill rate and burst capacity from your expected and peak request rates, sets per-minute quotas and burst ceilings for each pricing tier, and compares the four common throttling algorithms with a Redis memory estimate for each. It does not emit gateway configuration files.

What rate limits should this API have?

A token-bucket refill rate and burst capacity sized from your traffic, plus per-tier quotas and burst ceilings.

About 10 min · Calculator · Free

About this toolWhy it matters, common mistakes, FAQ

What Rate Limits Should This API Have?

Rate limits are a product decision wearing operational clothing. Set them too low and you break integrations you wanted; too high and one customer's retry loop becomes everyone's outage.

A single global limit is applied because it is easy to reason about. Real traffic is bursty and uneven — limits that work need to reflect tiers, endpoints and burst behaviour, or they throttle the wrong people.

Questions CTOs ask

What rate limiting algorithm should I use for my API?
The four main algorithms are fixed window (simplest, but allows burst at window boundaries), sliding window (smoother distribution, slightly more complex), token bucket (allows controlled bursts while maintaining average rate — best for most APIs), and leaky bucket (strictly smooth output, good for upstream protection). Token bucket is the most popular choice for public APIs because it accommodates legitimate traffic bursts while preventing abuse. For internal microservices, sliding window counters provide a good balance of accuracy and simplicity.
How do you set appropriate API rate limits?
Base rate limits on your infrastructure capacity divided by the number of expected consumers, with a safety margin of 20-30%. Analyze actual usage patterns to understand p50, p95, and p99 request volumes per client. Set burst limits at 2-5x the sustained rate to accommodate legitimate spikes. Implement tiered limits based on plan level (free, pro, enterprise) and communicate limits clearly via response headers (X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset). Start generous and tighten based on observed abuse patterns.

Related Reading