Distributed Systems8 min

How I Built a Distributed Rate Limiter

Atomic Redis Lua rate limiting with three algorithms, per-API-key tiers, Prometheus metrics and real k6 numbers.

By Rahul Gupta · RahulX Labs

Why in-process rate limits fail

A counter in memory works on one server. The moment you run multiple Node processes or multiple machines, each instance tracks its own budget — and clients can exceed the intended limit by hitting different instances.

Shared Redis state fixes the visibility problem, but GET → CHECK → INCR is still a race under concurrency. Two requests can both read “under limit” before either increments.

Atomic enforcement with Redis Lua

The limiter I built runs Sliding Window Counter, Sliding Window Log, and Token Bucket entirely inside Redis Lua scripts. The check and increment happen atomically inside Redis, so concurrent hammering of one API key cannot slip between processes or PM2 workers.

Each algorithm maps to a tier (anonymous / free / pro) so the tradeoffs are measurable in the dashboard and in k6 runs — not only described in comments.

  • Anonymous: Sliding Window Counter — cheap approximate budgets
  • Free: Sliding Window Log — exactness when request volume is lower
  • Pro: Token Bucket — burst capacity with steady refill

Observability as part of the design

Prometheus histograms and counters, /health vs /ready probes, structured JSON logs, and OpenAPI docs shipped with the limiter — not bolted on after the demo.

On SIGTERM the server drains in-flight requests and closes Redis cleanly, which matters for rolling deploys.

What the load tests showed

k6 runs against the anonymous tier produced 99.47% 429 responses when over budget, with 12.37ms p95 latency on enforcement. Pro-tier allowed traffic tracked theoretical refill closely (~349 allowed vs 300 + refill).

Those numbers live in the case study for the Distributed Tiered Rate Limiter — the article is the narrative; the case study is the architecture dump.