Designing a Redis-Based Task Queue
At-least-once delivery with Lua acquisition, heartbeat leases, an independent Janitor and a Dead Letter Queue — stress-tested at 10K concurrent jobs.
By Rahul Gupta · RahulX Labs
List-popping is not enough
Background job systems fail in boring ways: workers crash mid-task, networks partition, and poison messages loop forever. A queue that only does RPOP has a window where a crash loses work between pop and “I own this”.
The queue I built uses custom Redis Lua to combine RPOP + ZADD into one atomic acquisition into a processing Sorted Set scored by lease timestamp.
Heartbeats and an independent Janitor
Workers emit heartbeats to extend leases. An independent Janitor process sweeps expired leases and re-queues zombie tasks — recovery is not coupled to the workers that just died.
Tasks that exceed MAX_RETRIES move to a Dead Letter Queue so poison pills never block the hot path. The React dashboard exposes depth, throughput, DLQ size and manual reprocess controls.
Load results
Under burst load the system ingested ~15,800 tasks/sec and processed 10,000 concurrent jobs in 631ms with 0% task loss in the crash-recovery scenarios I tested.
Atomic Lua, a dedicated reaper and DLQ tooling turn failure from silent data loss into an operable workflow.