Distributed Systems9 min

Designing a Redis-Based Task Queue

At-least-once delivery with Lua acquisition, heartbeat leases, an independent Janitor and a Dead Letter Queue — stress-tested at 10K concurrent jobs.

By Rahul Gupta · RahulX Labs

List-popping is not enough

Background job systems fail in boring ways: workers crash mid-task, networks partition, and poison messages loop forever. A queue that only does RPOP has a window where a crash loses work between pop and “I own this”.

The queue I built uses custom Redis Lua to combine RPOP + ZADD into one atomic acquisition into a processing Sorted Set scored by lease timestamp.

Heartbeats and an independent Janitor

Workers emit heartbeats to extend leases. An independent Janitor process sweeps expired leases and re-queues zombie tasks — recovery is not coupled to the workers that just died.

Tasks that exceed MAX_RETRIES move to a Dead Letter Queue so poison pills never block the hot path. The React dashboard exposes depth, throughput, DLQ size and manual reprocess controls.

Load results

Under burst load the system ingested ~15,800 tasks/sec and processed 10,000 concurrent jobs in 631ms with 0% task loss in the crash-recovery scenarios I tested.

Atomic Lua, a dedicated reaper and DLQ tooling turn failure from silent data loss into an operable workflow.