Engineering notes

Distributed Task Queue

Stress-tested at 10,000 concurrent jobs with ~15,800 tasks/sec ingestion and 0% task loss under crash-recovery scenarios.

The challenge

Background job processing needs guarantees beyond list-popping. Workers crash, networks partition, and poison messages can block the queue forever. The system needed at-least-once delivery, automated recovery from worker failures, poison-pill isolation, and real-time visibility into queue health.

Important decisions

Lua scripts for atomic acquisition

Custom Redis Lua combines RPOP + ZADD into one atomic transaction, closing the microsecond gap where a crash could lose a task between pop and claim.

Independent Janitor (reaper)

A detached microservice sweeps queue:processing for zombie tasks whose heartbeat expired and safely re-queues them — recovery is not coupled to worker processes.

Heartbeat lease + Sorted Set state

Tasks move into a Sorted Set scored by timestamp. Workers emit heartbeats to extend the lease; expired leases become reclaimable by the Janitor.

Measured results

~15,800 tasks/sec

Ingestion

10,000 jobs / 631ms

Burst load

0%

Task loss

Full case study →Distributed systems service