Engineering notes
Distributed Task Queue
Stress-tested at 10,000 concurrent jobs with ~15,800 tasks/sec ingestion and 0% task loss under crash-recovery scenarios.
The challenge
Background job processing needs guarantees beyond list-popping. Workers crash, networks partition, and poison messages can block the queue forever. The system needed at-least-once delivery, automated recovery from worker failures, poison-pill isolation, and real-time visibility into queue health.
Important decisions
Lua scripts for atomic acquisition
Custom Redis Lua combines RPOP + ZADD into one atomic transaction, closing the microsecond gap where a crash could lose a task between pop and claim.
Independent Janitor (reaper)
A detached microservice sweeps queue:processing for zombie tasks whose heartbeat expired and safely re-queues them — recovery is not coupled to worker processes.
Heartbeat lease + Sorted Set state
Tasks move into a Sorted Set scored by timestamp. Workers emit heartbeats to extend the lease; expired leases become reclaimable by the Janitor.
Measured results
~15,800 tasks/sec
Ingestion
10,000 jobs / 631ms
Burst load
0%
Task loss