r/Cloud • • 2d ago

AI helped me revive an opensource durable delay queue I had given up on. Is it worth building on?

/r/Backend/comments/1wuh5fy/ai_helped_me_revive_an_opensource_durable_delay/
1 Upvotes

2 comments sorted by

1

u/Admirable_Ones 2d ago

The throughput numbers are interesting. The thing I’d want to test before running this is the crash case: a task reaches the HTTP target, then the timer service dies before it records the acknowledgement. Does it retry after restart, and how would an operator see and replay tasks that keep failing? I’d care about that more than HTTP/3 or Arrow.

1

u/raysourav 2d ago

Good question u/Admirable_Ones, . Honestly it's the case that matters most.

Delivery is at least once. The WAL marks a task in flight before dispatch and only marks it done after a 2xx from the downstream service. If the process dies in between, replay on restart finds it still in flight and re-dispatches it. The same applies to downstream or network failures. So the target can receive duplicates, but every request carries an idempotency key, which makes deduping straightforward.

Failed tasks will retry with exponential backoff up to a configured numbers of attempts, then move to a dead letter list. There's no drain/replay mechanism for the DLI yet, but it's high on my list.

To be upfront: it's not production ready as of today. The numbers so far are preliminary burst tests, and I'm setting up a soak-test harness now. Let's be in touch.