← blog · September 2, 2026

Queues and Async Job Processing: Where They Help, Where They Hurt

Pushing every slow operation into a queue does not remove complexity, it relocates it. A practical look at when async processing actually pays off, what delivery guarantees really mean, and why idempotency is not optional.

The problem: not everything can stay synchronous

When an HTTP request lands, the server has a limited window to respond. Users tolerate a few hundred milliseconds, put up with a few seconds, and abandon the page or the app past ten. But not every task fits inside that window: transcoding a video, sending a batch of emails, calling a third-party API that might be slow, generating a large report.

The obvious fix is to move the work off the request path: accept the request immediately, tell the user it was received, push the actual work onto a queue, and let a separate process (a worker) consume it. The catch is that adding a queue does not remove complexity, it relocates it. In synchronous code, a failure surfaces to the user immediately and they retry. In asynchronous code, a failure accumulates silently somewhere, nobody notices right away, and by the time it is discovered three days later, reconstructing what failed and when becomes its own investigation.

When a queue actually earns its place

What a queue buys you is not speed, it is decoupling: accepting work becomes independent from processing it. That is worth the cost when:

  • The work takes far longer than a user would ever wait for (multi-minute conversions, bulk imports).
  • The work needs to be retried independently of the caller (webhook delivery, email sending, notifying a payment provider).
  • Load arrives in bursts that need to be absorbed (a campaign triggers thousands of notifications at once, but the number of workers can stay fixed).
  • Different components need to scale at different rates. The web layer should respond fast, the processing layer should be slow but reliable; forcing both into the same process makes both worse.

The common thread is that the work does not need the caller's live connection. The caller can move on without waiting for the result and learning about it later, through a notification, polling, or a webhook, is acceptable.

When it hurts more than it helps

The most common mistake is reflexively queuing anything that feels slow. A few concrete cases:

If the user must see the result immediately, a queue is the wrong tool. Payment confirmation, a login attempt, form validation should stay synchronous; making them asynchronous and polling for the result both degrades the experience and forces an unnecessary state machine (pending / done / failed) into the UI.

If the operation is already fast, a queue only adds latency and operational overhead. If a database update takes 10 milliseconds, queuing it gains nothing and costs you one more worker process to monitor, one more queue to manage, and one more hop to trace through when debugging.

If ordering matters and the queue does not guarantee it, subtle bugs creep in. If a user's profile update events are processed in parallel across workers, a stale update that entered the queue earlier but finishes later can silently overwrite newer data. Most general-purpose queues (aside from within a single Kafka partition) offer no global ordering guarantee; if order matters, you enforce it either with a partition key or with application-level versioning.

If the debugging cost is ignored, the queue itself becomes a black box. A synchronous failure comes with a stack trace; an asynchronous failure usually shows up as "this job never showed up again" or "the same job ran five times," and tracing the source requires correlating logs across processes.

Delivery guarantees: at-least-once wins almost every time

Queue systems offer three theoretical delivery models: at-most-once (a message can be lost but never reprocessed), at-least-once (a message is never lost but can be processed more than once), and exactly-once (both guarantees at once, expensive in practice and rarely fully true in real systems). Nearly all production systems settle on at-least-once, because losing a message is usually worse than reprocessing one.

The cost is explicit: your worker can receive the same job twice. A worker processes a message, crashes before committing the result, the queue makes the message visible again once the visibility timeout expires, and another worker picks up the same job. This is not a theoretical edge case, it happens routinely in every production system.

Idempotency is a prerequisite, not an option

Once you accept at-least-once delivery, every worker function must be idempotent: running the same job twice must produce the same outcome as running it once. The standard way to achieve this is an idempotency key: the job carries a unique identity (for example "generate the 2026-09 invoice for user 42"), and the worker checks, before doing any work, whether that identity has already been processed, typically against a table with a unique constraint or a key in Redis. If it has, the worker exits silently.

Systems that skip this show the same recurring symptoms: a user receives the same email twice, a payment gets charged twice, a counter increments twice. All of them trace back to the same root cause, an unqualified side effect in the worker.

Retries, backoff, and dead-letter queues

When a job fails, there are three options: retry immediately, retry after a delay, or give up. Retrying immediately on a non-transient failure (corrupt data, for instance) sends the queue into a pointless loop and puts extra load on whatever downstream service is already struggling. The standard practice is exponential backoff: the first retry after 1 second, the next after 2, the next after 4, with a random jitter added on top so thousands of jobs do not retry in lockstep and create a new load spike.

Once a job crosses a retry limit (commonly 5-10 attempts), it should move to a dead-letter queue (DLQ) instead of retrying forever. The DLQ's purpose is not to discard the job but to pull it out of the automatic loop and expose it to human review. Teams that do not monitor their DLQ end up with failed jobs piling up unnoticed for months, which is why alerting on anything landing in the DLQ matters as much as setting the queue up in the first place.

Choosing a tool

Before standing up a full messaging system, the simplest option is often overlooked: if you already run a relational database, PostgreSQL's SELECT ... FOR UPDATE SKIP LOCKED lets you use an ordinary table as a queue, which is entirely sufficient for low-to-medium volume workloads. You avoid operating a separate messaging stack and get to reuse the transactional guarantees you already have.

As volume grows, or when you need to publish messages across independent services, a classic message broker like RabbitMQ (routing, DLQs, prioritization built in) or a lightweight Redis-backed queue library is a reasonable middle ground. If the scenario is genuinely high-throughput event streaming, with multiple consumers reading the same data at different speeds and a need to replay history, moving to a log-based system like Kafka makes sense, but it brings operational overhead that most applications simply do not need.

Watching queue depth

The deceptive part of queue-based systems is how long a problem can stay invisible. When worker throughput falls behind the rate work is produced, the queue grows silently, throws no errors, and the only symptom is that jobs slowly take longer to complete. The metric that matters is not queue length but the age of the oldest job waiting in it. Ten thousand queued jobs is not a problem if each clears in two seconds; fifty queued jobs is a real problem if the oldest one has been waiting six hours.

When to skip this approach

If the team is small and runs a single monolithic application, the operational cost of running a separate fleet of workers and a messaging stack often outweighs what it buys. A simple scheduled task running inside the same process, backed by the database, covers most "this should run in the background" needs without standing up a dedicated queue. Moving to a queue-based system when real scale or independent retry requirements actually appear costs less than building and maintaining one from the start.