The Request Ended. The Work Did Not.

The order exists. Payment succeeded. The database transaction committed.
Now the API waits three more seconds for an email provider to accept the confirmation message.
Sometimes it waits ten. Sometimes the provider times out after doing the work, so the backend does not know whether the email was accepted. The user sees a failed checkout and presses the button again, even though the expensive part already succeeded.
The email was not part of the request's answer. We made it part of the request's lifetime.
That is the boundary background jobs are meant to fix. The API completes the durable work needed to answer the user, records what still has to happen, and returns. A worker continues separately.
Simple in one sentence. Much less simple once "separately" has to remain reliable.
What belongs outside the request
Good background-job candidates share a few traits:
- the work is slow or has unpredictable latency;
- the user does not need its final result before continuing;
- it calls an external service that may be unavailable;
- it can be retried safely;
- it may need independent scaling.
Confirmation emails fit. So do image processing, CSV exports, webhook delivery, search indexing, analytics enrichment, and invoice generation.
Creating the order itself does not fit. The API should not tell the browser "order created" before the order is durably recorded. Queueing the core write would turn a clear result into "we accepted your wish and may create it later."
Sometimes that asynchronous product is intentional. Then the API should say so with a job or operation id and a pending status. The important part is not pretending pending work is complete.
A queue separates acceptance from execution
The request path becomes:
browser -> create order in PostgreSQL -> enqueue confirmation job -> return order worker -> receive job -> send email -> mark job complete
The queue buffers work between producers and consumers. The API can accept jobs quickly. Workers process them at the speed the downstream system allows. If email slows down, checkout does not have to slow down with it.
That buffer also lets workers scale independently. Ten image processors can consume from the same queue during a spike, then scale back later.
But a queue is not just an array somewhere on the server. It needs durable messages, acknowledgement, retry behavior, visibility into pending work, and a strategy for jobs that never succeed.
The database-to-queue gap
Here is a subtle failure:
const order = await database.orders.create(input); await queue.publish('order-confirmation', { orderId: order.id }); return order;
The database write succeeds. The process crashes before publishing the message. The order exists, but no confirmation job was recorded.
Reversing the order does not help:
await queue.publish('order-confirmation', { orderId }); await database.orders.create(input);
Now the worker may receive a job for an order that does not exist because the database write failed.
Two reliable systems do not automatically form one atomic operation.
The transactional outbox pattern solves this by writing the order and an outbox record in the same database transaction:
BEGIN; INSERT INTO orders (id, customer_id, status) VALUES ($1, $2, 'confirmed'); INSERT INTO outbox_events (id, type, payload, created_at) VALUES ($3, 'order.confirmed', $4, now()); COMMIT;
A separate publisher reads unsent outbox rows, publishes them to the queue, and marks them as sent. If it crashes after publishing but before marking the row, it may publish the same event again.
That is acceptable only if the consumer handles duplicates.
At-least-once means duplicates are normal
Most practical queue systems give us at-least-once delivery: a job will be delivered, but under failure it may be delivered more than once.
The worker receives a job, sends the email, then crashes before acknowledging completion. The queue cannot know the email was sent, so it delivers the job again.
Exactly-once delivery is an attractive phrase. End-to-end exactly-once effects across a queue, worker, database, and external email provider are much harder than the phrase suggests.
Design the job to be idempotent instead.
Give the effect a stable identity:
type SendOrderConfirmationJob = { jobId: string; orderId: string; templateVersion: number; };
Before sending, record or check whether that specific confirmation has already been completed. If the email provider supports an idempotency key, send the stable job id there too. If it does not, decide whether an occasional duplicate email is acceptable or whether stronger local coordination is worth the complexity.
Idempotency is not "ignore every repeated order id forever." A resend requested by support is a new intent and should have a new job identity.
A retry is the same intent attempted again. A new action deserves a new identity.
That is the same rule we used for checkout requests. Queues make it impossible to avoid thinking about it.
Retry only failures that may change
A timeout from an email provider may succeed on the next attempt. An invalid recipient address will not become valid because we tried it twelve times.
Classify failures:
- Transient: timeout, connection reset, provider rate limit, temporary outage.
- Permanent: invalid input, unsupported format, missing required data.
- Unknown: the remote side may have completed the operation before the connection failed.
Transient failures get retries with exponential backoff and jitter. Permanent failures stop immediately and become visible. Unknown outcomes need idempotency or reconciliation before retrying blindly.
const delayMs = Math.min(60_000, 1000 * 2 ** attempt); const jitterMs = Math.floor(Math.random() * 500); scheduleRetry(job, delayMs + jitterMs);
Backoff gives the dependency time to recover. Jitter prevents thousands of failed jobs from retrying in perfect synchronization and creating another outage the moment the provider returns.
A maximum attempt count is not enough by itself. A job that gives up needs somewhere to go and someone to notice.
Dead letters are unfinished product work
After repeated failure, a job may move to a dead-letter queue. That name sounds final. It should not mean invisible.
The job still represents a promise: send a confirmation, generate an export, deliver a webhook. Moving it aside protects healthy work from one poison message, but the product obligation remains.
For each job type, decide:
- who is alerted;
- what context is safe to inspect;
- whether the job can be replayed;
- whether the user should see a failed status;
- when the underlying operation needs manual recovery.
A dashboard that proudly shows "queue healthy" while 4,000 jobs sit dead is measuring the machinery instead of the outcome.
Queue depth, oldest-job age, processing duration, retry count, and terminal failures tell a more useful story than worker CPU alone.
Job payloads should be small references
It is tempting to put the complete order, customer, and rendered email content into the message. Then the worker has everything it needs.
It also freezes a large data snapshot inside the queue, exposes more sensitive information, and makes schema changes harder. A delayed job may execute days later with a payload produced by old code.
Prefer stable identifiers and intentional snapshots:
type GenerateInvoiceJob = { jobId: string; orderId: string; invoiceVersion: number; };
The worker loads current durable data by orderId. If a value must reflect the moment of purchase, such as the charged price, that snapshot already belongs in the order model. The message does not need to invent another source of truth.
There are exceptions. A webhook event may intentionally preserve the event-time payload. The principle is to choose snapshot semantics, not accidentally serialize half the database into every job.
Workers need deployment compatibility too
A queued message can outlive the code that produced it.
Deploy version two of a worker while version-one messages are still waiting, and the new worker must understand the old payload. Roll back the producer, and old code may start publishing the previous shape again.
This is the migration problem from the previous article, now applied to messages.
Use additive payload changes, sensible defaults, and explicit versions when semantics change. Do not rename a required field and assume the queue emptied during deployment. Measure that it did, or support both versions.
Long-running jobs also need graceful shutdown. When a worker receives a termination signal during deployment, it should stop accepting new work, finish or safely release the current job, acknowledge only after completion, and exit within the platform's deadline.
Otherwise every deployment creates retries by killing work halfway through.
The frontend needs a model for pending work
Some background work is invisible to the user. An analytics enrichment job does not need a progress bar.
An export does.
The API can return 202 Accepted with an operation id:
{ "operationId": "export_7f2a", "status": "queued" }
The frontend polls an operation endpoint, subscribes through server-sent events or WebSockets, or lets the user leave and returns the finished export through notifications.
The important state is not only loading and done. It may be queued, running, retrying, completed, failed, or cancelled. Those states belong to the product contract even if a queue implementation uses different internal terminology.
Do not show an eternal spinner because the worker died. Persist operation state somewhere durable, define a timeout or terminal failure, and give the user a next action.
Async architecture becomes good UX only when uncertainty is visible without becoming alarming.
Backpressure is a product decision
If requests create work faster than workers can finish it, queue depth grows. The queue is doing its job, but the delay may eventually violate the product promise.
We can add workers until a downstream limit stops us. We can prioritize urgent jobs. We can reject or postpone new work. We can reduce expensive processing. What we cannot do is treat an ever-growing queue as healthy because no messages were lost.
For confirmation email, five minutes may be embarrassing but survivable. For a fraud check that gates order fulfilment, five minutes may block the business. "How old can the oldest job become?" is a product question with an operational metric attached.
Queues buy time. They do not create capacity.
Keep the promise smaller than the machinery
Background jobs are useful because they let an HTTP request end before every side effect is complete. That separation improves latency and isolates unreliable dependencies.
But the queue is not the guarantee. The guarantee comes from the whole path:
- durable intent recorded with the business change;
- messages published despite process failure;
- consumers safe under duplicate delivery;
- retries limited to failures that may recover;
- dead jobs visible and recoverable;
- payloads compatible across deployments;
- product states that admit work is still pending.
Once those pieces are explicit, a queue stops being mysterious infrastructure. It becomes a controlled handoff between "the user can continue" and "the system still owes some work."
The next part of this guide moves back to the public edge: domains, DNS, HTTPS, and reverse proxies, and how a name typed into a browser reaches the right application without exposing every internal service directly.
If your application returns success before all the work is finished, write down what it still owes. That list is the beginning of the background-job design, not an implementation detail after it.
More than a blog post
I share frontend news and the reasoning behind it throughout the day. Pick the language that feels natural to you.