A practical guide to queues and workers for long-running AI tasks
Learn when to use queues and workers for AI tasks, how to model state, idempotency, retries, concurrency, and cancellation, and when to keep a direct request.
A queue does not make AI smarter or faster. It turns a long-running execution into work that can be identified, observed, and recovered.
An AI integration often starts inside an HTTP request: the frontend sends data, the API calls the model, and the response comes back. For a short and predictable task, this path is easy to understand, test, and operate.
The problem appears when the work takes a long time, involves several files, depends on external tools, or fails transiently. The connection may time out, the user may close the page, and a retry may repeat an expensive operation. In that situation, separating request intake from processing starts to make sense.
But a queue is not a scale button. It introduces components, state, and new failure modes. The useful question is not “does it use AI, so does it need a worker?” but “does this task need to keep existing and remain recoverable independently of the request that started it?”
When a direct request is still enough
- The duration fits comfortably within the time expected by the interface and infrastructure.
- The operation has few steps and does not need to survive the page being closed.
- A failure can be shown and retried without a relevant side effect.
- Volume and concurrency are low or predictable.
- There is no partial result, resumption, or cancellation state that needs to be preserved.
- The operational complexity of a queue would be greater than the problem.
Before adding a queue, I would measure the real workflow. If most requests finish quickly and reliably, a well-defined timeout, useful feedback, and error handling may already solve the problem.
Signals that asynchronous processing may help
- The task frequently exceeds the comfortable duration of a request.
- Processing has several stages or a batch of independent items.
- Transient failures need controlled retries.
- The application needs to limit concurrency, cost, or usage per user.
- The user should be able to leave the page and return later.
- Partial results, resumption, cancellation, or reprocessing have value.
- The system needs to absorb spikes without sending everything to the provider at once.
The HTTP response accepts the work; it does not promise the result
When processing will continue in the background, the API can return `202 Accepted`, an identifier, and a monitoring URL. This status says the request was accepted, but it may still fail, be canceled, or not have started yet. The interface contract needs to reflect that distinction.
1{2 "jobId": "job_123",3 "status": "queued",4 "statusUrl": "/jobs/job_123"5}The minimum job workflow
- The frontend submits the request.
- The API validates authentication, authorization, format, and limits.
- The application creates a job with an identifier and `queued` state.
- A small message references that job in the queue.
- A worker reserves the message and changes the state to `running`.
- The worker performs the work, validates it, and stores the result.
- The job ends as `succeeded`, `failed`, or `canceled`.
- The interface polls or receives updates from the persisted state.
States need to represent facts
- `queued`: accepted and waiting for capacity.
- `running`: reserved by a worker and being processed.
- `succeeded`: validated result is available.
- `failed`: finished without an acceptable result.
- `cancel_requested`: cancellation requested but not yet confirmed.
- `canceled`: processing safely stopped or skipped.
In a batch, the overall state does not need to erase the state of each item. A job can finish with partial errors and still preserve valid results. Transitions also need control: a completed task should not silently return to `running`, and a retry should leave a history.
Idempotency comes before retries
Delivery guarantees depend on the broker and its configuration. Common queues may deliver the same message again after failures or delayed acknowledgements. Consumers therefore need to follow the actual contract, and when duplicates are possible, repeating the same intent must not create duplicate charges, writes, or actions.
- Use an idempotency key when creating the job.
- Check state before starting or repeating a stage.
- Use an expiring lease or lock when needed.
- Store results per stage to enable resumption.
- Protect external side effects from repetition.
- Retry only recoverable failures, with a limit and backoff.
A retry without idempotency turns a transient failure into a risk of duplicated work and cost.
Concurrency and cost still exist
A queue absorbs a spike, but it does not remove the work. If ten thousand tasks arrive in a few minutes, they still need to be processed. When the average producer rate exceeds consumer capacity, the queue grows and waiting time increases.
- Provider and internal service limits.
- Budgets per period, user, or feature.
- Concurrent job count and batch sizes.
- Queue depth and age of the oldest job.
- Maximum duration, expiration, and cancellation policies.
- Error rate, retries, and cost per completed task.
The interface cannot hide processing
Replacing a direct request with a queue changes the experience contract. The interface needs to show that the request was received, whether it is waiting or running, whether verifiable progress exists, whether the user can leave and return, how to cancel, and what happened when part of the work failed.
Polling, server-sent events, or notifications are ways to transport updates. None of them replaces the persisted source of truth. Cancellation also needs to be honest: changing the interface to “canceled” does not automatically stop an external request that is already running.
Example: analyzing several documents
Imagine a task that receives ten documents and extracts structured data. The API validates types, sizes, permissions, and quantity, creates an idempotent job, and records ten pending items. The queue message carries only the necessary identifier, avoiding copies of documents or sensitive data.
The worker reserves a few items according to the concurrency limit and records every attempt. Output goes through schema validation. An outage may trigger a retry with backoff; an invalid file produces a permanent, readable error. The interface shows `7 of 10 completed`, preserves those seven results, and allows only corrected items to be retried.
Observability and security
- How many jobs are waiting, and for how long?
- Which stage concentrates failures and retries?
- How much does a completed task cost?
- Which jobs became stuck in `running`?
- Is a source, user, or payload causing saturation?
- Did cancellation actually prevent new stages?
Logs should use correlation identifiers and avoid full prompts, documents, or responses by default. The worker may need to revalidate authorization because being in a queue does not turn an earlier permission into permanent access. Jobs, results, and logs also need retention that matches recovery needs and data sensitivity.
Decision checklist
- Does the task really exceed the comfortable limit of a request?
- Does it need to continue after the user leaves the page?
- Do retries, partial results, or resumption create value?
- Is there a unit of work with defined states?
- Can the operation be repeated safely?
- Do concurrency and budget have explicit limits?
- Can the interface track, cancel, and recover the task?
- Do logs and metrics cover the queue and worker?
- Have data in payloads, results, and logs been minimized?
- Does the benefit justify the operational complexity?
Limits and conclusion
A queue does not guarantee exactly-once processing, does not fix missing idempotency, and does not solve an unstable integration by itself. Some workflows can use asynchronous processing without a dedicated broker; others need more advanced orchestration. Architecture should grow with observed requirements and failures.
When state, idempotency, limits, and the tracking experience are not defined, adding a broker only moves the problem. When those elements are necessary, a queue and worker create a valuable separation between receiving intent and completing work safely.
