Lesson 4: Background Work
When a request must return before the work finishes
Some work outlives a reasonable request. Extracting from a 200-page document, re-embedding a corpus after a model change, or generating a report over ten thousand records all take minutes to hours.
Three signs the work belongs in the background: it takes longer than a client will wait, which is roughly thirty seconds; it should survive the client disconnecting; or it should be retried independently of the request that started it.
FastAPI's built-in background tasks are not the answer for these. They run in the same process after the response is sent, so they die with a deployment, cannot be retried, cannot be monitored, and have no status. They are fine for a fire-and-forget side effect such as sending a notification, and wrong for work whose completion matters.
Task queues, status endpoints, and callbacks
The pattern has three parts: submit and return an identifier, expose status, and optionally notify on completion.
@router.post("/v1/jobs/extract-corpus", status_code=status.HTTP_202_ACCEPTED)
async def submit_corpus_extraction(
body: CorpusExtractionRequest,
deps: Dependencies = Depends(get_deps),
user: AuthenticatedUser = Depends(require_user),
) -> JobAccepted:
"""Queue a long-running extraction and return immediately."""
job = Job(
id=str(uuid.uuid4()),
kind="extract_corpus",
payload=body.model_dump(),
user_id=user.id,
status="queued",
idempotency_key=body.idempotency_key,
)
existing = await deps.jobs.find_by_idempotency_key(job.idempotency_key)
if existing is not None:
return JobAccepted(job_id=existing.id, status=existing.status)
await deps.jobs.enqueue(job)
return JobAccepted(job_id=job.id, status="queued")
202 Accepted, not 200, because the work has not happened yet and the status code should say so.
The idempotency key is Module 5's mechanism applied to job submission. A client retrying a submission that succeeded but whose response was lost must not start a second job.
Status endpoint:
@router.get("/v1/jobs/{job_id}")
async def get_job(job_id: str, deps=Depends(get_deps), user=Depends(require_user)) -> JobStatus:
job = await deps.jobs.get(job_id)
if job is None or job.user_id != user.id:
raise HTTPException(status_code=404, detail="job not found")
return JobStatus(
job_id=job.id,
status=job.status, # queued, running, succeeded, failed
progress=job.progress, # 0.0 to 1.0 where known
result_url=job.result_url,
error=job.error,
cost_usd=str(job.cost_usd) if job.cost_usd else None,
)
Note the ownership check combined into the 404 rather than a 403, so a user cannot learn that someone else's job exists.
Webhook callbacks notify a client rather than making them poll.
async def notify_completion(job: Job, deps: Dependencies) -> None:
"""Call the client's webhook, signed and with retries."""
if not job.callback_url:
return
payload = {"job_id": job.id, "status": job.status, "completed_at": _now_iso()}
body = json.dumps(payload)
signature = hmac.new(deps.webhook_secret, body.encode(), hashlib.sha256).hexdigest()
await post_with_retry(
job.callback_url,
content=body,
headers={"content-type": "application/json", "x-signature": signature},
)
Three requirements. Sign the payload so the receiver can verify it came from you. Retry with backoff, since the receiver may be briefly down, using Module 5's decorator. And keep the payload minimal, carrying an identifier rather than the result, so the client fetches it over an authenticated channel rather than receiving data over a URL they gave you.
The worker. Whether you use a dedicated queue system or a database-backed queue, the requirements are the same and they are all things earlier modules built: idempotent processing so a redelivered message is harmless, checkpointing from Module 4 so a long job resumes, bounded concurrency from Module 8 so workers do not exhaust the provider, cost attribution from Module 10 so job spend is visible, and graceful shutdown so a deployment does not lose in-progress work.