All articles
EngineeringOctober 2026

How We Built AI Control's Durable Capture Job Queue

By Daniil Shastovsky·· 14 min read

The 'Check Now' Button That Promises More Than It Can Deliver

Picture a simple button in the interface: "Check what ChatGPT, Perplexity, and Gemini are saying about my brand." A user clicks it expecting to wait a second or two, the way a regular search feels. On paper, this looks like exactly one API call: send a prompt, get an answer, render it on screen.

That's roughly how we thought about it too, in the first version of AI Control's brand-visibility checks. The hard part, we assumed, was building good prompts and parsing messy answers. The real difficulty turned out to live somewhere else entirely.

AI Control doesn't make one request — it makes a batch of them. A typical check covers several prompts at once (we wrote about how to pick them in our piece on prompt monitoring), run against several real AI systems: ChatGPT, Perplexity, Copilot, Gemini, Google AI. Each one is a separate external service that doesn't answer instantly — the model has to generate a response, sometimes search for sources first, sometimes both. In our experience, a single call to a single system can take anywhere from a few seconds to a couple of minutes, depending on the system, the prompt, and how busy it happens to be at that moment.

Multiply that by a dozen prompts across five systems, and a single "check" can add up to minutes of total work — most of it the server doing nothing but waiting on someone else's response. That's exactly where the naive design — hold the connection open until every answer comes back — starts to fall apart.

And this isn't an abstract engineering nicety. The results of these checks later feed historical analytics: brand visibility over time, Share of Voice against competitors, the actual list of sources an AI system cites. If a check occasionally gets lost halfway through, that's not just a bad UX moment — it's a gap in the data you can't reconstruct after the fact.

Why 'Just Wait For It' Doesn't Hold Up

The naive version looks simple: the user clicks Check, the browser sends a request, the server synchronously queries every AI system — in sequence or in parallel — collects the answers, and returns them all at once. That works fine in a demo with one prompt. With several prompts across several providers, it breaks in three specific places.

  • A long-running request tied to a browser tab. Close the tab, lock the phone, lose Wi-Fi for a second, and the connection is gone. Whatever work had already been done disappears with it; the check has to start over from zero.
  • Job state living only in server memory. Even if you make it asynchronous and show a progress bar, if "what's happening with job #482" only exists in a running process's RAM, a server restart — a deploy, a scale-out event, a crash — silently wipes every job that was mid-flight. The user is left staring at a progress bar frozen at 40% forever.
  • No concurrency control. If every incoming prompt immediately fires its own call to an AI system, a handful of users starting checks at the same moment can easily generate dozens of parallel calls to the same external providers. That's precisely the kind of pattern providers themselves try to discourage: too many simultaneous calls raise the odds of slow or outright rejected answers, and not just for the user who triggered it.

Each of these is survivable on its own. Together, they mean a user will occasionally lose a check's results with no real explanation — just because they closed a tab, or because the backend happened to restart during a routine deploy.

We didn't learn this in theory — we learned it from the first support tickets: "I ran a check, closed the tab, came back, and there's nothing there," "the check has been stuck in the same place for ten minutes, what's going on." Every one of those was a symptom of one of the three causes above, not a separate problem that needed its own one-off patch.

What's telling is that none of these three causes were bugs in the usual sense — nothing a one-line fix would catch. They're the consequence of a single architectural assumption: we'd designed the system as if the unit of work were a request, not a job. A request only lives as long as its specific connection and the specific process serving it do. A job doesn't have to live that way — and that distinction is exactly where the rethink started.

The Fix, in Plain Terms: Write State Before You Act

Patching the naive version piecemeal — add a retry here, bump a timeout there — is tempting, but it doesn't address the real issue: the unit of work in the system was an HTTP session, and it should have been a database row.

We changed the order of operations around one deceptively small step: before any real work starts, the backend writes a row to the database — a job — with a status of "pending". Only after that does anything resembling a call to an AI system happen. That one change matters enormously: the unit of work now exists independently of whether the browser tab is still open, whether the specific server process that first handled the request is still alive, or whether the user is even connected to the internet anymore.

From there, a small, fixed pool of background workers — not separate servers, just processes inside the same backend — continuously polls the database for pending jobs. When a worker finds one, it atomically marks it as claimed (atomically, so two workers can't grab the same job at once), and only then goes off to make the actual external call. The number of workers running at once is deliberately capped: think a small fixed handful of jobs in flight at a time, not "however many requests happen to arrive".

If the backend restarts while some jobs are mid-flight, that's no longer a disaster: on startup, the process scans the jobs table, finds anything stuck in an in-progress status, and puts it back into the queue for a worker to pick up again. Finished jobs — successful or failed — get cleaned up after a while so the jobs table doesn't grow forever and become a bottleneck in its own right.

A useful analogy is an open kitchen. In the naive version, a single waiter keeps every order in their head and runs back to cook each dish personally — if that waiter goes home, the orders go with them. In the durable version, an order becomes a ticket, the ticket goes on the rail, and any free cook can pick it up — if one cook steps away, the ticket is still sitting there for the next one.

Here's a rough example: a user runs a check across 8 prompts on 4 AI systems — that's 32 individual jobs. Done strictly sequentially and synchronously, at an average response time of roughly 30–40 seconds per call, the whole check could stretch into 15–20 minutes of continuous HTTP connection. With a queue and several workers, that same total amount of work happens in parallel: some jobs finish while others are still running, and the user sees the first results almost immediately instead of waiting on the slowest prompt before seeing anything at all.

There's a real cost here worth stating plainly: any individual job starts actually running a fraction of a second later — sometimes a second or two later — than it would under a direct synchronous call, because a worker first has to find it in the database. For a check that's already going to take a minute or two, that difference is invisible. For a hypothetical case where the answer truly has to be instant, this approach would be overkill — but that's not the situation we're in.

The Architecture, Layer by Layer

Put together, the system has five layers — from the incoming request down to the actual call to an external AI system.

API endpointaccepts the check request and responds immediately, without waiting for a result
Database writea job row is created with status "pending" before any real work begins
Bounded queue and worker poolbuffers jobs waiting for one of a small, fixed number of background workers
External AI systemthe real call to ChatGPT, Perplexity, Copilot, Gemini, or Google AI

Layers one and two are the only part that has to run synchronously and fast: accept the request, write a row. Everything after that — the queue, the workers, the actual call to the AI system — no longer blocks the user or the HTTP request that kicked things off.

One Job's Life Story

Zoom in from the architecture to a single job — say, "check prompt X against Perplexity" — and it moves through a handful of states.

Queued
Claimed by worker
Running
Done
Failed → retried

"Done" and "Failed" aren't always final. If a call to an AI system fails for what looks like a transient reason — the external service returning something that reads as temporary overload rather than "this prompt can never be processed" — the job goes back to "queued" with its retry counter incremented, instead of being marked dead on the spot. Retries aren't unlimited, though: after a handful of failed attempts, a job is marked as permanently failed, and that shows up in the interface rather than disappearing silently.

What This Looks Like in Code

To keep this from staying abstract, here's simplified pseudocode — not real AI Control production code, but a teaching version of the same logic: write the job, have workers poll for it, claim it atomically, do the work, record the outcome.

Pseudocode: writing a job and the worker loop
# simplified pseudocode, not real production code

def enqueue_check(prompt, provider):
    job = db.insert("capture_jobs", {
        "status": "pending",
        "prompt": prompt,
        "provider": provider,
        "attempts": 0,
        "created_at": now(),
    })
    return job.id

def worker_loop():
    while True:
        job = db.claim_one(
            "UPDATE capture_jobs SET status='claimed', claimed_at=now() "
            "WHERE id = (SELECT id FROM capture_jobs WHERE status='pending' "
            "ORDER BY created_at LIMIT 1 FOR UPDATE SKIP LOCKED) "
            "RETURNING *"
        )
        if job is None:
            sleep(1)
            continue

        try:
            db.update(job.id, status="running")
            result = call_ai_system(job.provider, job.prompt)
            db.update(job.id, status="done", result=result)
        except TransientError:
            db.update(job.id, status="pending", attempts=job.attempts + 1)
        except FatalError as e:
            db.update(job.id, status="failed", error=str(e))

def on_startup():
    # anything still "running" after a restart counts as interrupted
    db.execute(
        "UPDATE capture_jobs SET status='pending' WHERE status='running'"
    )

Two details matter more than they might look. First, claiming a job is one atomic SQL statement, not a "read, then write" pair of steps — that's what stops two workers from grabbing the same job in the rare but inevitable race. Second, on_startup is the entire recovery story: nothing beyond "find anything stuck and requeue it" is actually required.

Notice what's missing from this pseudocode: no "remember the job in an in-memory list," no global counter of active requests, no logic along the lines of "if the server restarted, try to reconstruct state from the last log line." All the durability comes from treating one table row as the single source of truth, rather than anything that only lives inside the process.

There was also a side benefit we didn't plan for but noticed after the fact: a fixed worker pool pulled all the error-handling logic for external systems into one place. Previously, code for "what do we do if an AI system comes back with an overload error" could have ended up scattered across every call site that happened to reach out to an external API. Now it lives in exactly one worker loop and behaves the same regardless of which screen or button kicked off the check — which also means there's exactly one place to update it, not a codebase-wide search.

Naive vs Durable: What Happens at the Edges

The difference between "wait synchronously for the answer" and "write a job and process it in the background" sounds like an internal implementation detail. In practice, it's exactly the edge cases where users notice it — and exactly where the naive version fails most often.

ScenarioNaive synchronous callDurable job queue
User closes the tabConnection drops, the answer is lost, the check has to restart from scratchThe job is already saved in the database — work continues in the background, and progress is visible on the next visit
Server restarts mid-checkWhatever the server was holding in memory disappears without a traceOn startup, the backend finds unfinished jobs and puts them back in the queue
Two similar requests arrive close togetherBoth fire duplicate calls to the AI systems — wasted cost and a race conditionA job-level idempotency check recognizes the duplicate and avoids doing the work twice
The external AI system doesn't respond in timeThe whole HTTP chain hangs along with the user's requestThe job is marked as timed out or retried, while the rest of the queue keeps moving

None of these scenarios are exotic — they're what happens every day under any reasonably active usage: someone checks brand visibility from a phone on the subway where the connection drops, a routine deploy goes out, someone double-clicks Check before the interface has had a chance to react. A durable architecture doesn't eliminate these situations — it makes them boring, predictable, and not something anyone has to fix by hand.

Questions We Actually Get About This Design

Why not just use an off-the-shelf queue like Celery or RabbitMQ?

At our volume, a separate message broker would have meant one more piece of infrastructure to deploy, monitor, and fix separately from everything else. The database we already use for everything else can reliably store and atomically update rows, and that's enough to implement a queue without adding a new service. The tradeoff is extra load on the database at very high job volumes — at our scale, that's a reasonable, deliberate trade, not a free lunch.

What happens if a worker itself crashes mid-job?

The job is left sitting in a running status in the database rather than vanishing or hanging in limbo. A check at startup, or on a timer, finds jobs that have been stuck in that status too long and puts them back in the queue, the same way it would after a full server restart. The one extra requirement this creates for the processing code is that it has to be safe to repeat, or a retry could double up the result or double-charge usage.

How does the user even know a check is progressing if it's all happening in the background?

The interface periodically asks the backend for job status and renders progress based on what's in the database, not on a live connection. That's exactly what makes it safe to close the tab and come back later from a different device — progress is reconstructed from the same source of truth the workers themselves use, not from anything that only lived in the browser.

Wouldn't it be simpler to just allow more parallel requests to the AI systems?

That would feel faster for any one user, but under real load, several people running checks at the same time would quickly hit the external providers' own limits — answers would start coming back slower or with errors, for everyone, not just the user who triggered the spike. A small, fixed worker pool is a deliberate tradeoff between the speed of a single check and the stability of the service as a whole.

Checklist: Building Your Own Durable Job Queue

None of this is specific to checking AI systems — it's a general pattern for anything that does something slow and external: sending emails, generating reports, calling someone else's API, processing an uploaded file. If you're building something similar — even if a good chunk of the code is written with AI assistance, the way we described in our piece on vibecoding for SEO specialists — here's where to start.

  1. Write job state before you act on it, not after — the database should be the source of truth, not a running process's memory.
  2. Cap concurrency on purpose. A small, fixed worker pool protects both your own service and the external systems it talks to.
  3. Design for restart-safety from day one, not after the first incident — on startup, look for jobs stuck in an in-between status and put them back in the queue.
  4. Make retries idempotent. Processing the same job twice should never produce a doubled result or a double charge.
  5. Clean up finished jobs. Otherwise the queue table grows forever and becomes a bottleneck of its own.
  6. Make progress observable from outside the process. If a user or a dashboard can't check status without reaching into server memory, recovering from a failure will always be painful.
Related reading

New to the idea of checking what AI says about a brand? Start with our guide to AI visibility monitoring, and for how the prompts themselves get collected, see our piece on prompt monitoring.

Want to check this in your market?

AI Control regularly collects AI responses, brand positions, competitors and cited sources for your prompt library.

Explore AI Control

Is your own content set up for stories like these?

Use 50 welcome credits for an AI Readiness check — retrieval, extractability, schema.org signals, and a prioritized rewrite brief, scored the way an AI assistant actually reads your page.

Get 50 credits

Don't just read about AI search. Check your own pages against it.

The same AEO/GEO signals covered above — schema, retrieval, direct answers, citations — are exactly what AI Readiness scores on any page you give it. New accounts receive 50 shared credits.