Scaling Agents — Concurrency, Queues, Fan-out, Rate Limits & Cost Budgets
Scaling agents is a distributed-systems problem: horizontal workers, job queues, per-tenant quotas, backpressure, and idempotent tool side effects. Every parallel tool call is a budgeted fork-join. Fan-out multiplies capability and cost together.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why is an agent run a job, not a request thread?
Answer
It can outlive one HTTP call, it needs a checkpoint, and it must survive a worker restart.
L2
What does the queue buy you?
Answer
A buffer, fair scheduling by tenant, and a place to apply backpressure when lag exceeds the SLO.
L3
Session affinity or work stealing?
Answer
Affinity is simple and creates hot workers. Stealing needs a durable checkpoint so any worker can resume.
L4
What is a safe fan-out?
Answer
Independent reads with a cap and a partial-failure policy. Money writes stay serialized or idempotent.
L5
How do you stop a 429 storm?
Answer
A shared limiter, per-tenant quotas, jittered backoff, and shedding low-priority jobs before the provider sheds you.
L6
What is the idempotency key?
Answer
A stable id for one side effect, often the run id joined with the step id. A retry returns the first receipt.
L7
How do you state a cost SLO?
Answer
Dollars per successful task, a kill switch at max dollars per run, and a monthly tenant budget. Token counts are an input, not the SLO.
Failure modes
Retry doubles a write
The worker crashes after the tool succeeds and runs the same step again without an idempotency key.
Fan-out of writes
Parallel refunds or emails partially commit and cannot be joined safely.
Noisy tenant
One customer fills the worker pool and every other tenant waits.
429 retry storm
Every worker retries immediately and the provider stays saturated.
Hot affinity
One worker owns a long run and the pool looks idle while that run stalls.
Misconceptions
More agents means more scale.
Extra agents multiply tokens and merge conflicts. Scale the worker pool and cap fan-out.
The provider's rate limit is your quota system.
You need your own per-tenant buckets so one customer cannot spend the shared pool.
Token count is the cost SLO.
The SLO is dollars per successful task. Tokens are how you estimate it.
Interviewer traps
Drawing a Kubernetes diagram and skipping idempotency.
Name the queue, the quota, the limiter, and the key on the write. Placement is secondary.
Treating saga orchestration as this lesson.
Distributed-transaction sagas are their own cluster. Here, orchestration means the agent worker and its fan-out.
Design scenario
Same prompt for every reader.
Requirements
Per-tenant fairness, a dollar cap per run, idempotent refunds, and a degrade path when queue lag exceeds the SLO.
Traffic / scale
About 5 agent starts per second per tenant, with bursts when a campaign lands.
Latency
Queue wait for an interactive acknowledgement stays under 2 seconds at p95. Deep tool runs may take longer behind the acknowledgement.
Consistency
A refund with the same idempotency key applies once. Read fan-out may return a partial set.
Availability
If the provider returns 429, jobs wait or shed. They do not stampede. A worker crash resumes from the checkpoint.
Failure assumptions
- Workers die after a tool succeeds and before the checkpoint.
- One tenant can submit a burst.
- Parallel reads can fail independently.
Constraints
- Writes are not fanned out without a coordinator.
- Each tenant has a quota separate from the global pool.
- A run stops at a max dollar amount.
Prompt
Scale a support agent to about 2000 concurrent sessions without melting the model provider.
API
What is enqueued, and what does the client get before the run finishes?
Data
Where do the checkpoint, the idempotency key, and the accrued cost live?
Architecture
Where do the quota check, the worker pool, and the provider limiter sit?
Two thousand support sessions arrive together
Prefer
A queue, a quota, and a capped fan-out
The HTTP handler enqueues and returns an acknowledgement. Workers pull fairly. Parallel reads have a join policy. Writes stay single-threaded per run.
- One tenant cannot fill the pool.
- A 429 waits in the queue instead of multiplying.
- A crashed worker resumes from a checkpoint.
- The dollar meter can stop the run.
Alternative
A thread per session, tools called in parallel, retries immediate
It works on a laptop. Under a campaign spike the provider locks you out and a retried refund pays twice.
- No fairness between tenants.
- Fan-out width is whatever the model asked for.
- Crash recovery depends on the process still being alive.
Overview
An agent run is a job. It can take many model calls, it must survive a process restart, and its cost is variable. Scale it the way you scale any job system, then add two agent-specific rules: cap fan-out, and make tool side effects idempotent.
Flow
- 1
1. API accepts goal
- next2. Enqueue the run
- 2
2. Enqueue the run
- next3. Tenant quota
- 3
3. Tenant quota
- next4. Worker takes job
- 4
4. Worker takes job
- next5. Model and tools
- 5
5. Model and tools
- next6. Limit and breaker
- 6
6. Limit and breaker
- next7. Checkpoint state
- 7
7. Checkpoint state
Lesson map
Scaling Agents — Concurrency, Queues, Fan-out, Rate Limits & Cost Budgets
Scaling agents is a distributed-systems problem: horizontal workers, job queues, per-tenant quotas, backpressure, and idempotent tool side effects. Every parallel tool call is a budgeted fork-join. Fan-out multiplies capability and cost together.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB w["Worker"] a["Tool A"] b["Tool B"] c["Tool C"] w -->|Read order| a w -->|Read user| b w -->|Read policy| c a -->|Result or error| w b -->|Result or error| w c -->|Result or error| w
The acknowledgement to the user can leave at step 2. The deep work happens on the worker. When the limiter is open, the worker checkpoints and waits. It does not tight-loop the provider.
| Component | Role | Scale knob |
|---|---|---|
| Queue | Buffer jobs and schedule fairly | Partitions or a fairness key by tenant |
| Workers | Run the loop | Autoscale on lag, not on CPU alone |
| Rate limiter | Protect providers and tools | Token bucket per key |
| State store | Checkpoint the run | Sharded by run id |
| Cost meter | Accrue dollars and tokens | Kill switch |
Concurrency inside the pool
| Pattern | Meaning | Risk |
|---|---|---|
| Session affinity | One worker owns a run | Hot workers, poor failover |
| Work stealing | Any worker resumes a checkpoint | Needs durable state and a fence against two workers |
| Fan-out of tools | Parallel independent reads | Cost spike and partial failure |
| Multi-agent fan-out | Parallel specialists | Merge conflicts and duplicate side effects |
Affinity is fine for a sticky prototype. At a few thousand sessions, prefer a checkpoint any worker can load, with a lease so two workers do not execute the same step. That lease is the same idea as any job system. Do not invent a second scheduler inside the model.
Fan-out and fan-in
Sequence
- 1
Worker → Tool A
Read order
- 2
Worker → Tool B
Read user
- 3
Worker → Tool C
Read policy
- 4
Tool A → Worker
Result or error
- 5
Tool B → Worker
Result or error
- 6
Tool C → Worker
Result or error
Three reads, one join. The join policy is part of the design:
- Fail-fast. Any error aborts the step. Use this when the answer is wrong without all three.
- Best-effort. Return a degraded answer and say which read failed. Use this for a summary, never for a payment.
Money tools do not fan out. Serialize writes, or give each write its own idempotency key and a coordinator. Partial commits are how you refund twice and email once.
Cap N. “The model asked for 40 browser calls” is not a capacity plan. Map-reduce research, a critic pass, and a specialist router are valid patterns only with a width limit and a merge rule. Do not spawn agents to hide a missing product requirement.
Quotas, backpressure, and cost
| Lever | Example | Effect |
|---|---|---|
| Per-tenant start rate | 5 agent starts per second | Fairness |
| Max in-flight runs | 50 per tenant | Tail latency |
| Max steps per run | 20 | Loop bound |
| Max dollars per run | 0.50 USD | Cost SLO |
| Provider RPM and TPM | Shared pool with priority | Avoid 429 storms |
When queue lag exceeds the SLO, shed low-priority jobs or switch to the answer-only degrade path from the hub checklist. Retrying harder is how you stay in 429.
Backoff uses jitter. A shared circuit breaker, described as a concept by Fowler, opens when the error rate trips. The reliability lesson owns the breaker implementation sketch. This page owns the decision to stop admitting work.
Idempotent side effects
Retries happen. The tool, or the adapter in front of it, must treat a repeated key as the same effect. A practical key joins the run id and the step id.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The second call does not add another 40. It returns the first receipt. That is the whole interview point.
A token bucket
export class TokenBucket {
tokens: number;
constructor(
private capacity: number,
private refillPerSec: number,
private last = Date.now(),
) {
this.tokens = capacity;
}
tryTake(n = 1): boolean {
const now = Date.now();
const elapsed = (now - this.last) / 1000;
this.last = now;
this.tokens = Math.min(this.capacity, this.tokens + elapsed * this.refillPerSec);
if (this.tokens < n) return false;
this.tokens -= n;
return true;
}
}Press Run. Snippets must be self-contained — no network, files, or native modules.
Use one bucket for the provider pool and another per tenant. The tenant bucket is what stops a noisy neighbor. The global bucket is what stops you from exceeding the contract.
Pitfalls
- Autoscaling workers because CPU is low while the provider is already returning 429. The scarce resource is the limit, not the CPU.
- Fan-out of writes with “we will reconcile later” and no ledger.
- A single shared API key and no per-tenant quota.
- Measuring success in tokens. Finance asked for dollars per resolved ticket.
- Calling this page a saga tutorial. Multi-step business undo has its own distributed-systems cluster. Here you only need idempotent tool effects and a checkpoint.
A campaign submits 100 starts in a second for one tenant. Walk what the quota does, what the queue lag does, and what a worker does when the provider returns 429 on the third model call. Say whether a refund step that already succeeded is safe to retry.
Interview Q&A
Workers scale, but tools return 429. What do you change first?
Answer
Global and per-tenant rate limits, queue backpressure, exponential backoff with jitter, and a circuit breaker on the unhealthy tool. Cache hot reads if the caching lesson says they are safe. Do not add workers. More workers make the storm worse.
Why is fan-out of writes dangerous?
Answer
Duplicate side effects and partial commits. One refund lands, the email fails, the retry refunds again. Serialize writes or use an idempotent ledger. Reads can fan out. Money cannot.
How do you express a cost SLO?
Answer
Dollars per successful task, a kill switch at max dollars per run, and a monthly budget per tenant. Token counts feed the estimate. They are not the number you page on.
Affinity or work stealing?
Answer
Stealing, once you have a checkpoint and a lease. Affinity is easier and pins a long run to a hot worker. Two workers executing the same step without a lease is worse than affinity, so the lease is part of the answer.
What is a partial-failure policy?
Answer
The rule for a join. Fail-fast aborts when any read fails. Best-effort returns a degraded answer and names the miss. Pick one before you parallelize. Do not discover it in production when policy lookup timed out and the model guessed.
What does the idempotency key look like?
Answer
Something stable across retries of the same step. A common shape joins the run id and the step id. The tool stores the first successful result under that key and returns it on the duplicate. A random UUID per attempt is not idempotency.
When do you shed load?
Answer
When queue lag breaks the acknowledgement SLO, or when the dollar meter says the tenant is over budget. Shed low-priority runs first. The hub’s degrade mode is the user-visible form: answer without tools, or queue a human.
Why not spawn a dozen specialist agents?
Answer
Each one is another model bill and another writer. Use a specialist only when the tool allowlists are disjoint and a router can mis-route safely. A critic pass is one extra call with a cap, not a swarm. Missing product rules are not fixed by more agents.
Go Deeper
- Stripe idempotent requests for the key semantics you want on refund tools.
- OpenAI rate limits and Anthropic rate limits for RPM and TPM. Your tenant buckets sit in front of those.
- Amazon SQS developer guide for the queue shape. Fairness is still your partition key.
- Stripe rate limiters and CircuitBreaker for the limiter and the open state.