Docs / Guides

Concurrency limits and queues

How Tend caps the number of simultaneously running jobs per project, how excess jobs wait in queues, and how to tune throughput safely.

Last updated May 14, 2026

#How concurrency works

Every Tend project has a concurrency limit: the maximum number of job runs that may be in the running state at the same moment. A run is running from the instant Tend begins delivering it to your handler until your handler acknowledges it, fails, or hits the maximum job runtime of 15 minutes. Runs that are due but cannot start because the limit has been reached are not dropped and do not fail. They wait in a queue and start as soon as a slot frees up.

The limit is applied per project and per environment. Your tnd_dev_ keys and your tnd_live_ keys draw from separate pools, so a runaway development script cannot starve production. Within one environment, every job and every schedule shares the same pool unless you assign a job to a named queue, described below.

Concurrency is counted in runs, not in jobs. A cron schedule that fires every 30 seconds and whose runs take 90 seconds each will, in steady state, hold three slots by itself. This surprises people exactly once.

#Limits by plan

PlanConcurrent runsActive schedulesRequests per minute
Hobby51060
Pro50250600
Scale5005,0003,000
Enterprise2,000 (dedicated pools available)Custom10,000 (negotiable)

#Named queues

By default every job lands in a queue called default. You can create additional queues to partition your concurrency budget, so that a burst of low-priority work (say, nightly report generation) cannot crowd out latency-sensitive work (say, sending a password-reset email). Each queue has an optional max_concurrency value. The sum of all queue limits may exceed your plan limit, but the plan limit is always the hard ceiling across all queues combined.

Queues are created implicitly the first time you reference one, and configured explicitly with a PUT to /queues/{name}. Queue names are lowercase, may contain letters, digits and hyphens, and are limited to 64 characters.

cURL
curl -X PUT https://api.tendcomputer.com/v2/queues/reports \
  -H "Authorization: Bearer $TEND_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"max_concurrency": 10}'

#Assigning jobs to queues

Pass the queue field when creating a job or schedule. Runs created by a schedule inherit the schedule's queue. If a job references a queue that has never been configured, it is created with no max_concurrency of its own and shares the plan-wide pool.

cURL
curl -X POST https://api.tendcomputer.com/v2/jobs \
  -H "Authorization: Bearer $TEND_API_KEY" \
  -H "Idempotency-Key: nightly-report-2026-05-14" \
  -H "Content-Type: application/json" \
  -d '{
    "queue": "reports",
    "run_at": "2026-05-15T02:00:00Z",
    "target": "https://app.example-shop.dev/hooks/build-report",
    "payload": {"report": "inventory", "date": "2026-05-14"}
  }'

#Ordering and fairness

Within a single queue, Tend starts runs in order of their scheduled time, oldest first. Ties are broken by job creation time. Across queues, free slots are handed out round-robin, weighted by each queue's unused headroom, so one deep queue cannot monopolize capacity that another queue is entitled to.

Ordering describes start order only. Because runs execute concurrently and can retry, Tend does not guarantee that run A finishes before run B, even when A started first. If your work requires strict serialization, set the queue's max_concurrency to 1 and make your handler idempotent anyway. Retries can still reorder effects.

  • Start order is by scheduled time within a queue, then by creation time.
  • A retrying run re-enters its queue at its retry time, not at the front.
  • Setting max_concurrency to 1 serializes starts, not effects.
  • A run that exceeds 15 minutes is terminated, marked failed, and its slot is freed immediately.

#Backpressure and monitoring

A queue that grows without bound is a signal, not a feature. Fetch queue statistics to see how many runs are waiting and how long the oldest has been waiting. The oldest_waiting_seconds field is the single most useful number for alerting: if it keeps climbing, your handlers are slower than your arrival rate.

JSON
{
  "name": "reports",
  "max_concurrency": 10,
  "running": 10,
  "waiting": 1842,
  "oldest_waiting_seconds": 613,
  "project_running": 27,
  "project_limit": 50
}

If you consistently sit at your plan's concurrency limit with a growing backlog, either shorten handler runtime, split work across queues, or move up a plan. Pro allows 50 concurrent runs and Scale allows 500. Enterprise customers can request dedicated concurrency pools with capacity reserved for their project.