Skip to main content

Workers (Background Execution)

Workers are the engine room behind everything that happens when you're not looking. They're the background-execution layer that runs your Agents, generation pipelines, scheduled updates, and Mission ticks reliably, in parallel, with retries — so the platform's autonomous operation keeps humming whether you have one Work or a hundred.

You rarely manage Workers directly. They're the "who's actually doing the job" answer underneath the Agents and schedules you do manage.

What Workers run​

Job kindWhat it does
Agent heartbeatsWake each active Agent on its cadence, run its decision loop, record the run.
Agent tasks & chat repliesExecute work assigned to an Agent; reply when an Agent is mentioned.
Generation pipelinesBuild and refresh a Work's content and code.
Scheduled updatesRe-run a Work's pipeline on its cadence.
Mission ticksGenerate fresh Ideas for scheduled Missions.
Inbound emailTurn incoming mail into Tasks or conversations.
Ingest & extractionNormalize and extract uploaded Knowledge Base sources.
Community PR processingTriage and merge community contributions.
DigestsAssemble and deliver the daily and weekly briefings.
Memory consolidationRun the consolidation pass over Memory on a cadence, not only when you press the button.
Goal evaluation & advanceAsk "is the number there yet?" and "should we keep working on this?" for each Goal.
Event ingest & triggersProcess the event spine and fire inbound triggers.
Webhook & notification deliverySign, POST and retry outbound webhooks; deliver notifications to each channel.
Terminal sessionsHost the durable shell behind an Agent Terminal.
Long-running plugin callsExecute plugin operations routed as long-running (an explicit profile on the call, or the manifest).
Housekeeping sweepsReclaim stranded runs and leases, prune transcripts and merged branches, settle credit meters.

Scheduled jobs and their cadences​

These jobs are cron-driven: nothing enqueues them, they simply fire. Times are UTC.

JobCadenceWhat it does
agent-heartbeat-dispatcherEvery minute (configurable)Claims the Agents whose heartbeat is due and enqueues one agent-heartbeat run each. Interval = AGENT_DISPATCH_INTERVAL_MINUTES.
work-schedule-dispatcherEvery minute (configurable, 1–60)Finds Works whose scheduled update is due and enqueues the pipeline.
mission-tickEvery minuteSpawns fresh Ideas for Missions whose cron matches.
task-recurrence-dispatcherEvery minuteMaterializes recurring Tasks — per-minute so any RRULE granularity works.
goal-evaluate-dispatcherEvery minuteRe-measures each Goal on its own checkFrequencyMinutes cadence.
user-research-rerun-dispatcherEvery minuteRuns the scheduled Work-proposal batch — the recurring research behind proposed Ideas.
data-repo-sync-dispatcherEvery minute (configurable)Dispatches due data-repository syncs. Cron overridable with DATA_SYNC_DISPATCHER_CRON.
goal-advance-dispatcherEvery 5 minutesDrives the Goal execution loop — proposes and advances the next step.
event-ingest-tickEvery 5 minutesProcesses the event-ingest spine, including the pull path.
credits-meter-flushEvery 5 minutesResends metered usage rows the settlement path could not deliver. See Credits & Billing.
fleet-job-lease-sweeperEvery 5 minutes, at :03 / :08 / …Reclaims expired Fleet job leases so a queue does not freeze when every node goes away.
deploy-ready-pollerEvery 2 minutesFlips mid-deploy Works to READY once their health endpoint answers 200.
task-pr-status-syncEvery 2 minutesRefreshes cached PR and CI verdicts on Tasks whose pull request is still open.
agent-run-sweeperEvery 2 hours, at :23Reaps agent_runs rows stranded in queued / running by a killed worker.
model-account-healthEvery 6 hours, at :19Checks each enabled provider account's credential; marks accounts expiring, expired or rejected. Never changes their order.
credits-daily-grantDaily, 00:05Grants the daily credit allowance.
anonymous-user-cleanupDaily, 03:17Purges expired anonymous users and the files they uploaded.
terminal-transcript-gcDaily, 03:17Prunes terminal transcript chunks past each run owner's retention window.
kb-reconcileDaily, 03:42Reconciles Knowledge Base documents against their mirrored files.
task-branch-gcDaily, 04:41Deletes remote task/* branches for finished or abandoned Tasks, per the Work's branch-cleanup policy. See Repositories.
digest-dispatcherDaily, 07:15Sends the daily digests; on Mondays the weekly ones ride the same run.
memory-consolidation-tickDaily, 08:37Consolidates Memory across the workspace.
Why the odd minutes

The per-minute crons already own :00 of every minute, and the midnight and hourly jobs crowd the top of the hour. So every sweeper and daily job is deliberately offset — agent-run-sweeper at :23 past every second hour, fleet-job-lease-sweeper on the 3/5 pattern (:03, :08, :13 …), kb-reconcile at 03:42 rather than 03:00. It keeps a slow housekeeping pass from colliding with the dispatchers that must not be delayed.

On-demand jobs​

Everything else is enqueued by something you — or an Agent — did. The API returns immediately, usually 202 Accepted, and a Worker picks the job up.

JobEnqueued whenTime limit
work-generationA Work's content and code pipeline runs, manually or on schedule.5 hours
work-importYou import an existing repository or dataset.2 hours
work-onboardingThe zero-friction onboarding flow finishes registration.2 hours
kb-backfill-skeletonAn operator backfills the Knowledge Base skeleton for a list of Works.2 hours
agent-task-executeA Task is assigned to an Agent.60 minutes
idea-build-executeAn Idea is built into a Work, retried or rebuilt.60 minutes
run-plugin-operationA plugin operation is declared or called long-running.60 minutes
terminal-sessionYou open an Agent Terminal.60 minutes
template-customizationA template is customized for a Work.60 minutes
workflow-runPOST /api/workflows/:id/run creates the run row, then enqueues it.—
webhook-deliveryAn event matches an outbound webhook subscription.30 minutes
kb-reembed-workA Work's Knowledge Base needs re-embedding wholesale.30 minutes
kb-transcribeAn uploaded audio source needs a transcript.30 minutes
kb-normalize-videoAn uploaded video source needs normalizing.30 minutes
kb-normalize-audioAn uploaded audio source needs normalizing.15 minutes
kb-embed-documentA Knowledge Base document is created or updated.10 minutes
kb-mirror-documentA Knowledge Base mutation must be written to its .yml sidecar and .md body.10 minutes
kb-org-overlay-fanoutAn org-scope Knowledge Base document must fan out to the Works it covers.10 minutes
agent-heartbeatThe dispatcher claims a due Agent — or you fire one by hand.30 minutes (default)
agent-chat-replyAn Agent is mentioned in a Task chat thread.5 minutes
notification-channel-deliveryA notification has to reach a channel.5 minutes

Two of those limits are configurable rather than fixed: agent-heartbeat uses AGENT_MAX_RUN_DURATION_SECONDS (default 1800), and the platform-wide ceiling for any single job is five hours.

How Workers behave​

  • Parallel — many jobs run at once; a dispatcher claims due work in batches so thousands of Agents and schedules scale without stepping on each other. The Agent dispatcher claims up to AGENT_DISPATCH_MAX_BATCH Agents per tick (default 25).
  • Safe under contention — a single Agent's heartbeat can only be claimed by one Worker at a time (compare-and-set), so nothing runs twice.
  • Retried — transient failures (network blips, provider rate limits, upstream 5xx) are retried with backoff before a job is marked failed.
  • Bounded — runs have timeouts; an Agent that keeps failing auto-pauses rather than burning budget.
  • Observable — every run emits activity-log entries and surfaces on the relevant Dashboard, with cost attributed to the right Agent, Task, or Work.

Retries and backoff​

Retry policy is set per job, because a chat notification and a customer's webhook endpoint deserve very different patience.

PolicySchedule
Shared default3 attempts, 1s → 10s backoff (factor 2, jittered). Written explicitly in trigger.config.ts only when TRIGGER_DEV_ENABLE_RETRIES=true, which is also what switches retries on in local dev; without the flag the job runtime's own defaults apply, and those are the same numbers — so this is the production behaviour either way.
webhook-delivery30s → 2m → 10m → 1h → 6h → 1d → 1d, capped at 24h between attempts, up to 10 attempts by default (WEBHOOK_MAX_CONSECUTIVE_FAILURES), after which the subscription is dead-lettered.
notification-channel-delivery30s → 2m → 8m → 32m → 2h, 5 attempts, 6h cap. A chat or email notification has little value a full day late.
work-onboarding3 attempts — onboarding is long, but a failed run should not silently strand a new account.

When an Agent keeps failing​

Every failed heartbeat increments the Agent's error count. Once it reaches that Agent's Pause after failures threshold (default 3, editable on the Agent's Settings tab), the Agent flips to error status and its next heartbeat is cleared — the dispatcher stops waking it, so a broken Agent cannot burn its budget overnight. A successful run resets the counter. Resume the Agent once you have fixed the cause.

When a Worker dies mid-job​

A killed process — out of memory, an eviction, a deploy, a laptop going to sleep — cannot clean up after itself, so two sweepers do it instead:

  • agent-run-sweeper reaps agent_runs rows stranded in queued or running. Its cutoff is derived from the longest Agent task's duration ceiling and clamped to at least three times that ceiling, so a run that is legitimately still retrying is never reaped out from under itself.
  • fleet-job-lease-sweeper reclaims expired Fleet leases. A lease is a deadline, not a lock. Reclaim also runs inline on every lease poll, which covers the ordinary case of one node dying while its siblings keep polling; the cron covers the case inline reclaim structurally cannot — a fleet where every node went away at once. Detection lag is bounded by the lease TTL plus five minutes.

Where they run​

Workers are powered by the platform's background-jobs infrastructure. In the cloud, this is fully managed for you. When you self-host — or run the Desktop App — Workers run alongside the rest of the stack, and you can also point them at an external or self-hosted jobs backend.

Which engine actually executes a job is a plugin choice, not a hard-coded dependency. Six job runtimes ship with the platform:

RuntimePluginShape
Trigger.devjob-runtime-triggerThe default; managed cloud or self-hosted.
Temporaljob-runtime-temporalDurable workflow engine for teams that already run one.
BullMQjob-runtime-bullmqRedis-backed queues inside your own deployment.
pg-bossjob-runtime-pgbossQueues on the Postgres you already have — no extra service.
Inngestjob-runtime-inngestHosted, event-driven execution.
Fleet nodejob-runtime-nodeYour own machines leasing work through Fleet.

The instance-wide choice is the EVER_WORKS_JOB_RUNTIME environment variable, with a per-tenant overlay under Sidebar → Settings → Job Runtime.

Read the runtime page before you switch

Part of that seam is shipped and part of it is not — the selector is honoured on the Agent-run dispatch path, but the queue dispatchers still route to Trigger.dev in the bundled build, and per-tenant bring-your-own credentials are recorded without yet being injected into runs. Job Runtimes spells out exactly which half is which.

Watching and driving a job​

  1. Watch the feed. Activity is the workspace-wide log; every Agent, Task and Work also has its own Activity tab showing the runs that touched it. Failures land there with the reason attached.
  2. Open the Agent. The Agent detail page's Dashboard tab is the per-Agent cockpit: the status pill (draft, active, paused, running, error, archived) and tiles for Heartbeat (its cadence, or Manual), Idle behavior, Last run and Next heartbeat, over a health strip that turns red once the error count is non-zero. To fire a heartbeat immediately instead of waiting for the next dispatcher tick, call POST /api/agents/:id/run-now — it answers 202 Accepted and is rate-limited to 30 calls a minute — or ask the platform chat to run the Agent now, which is the run_agent_now tool posting to that same endpoint. The response says what happened: a run id when a run was enqueued, or a skip reason (already-claimed, inactive, concurrency-limit, agent-missing) when one was not.
  3. Read the run. Agent runs are listed on the Agent's Activity tab and across the workspace under Sidebar → Teams → Sessions, where you can open a run, follow its steps, and steer or interrupt it while it is live.
  4. Re-run a schedule by hand. A Work's Schedule card has its own Run now button (the schedule has to be active), and a Mission's detail page has one for its tick. Both enqueue exactly the job the cron would have.
  5. Check what it cost. Every run attributes spend to the Agent, Task or Work that caused it — see Budgets & Usage.
A job that never seems to run

Check three things, in order. The entity is active — a paused or errored Agent is skipped before anything is claimed. The dispatcher is enabled for the deployment — AGENTS_DISPATCHER_ENABLED is not set to false. And the cadence is what you think it is — an Agent with no heartbeat cadence shows Manual on its Dashboard tab and only ever runs when something asks it to.

See also​