delegate_tasks (display name “Delegate to Agents”), the native tool that lets one agent, mid-task, hand work to OTHER agents in the same workspace and act on their combined results.
Overview
Use this tool when an agent needs to fan work out to specialist agents and then continue once they’re all done — e.g. “research these 5 tickers” → spawn one task per ticker on a research agent, wait, then summarize. Key properties:- Independent runs. Each delegated task is a fully independent run of the target agent with the prompt you provide. The parent cannot steer a child mid-run.
- Barrier semantics (default). By default the parent’s tool call finishes only when every delegated task reaches a terminal state (
completed/failed/stopped— any outcome). The parent then resumes with each child’s status + output and decides what to do next. - Fire-and-forget (
async: true). Whenasyncis on, the parent starts the children and continues immediately in the same turn — it does not wait and never receives their results. No await group is created, so the parent staysrunning. Use this for work whose output the parent doesn’t need back (kick off a long background job, trigger notifications, etc.). See Async / fire-and-forget. - Fan-out. A single task entry can carry an
inputsarray to spawn one child per value (a{{item}}placeholder in the prompt/title is substituted per value). - Bounded. Recursion depth, per-call count, and wait timeout are all configurable (see Configuration). (
timeout_minutesis ignored whenasyncis on — there is nothing to wait for.) - Pausing tool (unless async).
delegate_tasksis registered as a pausing tool in the orchestrator: while its child awaits are pending, the parent stays inwaiting_for_inputand the agent loop does not advance. Anasynccall creates no awaits, so the orchestrator keeps the loop running instead of pausing (see Flow).
Per-agent delegation settings
The agent edit page (Settings tab → Agent delegation,AgentDelegationProfileSection) exposes the two directions of agent-to-agent collaboration:
- Delegate out (caller) — “Automatically choose from available agents”. Backed by
config.auto_delegate(nil/absent = on). When off, the agent can never spawn subtasks:agent.filterToolsByAgentConfigstripsdelegate_tasksfrom its tool list, which also drops the agent roster from the decision prompt (sincecanSpawnAgentskeys off that list). When on, the agent decides on its own whether to delegate. - Receive delegations (callee) — “Allow other agents to delegate to this one”. Backed by
config.delegation_enabled(nil/absent = on) plus thedelegation_profileeditor revealed beneath it (see Delegation profiles).
AgentAutoDelegateSection and AgentDelegationProfileSection).
Per-task override (home task runner)
The home task runner (/, TaskFeed home mode) fuses a “Let other agents help with this task” toggle into the bottom of the composer (FusedChatInput), on by default (the legacy “Assign to an agent” item is removed — the task runs on the default agent and this toggle governs collaboration) — so an ad-hoc task can spawn any of the user’s workspace + public agents out of the box (the default agent already has delegate_tasks + auto_delegate on, and the roster spans workspace + public). Turning it off sends disable_delegation: true on POST /api/chat/message, which is persisted at INSERT time into tasks.source_metadata ({"disable_delegation": true}). At decision time agent.taskDisablesDelegation strips delegate_tasks (base and configured instances) from that task’s tool list, so it can’t delegate regardless of the agent’s own config. This is a disable-only override: leaving it on respects the agent’s own delegation capability.
Flow
The suspend/resume is built on the existingtask_awaits + signals machinery; the only new signal source_type is child_task.
- The worker executes the
delegate_tasksstep. For each (expanded) target it resolves theagent_group_idto the runnable target agent and builds a child task (tasksrow withparent_task_id,parent_step_id,spawn_depth = parent+1,source = 'spawned') plus its first user message — without emittingnew_messageyet. - It creates one
task_awaitper child in a group (source_type = child_task,correlation_key = child task id,strategy = all_required,scheduled_resume_at = now + timeout). - Only after the awaits are committed does it emit
new_messageper child, so a fast child can’t reach a terminal state before its await row exists (which would otherwise leave the parent hung). - The parent task is left in
waiting_for_input. The chat feed renders a Delegated tasks card listing each child with a live status pill (polled every 2s). - When a child reaches any terminal state,
TaskTerminalNotifierenqueues achild_tasksignal (dedup_key = child id, so repeated terminal fires are idempotent). - The
SignalMatchermatches the signal to the parent’s pending await, completes it, and once the wholeall_requiredgroup is complete emits anawait_completedevent. - The orchestrator resumes the parent: it inserts ONE consolidated
task_messagesummarizing every child (title / agent / status / output) and adecide_next_actionstep, then flips the task back torunning. - On
scheduled_resume_attheAwaitTimeoutProcessorre-checks each child’s liveness. A child that is still non-terminal (running, or parked inwaiting_for_input— which a user can still answer on the child’s own task page) is not timed out: its await is rescheduled one window forward so the parent stays paused. The parent therefore resumes only once every child reaches a terminal state (via thechild_tasksignal). This guarantees a parent never completes while a delegated subtask is still pending.
Async / fire-and-forget
When the call setsasync: true, the parent does not block on the children:
- The worker still builds each child task and emits
new_message(steps 1 + 3 above) — the children run exactly as normal. - It skips the await group (step 2) entirely — no
task_awaitrows, no timeout, no resume path. - The tool returns a result with
status: "spawned_async"(and emptygroup_id/strategy). The orchestrator’sisNonPausingDelegatehelper reads that status and treats the completion as non-pausing: it creates the nextdecide_next_actionstep and leaves the taskrunning, so the agent loop continues in the same turn with the children already started.
source = 'spawned', parent_task_id set), so they still appear in the Delegated tasks card, the task list ”↳ subtask” badge, and the child’s “Delegated by …” back-ticker — the parent just isn’t waiting on them. Because no await exists, each child’s terminal child_task signal lands as no_match (harmless, the same path as a stopped parent).
Tool input
- No
inputs→ one child per task entry. inputsset → one child per value;{{item}}inprompt/titleis replaced (if absent, the value is appended). The per-call limit applies to the total expanded count.async→false(default) waits for all children and returns their combined output;truestarts them and continues immediately without their results. Exposed to the LLM, and surfaced in the configured-tool modal as a toggle (“Run in the background (don’t wait for results)”,x-display-name) that an author can pin to a static value.
Examples
1. Wait for results (default). Hand a sub-analysis to a specialist agent and use its output:waiting_for_input; once the child finishes it resumes with a consolidated summary message and decides what to do next.
2. Fan-out and wait. Run the same agent over a list (one child per value):
status: "spawned_async" and the parent continues in the same turn — e.g. it can immediately reply to the user “I’ve kicked off the re-index in the background” without waiting for it to finish.
Configuration
Runtime bounds (backend/config.yaml, env-overridable)
Delegation allowlist (optional)
Configure a scoped instance via the standard Create Configured Tool modal (Tools → ⚙): pick Delegate to Agents, give it a name, set any parameters, and use the “Agents this tool can delegate to” multi-select.- Leave all unchecked → any workspace agent (default).
- Select some → allowlist. The tool will only delegate to those agents; a task targeting any other agent is rejected at execution.
allowed_agent_group_ids binding (a comma-separated list of agent_group_ids, x-exclude-from-llm so the LLM never fills it). At execution the binding is injected into the tool args and enforced by parseAllowlist. The decision-time roster still lists all workspace agents (discoverability); enforcement is at call time.
Because the allowlist value itself is hidden from the model, two things keep it from picking a restricted delegate tool for an agent it can’t reach (which would otherwise fail at execution):
- Description hint (routing):
loadConfiguredToolsresolves the allowlistagent_group_ids to agent names and appends them to the tool’s description — e.g. “This delegate tool can ONLY delegate to: “Pricing Agent” (agent_group_id …). To delegate to any other agent, use the delegate_tasks tool instead.” So the model sees, before calling, which agents each configured delegate tool can reach (agent.describeDelegateAllowlist/buildAgentNameByGroup). - Actionable rejection (recovery): if it still targets a disallowed agent, the execution error names the allowed agents and points to the open
delegate_taskstool (spawn_agent_tasks.allowedAgentNames), so it recovers immediately instead of leaving a failed step.
Discoverability (how the agent knows who to delegate to)
Whendelegate_tasks is enabled, the decision-time system prompt gets an injected roster, parsed as its own “Delegatable Agents” section in the debug view:
@-mention agents in the instructions editor; resolution is name-based against this roster.
In the admin decision-log debug view (DecisionLogDetail), this roster renders under the System Prompt breakdown with one row per delegatable agent — the same way each tool schema is its own row — so individual agents can be inspected separately (parseDelegatableAgents splits the section on the - <name> — agent_group_id: bullets).
Delegation profiles
Each agent can have an LLM-generated delegation profile — a shortdescription + example_input — stored at config.delegation_profile and surfaced under each roster entry so decide_next_action can pick the right agent.
Roster eligibility (who shows up as a delegation target). An agent appears in another agent’s roster only when both hold (agent.buildAvailableAgentsSection):
- Delegation is on —
config.delegation_enabledis unset ortrue(the default). Setting it tofalseremoves the agent from every roster. - It has a profile description — an agent with no
config.delegation_profile.descriptionis skipped, since the decider has nothing to match on. (Delegation is allowed by default, but undescribed agents are invisible until a profile exists.)
is_public = TRUE, from ListPublicAgents) shared from other workspaces. Public agents are tagged (public) in the prompt and badged Public in the debug view. Workspace agents take precedence over a public duplicate (same agent_group_id), shown once and untagged. The current agent is always excluded (no self-delegation).
Cross-workspace execution. Delegating to a public agent creates the child in the caller’s workspace (BuildChildTask sets the caller’s workspace_id) — only the shared agent definition (config/prompt/tools) is reused; the child’s data, integrations, and credentials resolve from the caller’s tenant, so there’s no cross-tenant data exposure. At execution the tool’s allowed-target set is the union of agentRepo.List(workspaceID) and agentRepo.ListPublicAgents() — both trusted server-side queries; the LLM-supplied agent_group_id still can’t widen it. Eligibility only gates the decision-time roster; the allowlist + self checks are enforced independently at execution.
- Generate per agent: the agent’s Settings tab has a “Delegation profile” section (
AgentDelegationProfileSection) with a master Delegatable toggle (on by default), a Generate with AI button, and editable description/example saved on blur.delegation_enabledpersists via the standard agent update (config.delegation_enabled). - Backfill all:
POST /api/agents/delegation-profiles/backfill(workspace-scoped;?force=trueregenerates existing ones) fills in any agents missing a profile. - Single endpoint:
POST /api/agents/:id/delegation-profilegenerates + saves one. - Generation:
services.AgentProfileService.GenerateDelegationProfileprompts an LLM (gemini-2.5-flash) with the agent’s name, instructions, and tools; prompt text lives inprompts/files/agent/agent_delegation_profile_{system.txt,user.tmpl}. The call setsDisableReasoning: trueso thinking tokens don’t starve the small output budget (a thinking model otherwise returns truncated/empty JSON).
UI
- Parent task feed: a “Delegated tasks” card lists each child with a live status pill (“3 of 4 complete”); deleted children show as
deleted. Backed byuseChildTasks→GET /api/tasks/query?parentTaskId=…. For an async delegation (resultstatus = spawned_async) the header reads “Started tasks” instead, since the parent isn’t waiting on them. - Child task feed: a floating ”↳ Delegated by <Agent>” ticker links back to the parent (
feedresponse carries aparentref with the parent’s agent id). - Task list: spawned children appear in the main list with a ”↳ subtask” badge (click → filter to that parent’s children).
source=spawnedis included in the default view. A parent row shows an “N subtasks” toggle (TaskTitleCell). Clicking it expands the delegated-subtasks list (SubtaskList) as its own full-width row directly beneath the parent — so the parent row’s cell heights and alignment never change when it opens. Each child shows its live status, the agent name it was delegated to, and a link to its task page; the list loads lazily and live-polls only while open. Expansion state lives inTasksTable(via tablemeta.subtaskExpansion). The count comes fromsubtask_counton the task-query response (a correlatedCOUNT(*)of non-deleted children, workspace-pinned).
Stopping & deleting
Because children are independent runs, stopping/deleting a parent prompts to cascade:- Stop (chat): if the task has delegated subtasks, a dialog offers “Stop only this” vs “Stop this & N subtasks”. Cascade stops each active descendant, cancels their steps and their pending awaits (so an in-flight child signal can’t resume a just-stopped task —
delegate_tasksdoes not fire wake-up signals during teardown). Stopping an already-terminal parent with the cascade flag still stops orphaned children. - Delete (single row + bulk): same prompt; cascades a soft-delete to descendants (
include_subtasks). Note: delete is soft-delete only — it hides children but does not halt ones still executing.
Edge cases & guarantees
- Self / same-group delegation is rejected (infinite-loop guard), as is any
agent_group_idoutside the workspace (tenant isolation / BOLA) or outside the configured allowlist. - Recursion is capped by
max_depth; a child may delegate further only until the cap. - Duplicate resume is prevented by
dedup_key = child idon the signal + the race-safeCompleteAwaitCAS. - Spawn↔await race is closed by committing all child awaits before emitting any child’s
new_message. - Parent stopped/deleted while children run: the parent’s awaits are cancelled, so a child’s terminal signal finds no pending await (
no_match, harmless). - Async (
async: true): no await group is created, so the parent never blocks and the children’s terminalchild_tasksignals always land asno_match(harmless). The parent gets no results back and cannot be resumed by the children — choose this only when the output isn’t needed.timeout_minutesis ignored. The recursion-depth, allowlist, self / cross-workspace, and per-call-count guards all still apply (async changes only the wait, not who can be spawned).
Key code
Source mode — deterministic fan-out (source)
Setting source INSTEAD of tasks[] switches this tool to deterministic fan-out: code pages a source tool and creates exactly one child task per item. Always synchronous, single-target, mutually exclusive with tasks[]/async.
Rollout is per agent. The source args carry x-exclude-from-llm in the shared tool definition, so no agent’s model sees them by default. Toggling “Process whole lists reliably (deterministic delegation)” on the agent’s Settings tab (or setting config.settings.delegate_source_mode: true) reveals the group (and its guidance) to that agent only, and execution re-checks the same setting fail-closed. The batch_fanout.enabled runtime flag remains the fleet-wide kill switch. Piloting on one agent in prod = arm the flag once, then flip the setting on that agent; unticking it reverts instantly.
Overview
delegate_tasks asks the LLM to enumerate items and hand-type one child per input, which drops work at batch scale (in a production run, 247 encounters became 63 children — the model fetched one page and enumerated part of it). Source mode moves counting and hand-off into code. The model supplies intent — which source lists the items, which agent handles each one — never the item list itself.
Key properties:
- Complete by construction. The service loops every page of the source tool and creates exactly one child task per item. A missing/empty item key fails the whole batch rather than silently dropping the item.
- Idempotent re-runs. Each child carries a stable
item_key(fromitem_key_field), unique per (parent task, target agent group, source) over non-deleted rows (uq_tasks_parent_source_item_key). Re-running the same parent skips items whose child is already terminal (skipped) and re-awaits ones still running (attached). A re-awaited child is moved onto the new run (batch_run_idis reassigned), so the run that is waiting on it also owns it in the rollup and the feed card — the superseded run keeps only the children nothing re-claimed. A fresh parent task may process the same items again. - Derived rollup. A
batch_runsrow records what was enumerated/enqueued (total_items,created_children,skipped_count, page progress; statusrunning→enqueued/failed). Child outcomes are always derived live fromtasks(GROUP BY status WHERE batch_run_id = …) — no maintained counters to drift. - Barrier semantics. The parent pauses on an
all_requiredawait group (same machinery asdelegate_tasks) and resumes with one compact rollup message — totals plus only the non-completed children enumerated, so a 250-item batch does not flood the parent’s context. - Bounded.
max_pages(default 100),max_items(default 250, hard cap 1000, operator-clampable), spawn-depth guard, self-spawn block, an optional config-time target allowlist (allowed_agent_group_ids, same contract asdelegate_tasks), and the standard delegation wait timeout. Exceeding a cap fails the batch — it never truncates. Credentialed sources work: the runner reuses the WORKER’s own step resolver (configured bindings, integration discovery, secret injection), injected late-bound, so a source call resolves exactly as a real step does.
Authentication and enablement
- No external authentication; the tool is native and works entirely against the platform database plus the (already-enabled) source tool.
- Per-agent opt-in via
config.settings.delegate_source_mode(invisible and non-executable otherwise), plus the fleet-widebatch_fanout.enabledruntime flag. - Armed by a runtime flag:
batch_fanout.enabled(seeded off, read fail-closed). Until an operator enables the flag, the tool returns an error result and creates nothing. Flip it with the admin runtime-flags API/UI.
Inputs
Outputs
Structured content on the feed card:status—awaiting_child_tasks(parent paused),batch_already_complete(every item already handled; failures called out explicitly), orbatch_no_items(the source returned 0 items — named distinctly so wrong filters are never mistaken for a finished batch).batch_run_id,group_id,total_items,created_children,skipped_terminal(skipped items whose outcome is already decided — deliberately excludesother_run, which is live work under a concurrent run and must not read as finished),skipped(split by what happened to each skipped item:completed/failed/stopped/other= its child row is gone /other_run= a concurrent run of the same batch claimed it — live work, not lost),attached_active,start_failed(children that could not be woken; their awaits complete asstart_failedso the parent never hangs).children— a bounded echo (first 25) of child refs; the full set lives intasks(batch_run_id = ...).
Source-tool authorization
source_tool runs with the calling agent’s authority, checked fail-closed before the first page. A configured tool authorizes through its own agent/workspace-scoped row; a base tool must be granted to the agent by the same rules the agent loop applies — its tool list plus phase-declared tools, or, when the tool list is empty, the default-enabled set only. CodeAct agents never get the default-enabled fallback: their grant is always tool list + phase-declared tools + run_code, however empty the list. Default-template agents additionally get every default-enabled tool unioned in, on every branch, exactly as the loop does. delegate_tasks itself can never be a source, including via a configured alias.
source_args_json is filtered to the source’s declared schema properties before anything is resolved. The tool itself ignores undeclared arguments, but the credential resolver reads some by fixed name — python_code_tool’s merged input takes args["integration_id"] and resolves it through an unscoped lookup — and those names are filled server-side from configured-tool bindings, never by the model. Dropped keys are logged. Server-resolved bindings are merged after the filter, so they are unaffected.
Not callable from run_code
Source mode pauses the calling task on an await group, which in-sandbox code cannot do — run_code would carry straight past the pause and the group’s later completion would force a decide step against a mid-flight task. A delegate_tasks call carrying source is therefore refused from CodeAct with a message pointing at the agent loop. The tasks[] shapes (including async: true) stay callable from code as before. The check runs on the resolved base tool, so a configured alias can’t route around it.
Paging contract
The paging arg names are checked against the source’s input schema before the first call: declared, numeric, distinct, and not statically bound. An undeclared arg would be silently dropped and mis-page the source; a bound one would repeat a single page; a non-numeric or duplicated name means the argument is not a pagination control at all (pointing both names at one property would otherwise let any tool qualify as “pageable”). Nothing can prove a declaredpage argument really pages; the enumeration rules below are the actual guard. The service then pages 1, 2, … and every item must carry item_key_field.
Termination — a short page is never treated as the end, because a source that honours page but ignores limit returns short pages forever and stopping at the first one would drop the rest. Enumeration ends when:
- a page comes back empty; or
- a page repeats every item already seen — the source ignores paging entirely (
fetch_all-style argument, or no paging support) and hands over its whole result set per call. That set is accepted as complete only if it was short; a repeated full page fails the batch, since the source may be capping atpage_sizewith rows that can never be reached.
item_key_field isn’t unique per item, and it fails the batch rather than silently collapsing it.
Side effects
- Creates one
tasksrow + first user message per item (sourcespawned,item_key/batch_run_idset), abatch_runsrow, and anall_requiredawait group on the parent. - Child tasks bill their own steps; the fan-out call itself is free.
- Children currently bypass agent-group admission control (same as
delegate_tasks);max_itemsis the guardrail. Run large batches off-hours or against a dedicated agent group.
Mixed parents (classic + batch on one task)
Resumes are symmetric: a batch group’s summary covers only its own run’s children, and a classic group’s summary excludes batch children entirely — each batch reports through its own resume, so neither direction re-renders the other’s results into the parent’s context.Superseded runs
A re-run cancels every pending await left by a previous run of the same step — matched onbatch_run_id in the await context, and on the owning batch_runs.parent_step_id. This covers the crash-mid-emit case and items the earlier attempt awaited that the new enumeration no longer returns, which would otherwise sit pending until the timeout processor completed them and fired a resume at an already-busy parent.
The step match is what makes the cancel safe: two runs of the same step are a retry or a reclaimed step, where superseding is the intent, while two runs from different steps are separate live delegations whose awaits must survive. A concurrent delegate_tasks delegation carries no batch_run_id at all and is never touched. task_awaits has no workspace_id column, so the statement asserts tenant scope with an EXISTS against the owning task.
Concurrent runs on one parent
Two overlapping source-mode calls on the same parent are not a supported shape in v1. The supersede is scoped to the same step, so a second run from a different step leaves the first run’s pending awaits alone — deliberately, so a genuinely live sibling isn’t orphaned. The consequence is that if the second run dedups onto a child the first run is still awaiting, committing its await group hits the pending-unique index and fails that whole batch loudly (nothing partially emitted,batch_runs marked failed, safe to re-run once the first finishes).
That is the intended trade-off: fail a duplicate run rather than silently steal a child from a live one. A graceful “another run owns it” skip needs one-in-flight-run-per-parent enforced at the step/event layer, which is a filed follow-up.
Failure & resume
- Pagination/enqueue failure: the
batch_runsrow is markedfailedwith the error andcurrent_page. No delegated task is started, but children built before the failure are already committed as rows — inert, with no await and no start event, so none of them runs. Re-invoking with the same parameters is safe and is how they heal: dedup adopts those rows (re-attaching or skipping them) instead of duplicating, and the remaining pages fill in. The tool’s error text says exactly this rather than claiming nothing was created, and a unit test pins the end-state (rows committed, no await group, nothing emitted). No reaper collects them — a re-run is the intended recovery. - Child failures: visible in the resume rollup (derived from the child tasks) and, on re-runs, in the
skipped.failedsplit. No automatic per-item retry yet (see non-goals) — re-run the batch to mop up, or handle in the parent. - Interruption (deploy, scale-down, shutdown drain): when the fan-out’s own context dies mid-start, the run stops rather than mis-recording the remaining children — they stay pending with pending awaits, the run is marked failed as interrupted, and the swept step’s automatic retry adopts every child already created (started ones are not re-woken).
start_failedis reserved for a child the platform genuinely refused, or whose row is gone. - Attached child already terminal: a re-attached child that went terminal before the new await group committed can never re-fire its signal, so its await is completed inline — including when the batched reconcile read fails and the per-child fallback read discovers the terminal state. A transient read failure leaves the await to the child’s signal or the timeout processor’s liveness re-check (the child was live at prefetch); only a child whose row is genuinely gone is settled as
start_failed, so a possibly-running child’s outcome is never falsified. - Timeout:
timeout_minutesis a liveness window, not a deadline — at each expiry still-running children are re-checked and the wait extends; the parent resumes only when every child is terminal (incl.start_failed). - Unreadable await group on resume: the rollup is scoped by the completed await’s group and its message id is derived from that group, so if the await or group can’t be read the summary would be both unscoped and un-deduplicated. The event is requeued instead — for 15 seconds, since
orchestrator_eventshas neither an attempt counter nor a scheduled-at column, so a requeue is re-picked on the next tick with no backoff. Past that window (or on a non-retryable failure) the parent resumes with the unscoped summary — for a classic parent that is exactly what it always received; a batch parent gets every batch’s rollup block rather than only its own (bloated but truthful). Only if even that read fails does a one-line go-read-your-subtasks note go out, so the parent never wakes with no information at all. The message is keyed to the await group where one is known — and to the event id when the await itself is unreadable — so a crash-reclaimed retry of the same event can’t post it twice. - Skip & send now: skipping the wait relabels the feed card
skipped(standardAwaitSkipperbehavior); children keep running.
Expected errors
Deterministic fan-out is not enabled on this deployment (runtime flag "batch_fanout.enabled" is off)— ask an operator to arm the flag. A transient flag-read failure returnsCould not verify the ... runtime flaginstead (fail-closed, retryable).source tool "X" is not enabled for this agent— the source tool isn’t in the calling agent’s tool list (authorization is fail-closed; enable the tool on the agent first).source tool "X" is not available— the source tool isn’t registered.items_path ... not found / does not point to an array— wrongitems_pathfor this source tool’s result shape.item is missing item_key_field ...— wrongitem_key_field; the batch fails rather than dropping the item.item key "X" repeated on page N — item_key_field ... is not unique across items— pick a per-item-unique field; a source that ignores paging is detected separately (see the paging contract).source has more than N items (max_items)/more than N pages— deliberate anti-truncation guardrails; raise the cap or narrow the source filters.An agent cannot fan out tasks onto itself./ max spawn depth — same delegation guards asdelegate_tasks.
max_items / max_pages / page_size are clamped to the hard maxima (safe — a source exceeding the effective cap still fails the run, so nothing is silently truncated), and a tasks[] array sent alongside source is dropped with a warning on the result rather than rejecting the call.
PHI note
Keep child prompts to IDs: the default{{item}} substitution passes only the item key, and the child fetches full details itself at run time. include_item_json copies entire item payloads into child prompts/messages — leave it off for EHR/PHI sources.
Feed card at batch scale
The tool result echoes at most 25 children (maxChildEcho) — that result is replayed into every later LLM call on the task, so a 250-item echo would cost context on every step. The card therefore shows the echoed rows plus an honest N+ of M complete lower bound and a “Showing 25 of M” note.
View All N Tasks expands the card to every child of that batch, read from the live child-task query. That query is scoped server-side by batch_run_id (an optional batchRunId filter on the task-query endpoint, additive like parentTaskId with the workspace still pinned, applied in both the enriched and simple query paths and echoed back as batch_run_id on list items), so each card polls only its own children: a later, larger batch on the same parent cannot push an earlier batch’s children out of the window and freeze that card on stale statuses. Expanded, the progress count is exact.
Dedup namespace
The unique index is(parent_task_id, agent_group_id, batch_source_key, item_key). For a base tool source, batch_source_key is the resolved base tool name; for a configured tool it is the base name plus the configured tool’s row id (base#<uuid> — the id, not the display name, so a rename doesn’t orphan earlier children). Two dimensions of collision are closed: different sources from one parent (get_appointments then get_claims, both carrying id 101) never share a namespace, and neither do two configured instances of one base tool bound to different accounts — without the instance id, an item_key collision across accounts would classify the other account’s unprocessed item as “already done” and silently skip it. Re-running the same source still dedups, because the key is stable.
One consequence to know: running a source via its configured alias and then via the bare base tool are different namespaces, so items may be processed twice across that switch. That direction is chosen deliberately — duplicate work is visible and recoverable, a silently skipped item is not.
Migrations & lock profile
Seven pairs ship with source mode:batch_runs, the three tasks columns (metadata-only ALTERs), the dedup unique index (its own migration — built in the same transaction as the column ALTER it would inherit ACCESS EXCLUSIVE and block reads too; alone it holds SHARE), the batch_fanout.enabled flag seed, the batch_run_id lookup index, the partial task_awaits (group_id) index (its own migration — classic delegation benefits, so a batch_runs rollback must not regress it), and a partial orchestrator_events (task_id) WHERE event_type = 'new_message' index serving the re-run don’t-wake-twice check, otherwise a seq scan over ~19M rows in prod. Prod runbook: after the column migration lands, pre-create the two tasks indexes and the orchestrator_events index CONCURRENTLY — the transactional runner builds them inline under SHARE otherwise, which is seconds for tasks (~450K rows) but not acceptable for orchestrator_events.
The tasks columns themselves are metadata-only ADD COLUMNs, but the index builds are not free. The runner is transactional, so CREATE INDEX CONCURRENTLY is unavailable: each build takes SHARE on tasks and scans the whole table (matching zero rows makes the index empty, not the build cheap), blocking writes for the scan — seconds at ~412K rows / 520MB, and growing with the table. SET LOCAL lock_timeout = '5s' bounds lock acquisition, not the scan; re-run on timeout. Ops may pre-create either index CONCURRENTLY ahead of the deploy, which makes the IF NOT EXISTS build a no-op. The two indexes are deliberately in separate migrations so each transaction holds SHARE for a single scan and commits, rather than one transaction spanning both. The down-migrations set the same lock_timeout, so a rollback fails fast instead of queueing behind live traffic.
batch_runs is created with a plain CREATE TABLE — not IF NOT EXISTS. The table is new in this migration, so a name collision means a database in an unexpected state and should fail loudly rather than silently keep a shape the code does not match.
Non-goals (v1)
Per-item retry/backoff, worker batch-drain/throughput scaling, and batch-aware admission control are deliberate fast-follows, not part of this tool. Three related changes are also deliberately out of this PR, each reviewable on its own: reworking the sandbox admission gate so a configured alias to a loop-control tool can’t reachrun_code (a pre-existing gap — this PR only refuses the source shape); bounding classic delegate_tasks resume summaries, which are still unbounded in child count and per-child output; and an advisory guard that flags an agent hand-writing one tasks[] entry per paged row.