> ## Documentation Index
> Fetch the complete documentation index at: https://agents.nanonets.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Delegate to Agents

> Native tool that lets one agent, mid-task, hand work to OTHER agents in the same workspace and act on their combined results.

This document covers `delegate_tasks` (display name **"Delegate to Agents"**), the native tool that lets one agent, mid-task, hand work to OTHER agents in the same workspace and act on their combined results.

## Overview

Use this tool when an agent needs to fan work out to specialist agents and then continue once they're all done — e.g. "research these 5 tickers" → spawn one task per ticker on a research agent, wait, then summarize.

Key properties:

* **Independent runs.** Each delegated task is a fully independent run of the target agent with the prompt you provide. The parent cannot steer a child mid-run.
* **Barrier semantics (default).** By default the parent's tool call finishes only when **every** delegated task reaches a terminal state (`completed` / `failed` / `stopped` — any outcome). The parent then resumes with each child's status + output and decides what to do next.
* **Fire-and-forget (`async: true`).** When `async` is on, the parent **starts** the children and **continues immediately** in the same turn — it does not wait and never receives their results. No await group is created, so the parent stays `running`. Use this for work whose output the parent doesn't need back (kick off a long background job, trigger notifications, etc.). See [Async / fire-and-forget](#async--fire-and-forget).
* **Fan-out.** A single task entry can carry an `inputs` array to spawn one child per value (a `{{item}}` placeholder in the prompt/title is substituted per value).
* **Bounded.** Recursion depth, per-call count, and wait timeout are all configurable (see [Configuration](#configuration)). (`timeout_minutes` is ignored when `async` is on — there is nothing to wait for.)
* **Pausing tool (unless async).** `delegate_tasks` is registered as a pausing tool in the orchestrator: while its child awaits are pending, the parent stays in `waiting_for_input` and the agent loop does not advance. An `async` call creates no awaits, so the orchestrator keeps the loop running instead of pausing (see [Flow](#flow)).

It is **enabled by default** on agents that use the default tool set, but a per-agent switch can turn it off (see [Per-agent delegation settings](#per-agent-delegation-settings)).

## Per-agent delegation settings

The agent edit page (Settings tab → **Agent delegation**, `AgentDelegationProfileSection`) exposes the two directions of agent-to-agent collaboration:

1. **Delegate out** (caller) — "Automatically choose from available agents". Backed by `config.auto_delegate` (nil/absent = on). When **off**, the agent can never spawn subtasks: `agent.filterToolsByAgentConfig` strips `delegate_tasks` from its tool list, which also drops the agent roster from the decision prompt (since `canSpawnAgents` keys off that list). When **on**, the agent decides on its own whether to delegate.
2. **Receive delegations** (callee) — "Allow other agents to delegate to this one". Backed by `config.delegation_enabled` (nil/absent = on) plus the `delegation_profile` editor revealed beneath it (see [Delegation profiles](#delegation-profiles)).

These are two **separate sections** on the page (`AgentAutoDelegateSection` and `AgentDelegationProfileSection`).

### Per-task override (home task runner)

The home task runner (`/`, `TaskFeed` home mode) fuses a **"Let other agents help with this task"** toggle into the bottom of the composer (`FusedChatInput`), on by default (the legacy "Assign to an agent" item is removed — the task runs on the default agent and this toggle governs collaboration) — so an ad-hoc task can spawn any of the user's workspace + public agents out of the box (the default agent already has `delegate_tasks` + `auto_delegate` on, and the roster spans workspace + public). Turning it off sends `disable_delegation: true` on `POST /api/chat/message`, which is persisted at INSERT time into `tasks.source_metadata` (`{"disable_delegation": true}`). At decision time `agent.taskDisablesDelegation` strips `delegate_tasks` (base **and** configured instances) from that task's tool list, so it can't delegate regardless of the agent's own config. This is a **disable-only** override: leaving it on respects the agent's own delegation capability.

## Flow

The suspend/resume is built on the existing `task_awaits` + `signals` machinery; the only new signal `source_type` is `child_task`.

1. The worker executes the `delegate_tasks` step. For each (expanded) target it resolves the `agent_group_id` to the runnable target agent and **builds a child task** (`tasks` row with `parent_task_id`, `parent_step_id`, `spawn_depth = parent+1`, `source = 'spawned'`) plus its first user message — **without emitting `new_message` yet**.
2. It creates one `task_await` per child in a group (`source_type = child_task`, `correlation_key = child task id`, `strategy = all_required`, `scheduled_resume_at = now + timeout`).
3. **Only after the awaits are committed** does it emit `new_message` per child, so a fast child can't reach a terminal state before its await row exists (which would otherwise leave the parent hung).
4. The parent task is left in `waiting_for_input`. The chat feed renders a **Delegated tasks** card listing each child with a live status pill (polled every 2s).
5. When a child reaches any terminal state, `TaskTerminalNotifier` enqueues a `child_task` signal (`dedup_key = child id`, so repeated terminal fires are idempotent).
6. The `SignalMatcher` matches the signal to the parent's pending await, completes it, and once the whole `all_required` group is complete emits an `await_completed` event.
7. The orchestrator resumes the parent: it inserts ONE consolidated `task_message` summarizing every child (title / agent / status / output) and a `decide_next_action` step, then flips the task back to `running`.
8. On `scheduled_resume_at` the `AwaitTimeoutProcessor` re-checks each child's liveness. A child that is still **non-terminal** (running, or parked in `waiting_for_input` — which a user can still answer on the child's own task page) is **not** timed out: its await is rescheduled one window forward so the parent stays paused. The parent therefore resumes only once every child reaches a terminal state (via the `child_task` signal). This guarantees a parent never completes while a delegated subtask is still pending.

```mermaid theme={null}
sequenceDiagram
    participant W as Worker (parent)
    participant D as DB / Orchestrator
    participant C as Child tasks
    participant M as SignalMatcher

    W->>D: BuildChildTask × N (tasks + first messages, NO new_message)
    W->>D: CreateGroup(task_awaits, source_type=child_task, all_required)
    W->>C: EmitChildTaskStart × N (new_message)  %% only after awaits exist
    Note over W,D: Parent → waiting_for_input. Feed shows "Delegated tasks" card.

    C->>D: each child runs independently → terminal (completed/failed/stopped)
    D->>M: TaskTerminalNotifier enqueues child_task signal (dedup_key=child id)
    M->>D: FindPendingAwait(child_task, child id) → CompleteAwait
    M->>D: group all_required complete? → emit await_completed
    D->>W: insert consolidated summary message + decide_next_action; status → running
```

### Async / fire-and-forget

When the call sets `async: true`, the parent does **not** block on the children:

1. The worker still **builds each child task** and emits `new_message` (steps 1 + 3 above) — the children run exactly as normal.
2. It **skips the await group** (step 2) entirely — no `task_await` rows, no timeout, no resume path.
3. The tool returns a result with `status: "spawned_async"` (and empty `group_id` / `strategy`). The orchestrator's `isNonPausingDelegate` helper reads that status and treats the completion as **non-pausing**: it creates the next `decide_next_action` step and leaves the task `running`, so the agent loop continues in the same turn with the children already started.

The children are normal spawned tasks (`source = 'spawned'`, `parent_task_id` set), so they still appear in the **Delegated tasks** card, the task list "↳ subtask" badge, and the child's "Delegated by …" back-ticker — the parent just isn't waiting on them. Because no await exists, each child's terminal `child_task` signal lands as `no_match` (harmless, the same path as a stopped parent).

## Tool input

```jsonc theme={null}
{
  "tasks": [
    {
      "agent_group_id": "…", // a workspace agent (not self, not another workspace)
      "prompt": "Process {{item}}", // {{item}} substituted per input when fanning out
      "title": "optional",
      "inputs": ["c1", "c2", "c3"], // optional fan-out → one child per value
    },
  ],
  "timeout_minutes": 60, // optional; capped by config; ignored when async is true
  "async": false, // optional; true = start the tasks and continue without waiting (default false)
}
```

* **No `inputs`** → one child per task entry.
* **`inputs` set** → one child per value; `{{item}}` in `prompt`/`title` is replaced (if absent, the value is appended). The per-call limit applies to the **total expanded count**.
* **`async`** → `false` (default) waits for all children and returns their combined output; `true` starts them and continues immediately without their results. Exposed to the LLM, and surfaced in the configured-tool modal as a toggle (**"Run in the background (don't wait for results)"**, `x-display-name`) that an author can pin to a static value.

### Examples

**1. Wait for results (default).** Hand a sub-analysis to a specialist agent and use its output:

```jsonc theme={null}
{
  "tasks": [
    {
      "agent_group_id": "9c1b…",
      "title": "Q3 competitor pricing",
      "prompt": "Research current list pricing for Acme, Globex and Initech and return a short comparison table.",
    },
  ],
  "timeout_minutes": 30,
}
```

The parent pauses in `waiting_for_input`; once the child finishes it resumes with a consolidated summary message and decides what to do next.

**2. Fan-out and wait.** Run the same agent over a list (one child per value):

```jsonc theme={null}
{
  "tasks": [
    {
      "agent_group_id": "9c1b…",
      "title": "Research {{item}}",
      "prompt": "Research the latest financials for ticker {{item}} and return revenue, margin and a one-line outlook.",
      "inputs": ["MSFT", "GOOG", "AMZN"],
    },
  ],
}
```

Spawns 3 children; the parent resumes only after all 3 reach a terminal state.

**3. Fire-and-forget (async).** Kick off background work whose result the parent doesn't need:

```jsonc theme={null}
{
  "tasks": [
    {
      "agent_group_id": "4f7e…",
      "title": "Re-index knowledge base",
      "prompt": "Re-crawl the docs site and refresh the vector index. No reply needed.",
    },
  ],
  "async": true,
}
```

The tool returns `status: "spawned_async"` and the parent continues in the same turn — e.g. it can immediately reply to the user "I've kicked off the re-index in the background" without waiting for it to finish.

## Configuration

### Runtime bounds (`backend/config.yaml`, env-overridable)

```yaml theme={null}
spawn_tasks:
  max_depth: 3 # SPAWN_TASKS_MAX_DEPTH — caps recursion (top-level=0)
  max_children_per_call: 250 # SPAWN_TASKS_MAX_CHILDREN_PER_CALL — caps one call's children (classic and source); 0 = unlimited
  default_timeout_minutes: 60 # SPAWN_TASKS_DEFAULT_TIMEOUT_MINUTES
  max_timeout_minutes: 1440 # SPAWN_TASKS_MAX_TIMEOUT_MINUTES — ceiling for per-call override
```

### Delegation allowlist (optional)

Configure a scoped instance via the standard **Create Configured Tool** modal (Tools → ⚙): pick **Delegate to Agents**, give it a name, set any parameters, and use the **"Agents this tool can delegate to"** multi-select.

* **Leave all unchecked → any workspace agent** (default).
* **Select some → allowlist.** The tool will only delegate to those agents; a task targeting any other agent is rejected at execution.

The selection is stored as the configured tool's `allowed_agent_group_ids` binding (a comma-separated list of `agent_group_id`s, `x-exclude-from-llm` so the LLM never fills it). At execution the binding is injected into the tool args and enforced by `parseAllowlist`. The decision-time roster still lists all workspace agents (discoverability); enforcement is at call time.

Because the allowlist value itself is hidden from the model, two things keep it from picking a restricted delegate tool for an agent it can't reach (which would otherwise fail at execution):

* **Description hint (routing):** `loadConfiguredTools` resolves the allowlist `agent_group_id`s to agent names and appends them to the tool's description — e.g. *"This delegate tool can ONLY delegate to: "Pricing Agent" (agent\_group\_id …). To delegate to any other agent, use the delegate\_tasks tool instead."* So the model sees, before calling, which agents each configured delegate tool can reach (`agent.describeDelegateAllowlist` / `buildAgentNameByGroup`).
* **Actionable rejection (recovery):** if it still targets a disallowed agent, the execution error names the allowed agents and points to the open `delegate_tasks` tool (`spawn_agent_tasks.allowedAgentNames`), so it recovers immediately instead of leaving a failed step.

## Discoverability (how the agent knows who to delegate to)

When `delegate_tasks` is enabled, the decision-time system prompt gets an injected roster, parsed as its own **"Delegatable Agents"** section in the debug view:

```
# Agents you can delegate to
You can delegate tasks to the agents below using the delegate_tasks tool. ...
- Research Agent — agent_group_id: <uuid>
    Description: Researches a stock ticker and returns financials.
    Example input: { "ticker": "MSFT" }
- Summarizer — agent_group_id: <uuid>
```

The current agent is excluded (no self-delegation), as are agents with delegation switched off or with no profile description (see [Delegation profiles](#delegation-profiles)). Agent authors can also `@`-mention agents in the instructions editor; resolution is name-based against this roster.

In the admin **decision-log debug view** (`DecisionLogDetail`), this roster renders under the System Prompt breakdown with **one row per delegatable agent** — the same way each tool schema is its own row — so individual agents can be inspected separately (`parseDelegatableAgents` splits the section on the `- <name> — agent_group_id:` bullets).

### Delegation profiles

Each agent can have an LLM-generated **delegation profile** — a short `description` + `example_input` — stored at `config.delegation_profile` and surfaced under each roster entry so `decide_next_action` can pick the right agent.

**Roster eligibility (who shows up as a delegation target).** An agent appears in another agent's roster only when **both** hold (`agent.buildAvailableAgentsSection`):

1. **Delegation is on** — `config.delegation_enabled` is unset or `true` (the default). Setting it to `false` removes the agent from every roster.
2. **It has a profile description** — an agent with no `config.delegation_profile.description` is skipped, since the decider has nothing to match on. (Delegation is allowed by default, but undescribed agents are invisible until a profile exists.)

The roster spans the caller's **workspace agents plus published public agents** (`is_public = TRUE`, from `ListPublicAgents`) shared from other workspaces. Public agents are tagged `(public)` in the prompt and badged **Public** in the debug view. Workspace agents take precedence over a public duplicate (same `agent_group_id`), shown once and untagged. The current agent is always excluded (no self-delegation).

**Cross-workspace execution.** Delegating to a public agent creates the child **in the caller's workspace** (`BuildChildTask` sets the caller's `workspace_id`) — only the shared agent *definition* (config/prompt/tools) is reused; the child's data, integrations, and credentials resolve from the caller's tenant, so there's no cross-tenant data exposure. At execution the tool's allowed-target set is the union of `agentRepo.List(workspaceID)` and `agentRepo.ListPublicAgents()` — both trusted server-side queries; the LLM-supplied `agent_group_id` still can't widen it. Eligibility only gates the decision-time roster; the allowlist + self checks are enforced independently at execution.

* **Generate per agent:** the agent's **Settings** tab has a "Delegation profile" section (`AgentDelegationProfileSection`) with a master **Delegatable** toggle (on by default), a **Generate with AI** button, and editable description/example saved on blur. `delegation_enabled` persists via the standard agent update (`config.delegation_enabled`).
* **Backfill all:** `POST /api/agents/delegation-profiles/backfill` (workspace-scoped; `?force=true` regenerates existing ones) fills in any agents missing a profile.
* **Single endpoint:** `POST /api/agents/:id/delegation-profile` generates + saves one.
* Generation: `services.AgentProfileService.GenerateDelegationProfile` prompts an LLM (gemini-2.5-flash) with the agent's name, instructions, and tools; prompt text lives in `prompts/files/agent/agent_delegation_profile_{system.txt,user.tmpl}`. The call sets `DisableReasoning: true` so thinking tokens don't starve the small output budget (a thinking model otherwise returns truncated/empty JSON).

## UI

* **Parent task feed:** a "Delegated tasks" card lists each child with a live status pill ("3 of 4 complete"); deleted children show as `deleted`. Backed by `useChildTasks` → `GET /api/tasks/query?parentTaskId=…`. For an async delegation (result `status = spawned_async`) the header reads **"Started tasks"** instead, since the parent isn't waiting on them.
* **Child task feed:** a floating "↳ Delegated by \<Agent>" ticker links back to the parent (`feed` response carries a `parent` ref with the parent's agent id).
* **Task list:** spawned children appear in the main list with a "↳ subtask" badge (click → filter to that parent's children). `source=spawned` is included in the default view. A **parent** row shows an "N subtasks" toggle (`TaskTitleCell`). Clicking it expands the delegated-subtasks list (`SubtaskList`) as its **own full-width row** directly beneath the parent — so the parent row's cell heights and alignment never change when it opens. Each child shows its live status, the **agent name** it was delegated to, and a link to its task page; the list loads lazily and live-polls only while open. Expansion state lives in `TasksTable` (via table `meta.subtaskExpansion`). The count comes from `subtask_count` on the task-query response (a correlated `COUNT(*)` of non-deleted children, workspace-pinned).

## Stopping & deleting

Because children are independent runs, stopping/deleting a parent prompts to cascade:

* **Stop** (chat): if the task has delegated subtasks, a dialog offers "Stop only this" vs "Stop this & N subtasks". Cascade stops each active descendant, cancels their steps **and their pending awaits** (so an in-flight child signal can't resume a just-stopped task — `delegate_tasks` does not fire wake-up signals during teardown). Stopping an already-terminal parent with the cascade flag still stops orphaned children.
* **Delete** (single row + bulk): same prompt; cascades a soft-delete to descendants (`include_subtasks`). Note: delete is soft-delete only — it hides children but does not halt ones still executing.

## Edge cases & guarantees

* **Self / same-group delegation** is rejected (infinite-loop guard), as is any `agent_group_id` outside the workspace (tenant isolation / BOLA) or outside the configured allowlist.
* **Recursion** is capped by `max_depth`; a child may delegate further only until the cap.
* **Duplicate resume** is prevented by `dedup_key = child id` on the signal + the race-safe `CompleteAwait` CAS.
* **Spawn↔await race** is closed by committing all child awaits before emitting any child's `new_message`.
* **Parent stopped/deleted while children run:** the parent's awaits are cancelled, so a child's terminal signal finds no pending await (`no_match`, harmless).
* **Async (`async: true`):** no await group is created, so the parent never blocks and the children's terminal `child_task` signals always land as `no_match` (harmless). The parent gets no results back and cannot be resumed by the children — choose this only when the output isn't needed. `timeout_minutes` is ignored. The recursion-depth, allowlist, self / cross-workspace, and per-call-count guards all still apply (async changes only the wait, not who can be spawned).

## Key code

| Concern | Location |
| - | - |
| Tool (`delegate_tasks`) | `backend/internal/tools/spawn_agent_tasks/` |
| Child create / emit split | `services.TriggerService.BuildChildTask` / `EmitChildTaskStart` |
| Terminal → `child_task` signal | `services.TaskTerminalNotifier`, wired in `orchestrator.fireOnTaskTerminal` + API stop/cancel |
| Resume (consolidated summary) | `orchestrator.handleAwaitCompletedEvent` |
| Non-pausing delegate detect | `orchestrator.isNonPausingDelegate` (`delegate_pause.go`) — keeps the loop running for `spawned_async`, `batch_already_complete`, `batch_no_items` |
| Timeout | `orchestrator.AwaitTimeoutProcessor.completeChildTaskAwait` |
| Roster injection | `agent.buildAvailableAgentsSection` (+ `prompts/files/agent/agent_system_prompt.tmpl`) |
| Parent/children queries | `repositories.TaskRepository` (`ListChildren`, `GetChildStatusSummary`, `ListDescendantIDs`, `CountChildren`, parent filter) |
| Allowlist | `allowed_agent_group_ids` configured-tool binding → `tool.parseAllowlist` (enforced in `Execute`) |
| Frontend | `SpawnedTasksFeedItem` (async header), `AgentAllowlistField` + `BindingInput` boolean toggle (in `ConfiguredToolModal`), `useChildTasks`, `TaskFeed` ticker, `TasksTable` badge/filter |
| Migration | `backend/migrations/162_add_parent_task_id.{up,down}.sql` (`parent_task_id`, `parent_step_id`, `spawn_depth`) |

## Source mode — deterministic fan-out (`source`)

Setting `source` INSTEAD of `tasks[]` switches this tool to deterministic fan-out: code pages a source tool and creates exactly one child task per item. Always synchronous, single-target, mutually exclusive with `tasks[]`/`async`.

**Rollout is per agent.** The `source` args carry `x-exclude-from-llm` in the shared tool definition, so no agent's model sees them by default. Toggling **"Process whole lists reliably (deterministic delegation)"** on the agent's Settings tab (or setting `config.settings.delegate_source_mode: true`) reveals the group (and its guidance) to that agent only, and execution re-checks the same setting fail-closed. The `batch_fanout.enabled` runtime flag remains the fleet-wide kill switch. Piloting on one agent in prod = arm the flag once, then flip the setting on that agent; unticking it reverts instantly.

### Overview

`delegate_tasks` asks the LLM to enumerate items and hand-type one child per input, which drops work at batch scale (in a production run, 247 encounters became 63 children — the model fetched one page and enumerated part of it). Source mode moves counting and hand-off into code. The model supplies **intent** — which source lists the items, which agent handles each one — never the item list itself.

Key properties:

* **Complete by construction.** The service loops every page of the source tool and creates exactly one child task per item. A missing/empty item key fails the whole batch rather than silently dropping the item.
* **Idempotent re-runs.** Each child carries a stable `item_key` (from `item_key_field`), unique per (parent task, target agent group, source) over non-deleted rows (`uq_tasks_parent_source_item_key`). Re-running the same parent skips items whose child is already terminal (`skipped`) and re-awaits ones still running (`attached`). A re-awaited child is **moved onto the new run** (`batch_run_id` is reassigned), so the run that is waiting on it also owns it in the rollup and the feed card — the superseded run keeps only the children nothing re-claimed. A fresh parent task may process the same items again.
* **Derived rollup.** A `batch_runs` row records what was enumerated/enqueued (`total_items`, `created_children`, `skipped_count`, page progress; status `running` → `enqueued`/`failed`). Child outcomes are always derived live from `tasks` (`GROUP BY status WHERE batch_run_id = …`) — no maintained counters to drift.
* **Barrier semantics.** The parent pauses on an `all_required` await group (same machinery as `delegate_tasks`) and resumes with one compact rollup message — totals plus only the non-completed children enumerated, so a 250-item batch does not flood the parent's context.
* **Bounded.** `max_pages` (default 100), `max_items` (default 250, hard cap 1000, operator-clampable), spawn-depth guard, self-spawn block, an optional config-time target allowlist (`allowed_agent_group_ids`, same contract as `delegate_tasks`), and the standard delegation wait timeout. Exceeding a cap **fails the batch — it never truncates**. Credentialed sources work: the runner reuses the WORKER's own step resolver (configured bindings, integration discovery, secret injection), injected late-bound, so a source call resolves exactly as a real step does.

## Authentication and enablement

* No external authentication; the tool is native and works entirely against the platform database plus the (already-enabled) source tool.
* Per-agent opt-in via `config.settings.delegate_source_mode` (invisible and non-executable otherwise), plus the fleet-wide `batch_fanout.enabled` runtime flag.
* **Armed by a runtime flag**: `batch_fanout.enabled` (seeded off, read fail-closed). Until an operator enables the flag, the tool returns an error result and creates nothing. Flip it with the admin runtime-flags API/UI.

## Inputs

| Arg | Required | Description |
| - | - | - |
| `source_tool` | yes | Name of the tool that lists the items (e.g. `nanonetshealth_get_appointments`). Accepts a base tool name (must be enabled for the calling agent — fail-closed) or a **configured tool** display name (authorized via its agent/workspace-scoped row); tools that spawn or wait rather than list (`delegate_tasks` itself, `run_code`, and the await family) are never accepted as sources. The tool's schema must declare the page/page-size args and it must return an items array. |
| `source_args_json` | no | JSON object of fixed args for the source tool (filters, date ranges). Do not include page/limit. |
| `items_path` | yes | Dot-path to the items array inside the source result's structured content (e.g. `data` or `data.results`; `.` = the result itself is the array). |
| `item_key_field` | yes | Field on each item that uniquely, stably identifies it (e.g. `appointment_id`). Values are capped at 512 chars (they land in a unique index and in child titles); every key is validated during enumeration, **before** any child is created. |
| `target_agent_group_id` | yes | Agent to run once per item. Must be an available workspace/public agent and not the calling agent. |
| `prompt_template` | yes | Child instruction; `{{item}}` is replaced with the item's key (appended if omitted). |
| `title_template` | no | Child title template; defaults to the item key. |
| `include_item_json` | no | When true, `{{item_json}}` is replaced with the item's full JSON. **PHI caution below.** Default false. |
| `page_param` / `page_size_param` | no | Source arg names for pagination. Defaults `page` / `limit` (1-based pages). |
| `page_size` / `max_pages` / `max_items` | no | Paging knobs; defaults 100 / 100 / 250, hard ceilings 500 / 1000 / 1000 (values above a ceiling are clamped down to it, not rejected). `max_items` is further clamped by the operator's `spawn_tasks.max_children_per_call`. Exceeding the *effective* cap fails the batch — it never truncates. |
| `timeout_minutes` | no | Max wait for children; hard ceiling 10080 (7 days), then clamped to the delegation config (`spawn_tasks.max_timeout_minutes`), default from `spawn_tasks.default_timeout_minutes`. |

## Outputs

Structured content on the feed card:

* `status` — `awaiting_child_tasks` (parent paused), `batch_already_complete` (every item already handled; failures called out explicitly), or `batch_no_items` (the source returned 0 items — named distinctly so wrong filters are never mistaken for a finished batch).
* `batch_run_id`, `group_id`, `total_items`, `created_children`, `skipped_terminal` (skipped items whose outcome is already **decided** — deliberately excludes `other_run`, which is live work under a concurrent run and must not read as finished), `skipped` (split by what happened to each skipped item: `completed` / `failed` / `stopped` / `other` = its child row is gone / `other_run` = a **concurrent run of the same batch claimed it** — live work, not lost), `attached_active`, `start_failed` (children that could not be woken; their awaits complete as `start_failed` so the parent never hangs).
* `children` — a bounded echo (first 25) of child refs; the full set lives in `tasks` (`batch_run_id = ...`).

On resume, the parent receives one consolidated message: batch totals plus per-child lines only for failed/stopped/unfinished children (with structured output when present).

## Source-tool authorization

`source_tool` runs with the **calling** agent's authority, checked fail-closed before the first page. A configured tool authorizes through its own agent/workspace-scoped row; a base tool must be granted to the agent by the same rules the agent loop applies — its tool list plus phase-declared tools, or, when the tool list is empty, the default-enabled set only. **CodeAct agents never get the default-enabled fallback**: their grant is always tool list + phase-declared tools + `run_code`, however empty the list. **Default-template agents** additionally get every default-enabled tool unioned in, on every branch, exactly as the loop does. `delegate_tasks` itself can never be a source, including via a configured alias.

`source_args_json` is **filtered to the source's declared schema properties** before anything is resolved. The tool itself ignores undeclared arguments, but the credential resolver reads some by fixed name — `python_code_tool`'s merged input takes `args["integration_id"]` and resolves it through an unscoped lookup — and those names are filled server-side from configured-tool bindings, never by the model. Dropped keys are logged. Server-resolved bindings are merged *after* the filter, so they are unaffected.

## Not callable from `run_code`

Source mode pauses the calling task on an await group, which in-sandbox code cannot do — `run_code` would carry straight past the pause and the group's later completion would force a decide step against a mid-flight task. A `delegate_tasks` call carrying `source` is therefore refused from CodeAct with a message pointing at the agent loop. The `tasks[]` shapes (including `async: true`) stay callable from code as before. The check runs on the *resolved* base tool, so a configured alias can't route around it.

## Paging contract

The paging arg names are checked against the source's input schema before the first call: **declared, numeric, distinct, and not statically bound**. An undeclared arg would be silently dropped and mis-page the source; a bound one would repeat a single page; a non-numeric or duplicated name means the argument is not a pagination control at all (pointing both names at one property would otherwise let any tool qualify as "pageable"). Nothing can prove a declared `page` argument really pages; the enumeration rules below are the actual guard. The service then pages 1, 2, … and every item must carry `item_key_field`.

Termination — a **short page is never treated as the end**, because a source that honours `page` but ignores `limit` returns short pages forever and stopping at the first one would drop the rest. Enumeration ends when:

* a page comes back **empty**; or
* a page repeats **every** item already seen — the source ignores paging entirely (`fetch_all`-style argument, or no paging support) and hands over its whole result set per call. That set is accepted as complete **only if it was short**; a repeated **full** page fails the batch, since the source may be capping at `page_size` with rows that can never be reached.

**Partial** overlap between pages is neither case: it means `item_key_field` isn't unique per item, and it **fails the batch** rather than silently collapsing it.

## Side effects

* Creates one `tasks` row + first user message per item (source `spawned`, `item_key`/`batch_run_id` set), a `batch_runs` row, and an `all_required` await group on the parent.
* Child tasks bill their own steps; the fan-out call itself is free.
* Children currently bypass agent-group admission control (same as `delegate_tasks`); `max_items` is the guardrail. Run large batches off-hours or against a dedicated agent group.

## Mixed parents (classic + batch on one task)

Resumes are symmetric: a **batch** group's summary covers only its own run's children, and a **classic** group's summary excludes batch children entirely — each batch reports through its own resume, so neither direction re-renders the other's results into the parent's context.

## Superseded runs

A re-run cancels every **pending** await left by a previous run **of the same step** — matched on `batch_run_id` in the await context, and on the owning `batch_runs.parent_step_id`. This covers the crash-mid-emit case *and* items the earlier attempt awaited that the new enumeration no longer returns, which would otherwise sit pending until the timeout processor completed them and fired a resume at an already-busy parent.

The step match is what makes the cancel safe: two runs of the *same* step are a retry or a reclaimed step, where superseding is the intent, while two runs from *different* steps are separate live delegations whose awaits must survive. A concurrent `delegate_tasks` delegation carries no `batch_run_id` at all and is never touched. `task_awaits` has no `workspace_id` column, so the statement asserts tenant scope with an `EXISTS` against the owning task.

## Concurrent runs on one parent

Two overlapping source-mode calls on the same parent are **not** a supported shape in v1. The supersede is scoped to the same step, so a second run from a *different* step leaves the first run's pending awaits alone — deliberately, so a genuinely live sibling isn't orphaned. The consequence is that if the second run dedups onto a child the first run is still awaiting, committing its await group hits the pending-unique index and **fails that whole batch loudly** (nothing partially emitted, `batch_runs` marked `failed`, safe to re-run once the first finishes).

That is the intended trade-off: fail a duplicate run rather than silently steal a child from a live one. A graceful "another run owns it" skip needs one-in-flight-run-per-parent enforced at the step/event layer, which is a filed follow-up.

## Failure & resume

* **Pagination/enqueue failure**: the `batch_runs` row is marked `failed` with the error and `current_page`. **No delegated task is started**, but children built before the failure are already committed as rows — inert, with no await and no start event, so none of them runs. Re-invoking with the same parameters is safe and is how they heal: dedup adopts those rows (re-attaching or skipping them) instead of duplicating, and the remaining pages fill in. The tool's error text says exactly this rather than claiming nothing was created, and a unit test pins the end-state (rows committed, no await group, nothing emitted). No reaper collects them — a re-run is the intended recovery.
* **Child failures**: visible in the resume rollup (derived from the child tasks) and, on re-runs, in the `skipped.failed` split. No automatic per-item retry yet (see non-goals) — re-run the batch to mop up, or handle in the parent.
* **Interruption (deploy, scale-down, shutdown drain)**: when the fan-out's own context dies mid-start, the run **stops** rather than mis-recording the remaining children — they stay pending with pending awaits, the run is marked failed as interrupted, and the swept step's automatic retry adopts every child already created (started ones are not re-woken). `start_failed` is reserved for a child the platform genuinely refused, or whose row is gone.
* **Attached child already terminal**: a re-attached child that went terminal before the new await group committed can never re-fire its signal, so its await is completed inline — including when the batched reconcile read fails and the per-child fallback read discovers the terminal state. A **transient** read failure leaves the await to the child's signal or the timeout processor's liveness re-check (the child was live at prefetch); only a child whose row is genuinely **gone** is settled as `start_failed`, so a possibly-running child's outcome is never falsified.
* **Timeout**: `timeout_minutes` is a liveness window, not a deadline — at each expiry still-running children are re-checked and the wait extends; the parent resumes only when every child is terminal (incl. `start_failed`).
* **Unreadable await group on resume**: the rollup is scoped by the completed await's group and its message id is derived from that group, so if the await or group can't be read the summary would be both unscoped and un-deduplicated. The event is requeued instead — for **15 seconds**, since `orchestrator_events` has neither an attempt counter nor a scheduled-at column, so a requeue is re-picked on the next tick with no backoff. Past that window (or on a non-retryable failure) the parent resumes with the **unscoped summary** — for a classic parent that is exactly what it always received; a batch parent gets every batch's rollup block rather than only its own (bloated but truthful). Only if even that read fails does a one-line go-read-your-subtasks note go out, so the parent never wakes with no information at all. The message is keyed to the await group where one is known — and to the **event id** when the await itself is unreadable — so a crash-reclaimed retry of the same event can't post it twice.
* **Skip & send now**: skipping the wait relabels the feed card `skipped` (standard `AwaitSkipper` behavior); children keep running.

## Expected errors

* `Deterministic fan-out is not enabled on this deployment (runtime flag "batch_fanout.enabled" is off)` — ask an operator to arm the flag. A transient flag-read failure returns `Could not verify the ... runtime flag` instead (fail-closed, retryable).
* `source tool "X" is not enabled for this agent` — the source tool isn't in the calling agent's tool list (authorization is fail-closed; enable the tool on the agent first).
* `source tool "X" is not available` — the source tool isn't registered.
* `items_path ... not found / does not point to an array` — wrong `items_path` for this source tool's result shape.
* `item is missing item_key_field ...` — wrong `item_key_field`; the batch fails rather than dropping the item.
* `item key "X" repeated on page N — item_key_field ... is not unique across items` — pick a per-item-unique field; a source that ignores paging is detected separately (see the paging contract).
* `source has more than N items (max_items)` / `more than N pages` — deliberate anti-truncation guardrails; raise the cap or narrow the source filters.
* `An agent cannot fan out tasks onto itself.` / max spawn depth — same delegation guards as `delegate_tasks`.

Deliberately **not** errors, because weak models ping-pong between them instead of converging: over-large `max_items` / `max_pages` / `page_size` are clamped to the hard maxima (safe — a source exceeding the effective cap still fails the run, so nothing is silently truncated), and a `tasks[]` array sent alongside `source` is dropped with a warning on the result rather than rejecting the call.

## PHI note

Keep child prompts to IDs: the default `{{item}}` substitution passes only the item **key**, and the child fetches full details itself at run time. `include_item_json` copies entire item payloads into child prompts/messages — leave it off for EHR/PHI sources.

## Feed card at batch scale

The tool result echoes at most **25** children (`maxChildEcho`) — that result is replayed into every later LLM call on the task, so a 250-item echo would cost context on every step. The card therefore shows the echoed rows plus an honest `N+ of M complete` lower bound and a "Showing 25 of M" note.

**View All N Tasks** expands the card to every child of that batch, read from the live child-task query. That query is scoped **server-side** by `batch_run_id` (an optional `batchRunId` filter on the task-query endpoint, additive like `parentTaskId` with the workspace still pinned, applied in **both** the enriched and simple query paths and echoed back as `batch_run_id` on list items), so each card polls only its own children: a later, larger batch on the same parent cannot push an earlier batch's children out of the window and freeze that card on stale statuses. Expanded, the progress count is exact.

## Dedup namespace

The unique index is `(parent_task_id, agent_group_id, batch_source_key, item_key)`. For a **base tool** source, `batch_source_key` is the resolved base tool name; for a **configured tool** it is the base name plus the configured tool's **row id** (`base#<uuid>` — the id, not the display name, so a rename doesn't orphan earlier children). Two dimensions of collision are closed: different sources from one parent (`get_appointments` then `get_claims`, both carrying id `101`) never share a namespace, and neither do two configured instances of one base tool bound to **different accounts** — without the instance id, an `item_key` collision across accounts would classify the other account's unprocessed item as "already done" and silently skip it. Re-running the *same* source still dedups, because the key is stable.

One consequence to know: running a source via its configured alias and then via the bare base tool are **different namespaces**, so items may be processed twice across that switch. That direction is chosen deliberately — duplicate work is visible and recoverable, a silently skipped item is not.

## Migrations & lock profile

Seven pairs ship with source mode: `batch_runs`, the three `tasks` columns (metadata-only ALTERs), the dedup unique index (its **own** migration — built in the same transaction as the column ALTER it would inherit ACCESS EXCLUSIVE and block reads too; alone it holds SHARE), the `batch_fanout.enabled` flag seed, the `batch_run_id` lookup index, the partial `task_awaits (group_id)` index (its own migration — classic delegation benefits, so a `batch_runs` rollback must not regress it), and a partial `orchestrator_events (task_id) WHERE event_type = 'new_message'` index serving the re-run don't-wake-twice check, otherwise a seq scan over \~19M rows in prod. **Prod runbook**: after the column migration lands, pre-create the two `tasks` indexes and the `orchestrator_events` index `CONCURRENTLY` — the transactional runner builds them inline under `SHARE` otherwise, which is seconds for `tasks` (\~450K rows) but not acceptable for `orchestrator_events`.

The `tasks` columns themselves are metadata-only `ADD COLUMN`s, but **the index builds are not free**. The runner is transactional, so `CREATE INDEX CONCURRENTLY` is unavailable: each build takes `SHARE` on `tasks` and scans the whole table (matching zero rows makes the index empty, not the build cheap), blocking writes for the scan — seconds at \~412K rows / 520MB, and growing with the table. `SET LOCAL lock_timeout = '5s'` bounds lock *acquisition*, not the scan; re-run on timeout. Ops may pre-create either index `CONCURRENTLY` ahead of the deploy, which makes the `IF NOT EXISTS` build a no-op. The two indexes are deliberately in **separate migrations** so each transaction holds `SHARE` for a single scan and commits, rather than one transaction spanning both. The down-migrations set the same `lock_timeout`, so a rollback fails fast instead of queueing behind live traffic.

`batch_runs` is created with a plain `CREATE TABLE` — not `IF NOT EXISTS`. The table is new in this migration, so a name collision means a database in an unexpected state and should fail loudly rather than silently keep a shape the code does not match.

## Non-goals (v1)

Per-item retry/backoff, worker batch-drain/throughput scaling, and batch-aware admission control are deliberate fast-follows, not part of this tool.

Three related changes are also deliberately **out of this PR**, each reviewable on its own: reworking the sandbox admission gate so a configured alias to a loop-control tool can't reach `run_code` (a pre-existing gap — this PR only refuses the `source` shape); bounding classic `delegate_tasks` resume summaries, which are still unbounded in child count and per-child output; and an advisory guard that flags an agent hand-writing one `tasks[]` entry per paged row.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.