Skip to main content
Extracts structured fields from documents (PDF, Office, CSV, images, markdown, HTML, XML, text) using AI. Display name “Data Extraction”. On by default for new agents. This output is already presentation-ready; don’t reformat it with another AI step.

Authentication and enablement

No integration. Uses the agent’s extraction model. Password-protected files cannot be unlocked here — call file_utils with operation: "unlock" first, then pass the same file reference.

Inputs

  • file_url (required): document URL (${VAR_N} from an upload or prior step).
  • extraction_instructions (required): what to extract, in natural language.
  • output_schema — optional JSON Schema for the result shape. Set per configured tool, with two modes (see Choosing the output schema).
  • model — optional extraction-model override (configured tool).
  • temperature — optional, 0.0–2.0 (configured tool). How much the model varies its answer between runs. Blank leaves the provider’s own default in place. Set 0 when the same document must return the same result every time, such as tables and line items where a re-run silently returning one row more or fewer is the problem being solved. Provider ranges differ: Gemini accepts the full 0.0–2.0, Claude Sonnet accepts 0.0–1.0 and anything higher is clamped to 1.0, and Claude Opus rejects the parameter so the setting is ignored there.
  • seed — optional integer, 0 to 2147483647 (configured tool). Fixes the model’s random draw so an identical document and prompt replay the same output. Pair it with temperature: 0 for reproducible extraction; on its own it narrows run-to-run drift without removing it. Gemini extraction models only; other providers ignore it.
  • reasoning_effort — optional, default "" (configured tool). How deeply the model thinks before answering: "" (Automatic) leaves each model on its own default, medium (Reduced) makes it deliberate less and lowers latency at some risk to accuracy, low (Minimal) cuts thinking as far as the model allows, with the largest effect on both. Automatic is not “no thinking” — it is whatever the model does when nothing is pinned, which is what extraction has always used. This is a target level, not a limit: on a straightforward document, pinning a level can make the model think MORE than it would have, and no transport we use can express a true ceiling. How closely a model follows the setting varies by model, so the band actually sent is recorded on every call alongside the reasoning tokens spent — check it against your own traffic rather than assuming. Claude, Nanonets Lite and GLM have no thinking controls and ignore it entirely (a warning is logged). Thinking is billed as output tokens and shares the model’s output limit, so a deeper setting leaves less room for the answer. One scope note: on the spreadsheet agent lane (spreadsheet_agent_mode: on) the setting applies to each decision in a multi-turn loop rather than once per document, so its effect on cost and latency there is larger than on a single extraction call.
  • use_source_citations — optional, default off (configured tool). Changes how citations are produced for PDFs and images. Off, the model extracts values and a separate pass works out which words each value came from by searching the page text for that value — so two cells holding the same number cannot be told apart, and the highlight can land on the wrong one. On, the extraction prompt carries an index of every OCR word alongside the page images and the model reports which words it read each value from, which is then resolved directly. Citations become reproducible run to run, and repeated values keep separate highlights. The trade-off is latency: OCR has to finish before extraction starts instead of running alongside it, and the word index adds to the prompt. Falls back to the standard behaviour, with no change in output shape, when OCR is unavailable or when a very large document is re-extracted in page-range chunks.
  • spreadsheet_agent_mode — optional, default off (configured tool). Per-agent gate for the CSV/XLSX spreadsheet agent loop. on also requires the global sheet_agent.enabled runtime flag; if either gate is off, or the agent declines, today’s chunked lane runs. Has no effect on PDFs or images.
temperature, seed and reasoning_effort are sent to the provider only when set, so leaving them blank keeps the request exactly as it was.

Choosing the output schema

Each configured Data Extraction tool decides where its output_schema comes from:
  • Set a specific value for this field (default) — you write the JSON Schema once in the tool’s configuration and every run extracts exactly those fields. This is the right choice whenever the documents a tool handles share one shape.
  • Let your agent generate a value for this field — the tool leaves the schema open and the agent chooses it on each call, usually by passing along a schema an earlier step produced. Use this when one tool must handle document types whose fields genuinely differ, and no single schema covers them without becoming so broad that extraction quality drops.
Switching an existing tool to the second mode changes what the agent must supply: the schema becomes a required argument, so the tool cannot run until the step that produces the schema has run. A run that supplies no schema, or one that is not a JSON object or array, fails with an error the agent can correct and retry — it does not fall back to extracting without a schema. If the tool also uses Static checks, note that their field names can no longer be verified when you save. Nanonets normally checks each check’s target against the fields your schema declares and rejects a typo straight away; with no schema configured here there is nothing to check them against, so a misspelled field name is only discovered when the check silently matches nothing at run time. Name those fields carefully, or keep a specific schema on tools that rely on static checks. Every tool keeps its current mode until you change it.

HTML files

.html / .htm files are read as rendered text, not as markup: the page is walked into one numbered line per visible block (a paragraph, a list item, a table cell), and citations point at those lines. Script tags, event handlers and remote images are stripped before the file is read, so nothing in an untrusted page executes or phones home. The same lane reads the email body when a dedicated-inbox trigger has Attach the email body as a document enabled. That option adds email-body.html to the task — the message as written — so a value extracted from an email can be cited back to the sentence it came from, and the mail opens beside the review form. The agent’s context is unchanged: it still receives the plain-text body as before. Point file_url at email-body.html when the order details live in the covering note rather than in an attachment.

Output

Display-ready structured data, typically {value, word_id_groups} per field so citations work in the feed. The shape is the same whether or not use_source_citations is on — that setting changes where the citation comes from, not what the output looks like. Can merge/combine prior extraction results into one table when instructed.

Spreadsheets (CSV and XLSX)

Spreadsheet extraction defaults to the existing chunked LLM path. An optional spreadsheet agent lane is dual-gated and off by default (not GA):
  1. Ops arms the global runtime flag sheet_agent.enabled (fail-closed; admin Runtime Flags).
  2. The configured tool on a specific agent sets spreadsheet_agent_mode to on (config-only; default off).
When both are set, the model is shown the sheet’s cells and writes a Python program that reads the file. The program returns one of three things: the extracted data, a probe (something it wants to look at first, fed straight back so the next program can use it), or a decline. Derived values — totals, counts, averages — are always computed by the program, never typed by the model. If the agent declines, today’s chunked path runs. The workbook is read once, by the same parser the chunked path uses, so both lanes see the same text — a currency cell reads as $1,250.00 and a date as 14-Aug-26 in either. That is what makes comparing the two lanes on the same file meaningful. Code the agent writes opens the file itself and still sees the underlying numbers, so totals are computed on values, not on formatted text. Text recovered from pictures embedded in a workbook now reaches this lane too; previously only the chunked path saw it. CSV delimiters are sniffed rather than assumed, so a semicolon-separated export (common in European locales, where the comma is the decimal mark) is read as a table on both paths — it used to arrive as a single column. Generated code receives the file as a short-lived signed URL in input["file_url"]; the sandbox has no access to your storage credentials. Code that trips the sandbox’s safety scan (urllib, subprocess, eval, exec, …) is returned to the model as a fixable error and costs one Python attempt — it does not end the run. Metadata: extraction_lane, sheet_agent_model_calls, sheet_agent_probes, sheet_agent_python_runs, sheet_agent_used_python, sheet_agent_correction_rounds, sheet_agent_checks_still_failing, sheet_agent_uncovered_rows, plus sheet_agent_decline_reason and sheet_agent_decline_detail. Decline reasons include not_armed (tool mode on but runtime flag off), flag_unreadable (the flag store did not answer, so the lane failed closed — a different thing to investigate), workbook_too_large, sheets_filtered (a sheets filter is set; the agent sandbox would still see the full workbook, so the chunked lane is used instead), time_budget_spent (the lane ran past its wall-clock budget), and three that separate why a run ended: python_failed (the last program raised in the sandbox), bad_envelope (a program ran but returned something that is not the documented shape), contract_check_failed (the answer itself was wrong — it missed the schema, or its citations did not match the file) and answer_unrenderable (a well-formed answer past the size cap, which the chunked lane handles by splitting it). The per-attempt transcript is written to the worker log rather than into the result, so it does not sit in the calling agent’s context for the rest of the task. Before an answer is accepted it must agree with the file: every cited cell has to exist, a rows entry has to name real column letters with one row number per row, and a sample of the values the model says it copied is compared against those cells — per table, so a single mis-cited table is not diluted by the correct ones beside it. Dates are compared as dates rather than as text, so a column the program reformatted still matches while one cited a row out does not. A systematic mismatch — the shape a row-number offset produces — is sent back for correction instead of returned. Separately, an answer that cites nothing from a long run of rows is sent back once, asking for those rows to be extracted or for the gap to be explained — the same silent-row-loss detector the chunked path uses. If the gap remains, the result is still returned and the stretch is named in sheet_agent_uncovered_rows, since a sheet legitimately excludes totals blocks and footers. Budgets: 3 attempts — one attempt is one program, written, run, and either answering, probing or declining — plus 6 minutes of wall clock for the agent and 25 minutes for the agent and the chunked fallback together. A traceback, an answer that fails the contract, and an answer that fails the tool’s configured checks all come back as the next attempt, so the worst case is 3 model calls and 3 sandbox runs. The clock is what stops a single slow sandbox run from holding a worker; measured runs on real files finish in well under a minute. Decline → chunked lane unchanged, so exceeding the agent’s clock costs a slower answer rather than a lost one. Decisions are logged under [SheetAgent] with task/step/agent ids and budget counters. Free-text fields that may contain cell values (Python bodies, tracebacks, why, give-up reasons) are length-capped or omitted in logs and result metadata — use the transcript counters plus decline reason tokens for diagnosis, not a full replay of cell content.

Limits and side effects

  • Cannot open password-protected files.
  • Large documents may paginate internally; free-tier orgs can hit a page cap.
  • use_source_citations applies to PDF and image extraction only. Spreadsheet, markdown, HTML, XML, and plain-text lanes already have the extractor name its own source, so the setting has no effect there.
  • temperature and seed apply to document (PDF and image) extraction. Spreadsheet, markdown-OCR, HTML, XML, and plain-text extraction lanes do not read them yet.
  • No external writes.

Expected errors

  • Missing file_url or extraction_instructions.
  • Locked file (unlock with file_utils first).
  • Unsupported or empty document.
  • An HTML file with no visible text (for example a body that is only a tracking pixel) extracts nothing.
  • Missing or malformed output_schema, when the tool is set to let the agent generate it.