Skip to main content
Analyzes a PDF and produces a splitting schema (page groups, filenames, rationale). Display name “Document Splitter”. On by default for new agents. The user (or a later step) reviews the schema before pages are actually split. Password-protected PDFs: file_utils unlock first.

Authentication and enablement

No integration. Default-on.

Inputs

  • file_url (required): the PDF.
  • splitting_criteria (required unless page_ranges is used): natural-language split rules and per-document fields to extract.
  • page_ranges (optional): explicit spans to cut when you already know the pages — no document read, no LLM. An array of { doc_id?, new_filename?, start_page, end_page }; each entry is a contiguous, 1-indexed, inclusive span. Entries with different doc_ids become separate PDFs; entries that repeat the SAME explicit doc_id are merged into ONE PDF holding just those pages in ascending order, which is how you cut a document whose pages are not next to each other — [{"doc_id":"report","start_page":1,"end_page":2},{"doc_id":"report","start_page":4,"end_page":7}] yields a single PDF of pages 1, 2, 4, 5, 6, 7 with page 3 left out. Spans sharing a doc_id must not overlap, and a doc_id that collides with another entry’s positional default (doc_<n>) is still rejected as ambiguous. When provided, splitting_criteria is ignored and no page is analyzed. It is a standalone deterministic cut (no OCR, no schema generation, no field extraction), so it cannot be combined with any other splitting option — document_id_regex, validation_rules, static_checks, reviewers, exclude_blank_pages, or markdown_url — and doing so is rejected rather than silently ignored (splitting_criteria is the one exception: it is simply ignored). At most 500 ranges per call. Example: [{ "doc_id": "cash_flow", "start_page": 162, "end_page": 170 }]. Use this when an earlier step (e.g. a table-of-contents pass) has already located a section’s pages.
  • document_id_regex (optional): when each document starts on a page with a printed id (INV-[0-9]{6}). OCR misreads are repaired; do not loosen the pattern for them.
  • markdown_url (optional): if document_parser already ran, pass its markdown to avoid a second parse. Still pass file_url for citations and the real split.
  • exclude_blank_pages, validation_rules, reviewers — optional.

Output

A schema of page groups with suggested filenames and document_id when regex matching is used. On the page_ranges path, each distinct doc_id is returned as one document with its page_map, split_file_id, and split_file_url; a document merged from several spans carries the union of their pages, sorted ascending, and keeps the position of its first span.

Limits and side effects

  • Analysis + schema generation; actual page split follows the schema path.
  • page_ranges cuts directly (no OCR/LLM); each span is capped at 2000 pages and end_page must be within the document.
  • Validation failures can pause the task for review.

Expected errors

  • Missing file_url; missing splitting_criteria when page_ranges is not used.
  • page_ranges with a bad span (start_page < 1, end_page < start_page, span too large, or end_page beyond the document), a doc_id repeated without being set explicitly on every one of its entries, overlapping spans that share a doc_id, more than 500 ranges, or combined with any other splitting option (document_id_regex / validation_rules / static_checks / reviewers / exclude_blank_pages / markdown_url).
  • Locked PDF.
  • Unreadable file.