> ## Documentation Index
> Fetch the complete documentation index at: https://agents.nanonets.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Document Splitter

> Analyzes a PDF and produces a splitting schema (page groups, filenames, rationale).

Analyzes a PDF and produces a splitting schema (page groups, filenames, rationale). Display name **"Document Splitter"**. On by default for new agents.

The user (or a later step) reviews the schema before pages are actually split. Password-protected PDFs: `file_utils` `unlock` first.

## Authentication and enablement

No integration. Default-on.

## Inputs

* `file_url` (required): the PDF.
* `splitting_criteria` (required unless `page_ranges` is used): natural-language split rules and per-document fields to extract.
* `page_ranges` (optional): explicit spans to cut when you already know the pages — no document read, no LLM. An array of `{ doc_id?, new_filename?, start_page, end_page }`; each entry is a contiguous, 1-indexed, inclusive span. Entries with different `doc_id`s become separate PDFs; entries that repeat the SAME **explicit** `doc_id` are merged into ONE PDF holding just those pages in ascending order, which is how you cut a document whose pages are not next to each other — `[{"doc_id":"report","start_page":1,"end_page":2},{"doc_id":"report","start_page":4,"end_page":7}]` yields a single PDF of pages 1, 2, 4, 5, 6, 7 with page 3 left out. Spans sharing a `doc_id` must not overlap, and a `doc_id` that collides with another entry's positional default (`doc_<n>`) is still rejected as ambiguous. When provided, `splitting_criteria` is ignored and no page is analyzed. It is a standalone deterministic cut (no OCR, no schema generation, no field extraction), so it cannot be combined with any other splitting option — `document_id_regex`, `validation_rules`, `static_checks`, `reviewers`, `exclude_blank_pages`, or `markdown_url` — and doing so is rejected rather than silently ignored (`splitting_criteria` is the one exception: it is simply ignored). At most 500 ranges per call. Example: `[{ "doc_id": "cash_flow", "start_page": 162, "end_page": 170 }]`. Use this when an earlier step (e.g. a table-of-contents pass) has already located a section's pages.
* `document_id_regex` (optional): when each document starts on a page with a printed id (`INV-[0-9]{6}`). OCR misreads are repaired; do not loosen the pattern for them.
* `markdown_url` (optional): if `document_parser` already ran, pass its markdown to avoid a second parse. Still pass `file_url` for citations and the real split.
* `exclude_blank_pages`, `validation_rules`, `reviewers` — optional.

## Output

A schema of page groups with suggested filenames and `document_id` when regex matching is used. On the `page_ranges` path, each distinct `doc_id` is returned as one document with its `page_map`, `split_file_id`, and `split_file_url`; a document merged from several spans carries the union of their pages, sorted ascending, and keeps the position of its first span.

## Limits and side effects

* Analysis + schema generation; actual page split follows the schema path.
* `page_ranges` cuts directly (no OCR/LLM); each span is capped at 2000 pages and `end_page` must be within the document.
* Validation failures can pause the task for review.

## Expected errors

* Missing `file_url`; missing `splitting_criteria` when `page_ranges` is not used.
* `page_ranges` with a bad span (`start_page < 1`, `end_page < start_page`, span too large, or `end_page` beyond the document), a `doc_id` repeated without being set explicitly on every one of its entries, overlapping spans that share a `doc_id`, more than 500 ranges, or combined with any other splitting option (`document_id_regex` / `validation_rules` / `static_checks` / `reviewers` / `exclude_blank_pages` / `markdown_url`).
* Locked PDF.
* Unreadable file.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.