file_utils unlock first.
Authentication and enablement
No integration. Default-on.Inputs
file_url(required): the PDF.splitting_criteria(required unlesspage_rangesis used): natural-language split rules and per-document fields to extract.page_ranges(optional): explicit spans to cut when you already know the pages — no document read, no LLM. An array of{ doc_id?, new_filename?, start_page, end_page }; each entry is a contiguous, 1-indexed, inclusive span. Entries with differentdoc_ids become separate PDFs; entries that repeat the SAME explicitdoc_idare merged into ONE PDF holding just those pages in ascending order, which is how you cut a document whose pages are not next to each other —[{"doc_id":"report","start_page":1,"end_page":2},{"doc_id":"report","start_page":4,"end_page":7}]yields a single PDF of pages 1, 2, 4, 5, 6, 7 with page 3 left out. Spans sharing adoc_idmust not overlap, and adoc_idthat collides with another entry’s positional default (doc_<n>) is still rejected as ambiguous. When provided,splitting_criteriais ignored and no page is analyzed. It is a standalone deterministic cut (no OCR, no schema generation, no field extraction), so it cannot be combined with any other splitting option —document_id_regex,validation_rules,static_checks,reviewers,exclude_blank_pages, ormarkdown_url— and doing so is rejected rather than silently ignored (splitting_criteriais the one exception: it is simply ignored). At most 500 ranges per call. Example:[{ "doc_id": "cash_flow", "start_page": 162, "end_page": 170 }]. Use this when an earlier step (e.g. a table-of-contents pass) has already located a section’s pages.document_id_regex(optional): when each document starts on a page with a printed id (INV-[0-9]{6}). OCR misreads are repaired; do not loosen the pattern for them.markdown_url(optional): ifdocument_parseralready ran, pass its markdown to avoid a second parse. Still passfile_urlfor citations and the real split.exclude_blank_pages,validation_rules,reviewers— optional.
Output
A schema of page groups with suggested filenames anddocument_id when regex matching is used. On the page_ranges path, each distinct doc_id is returned as one document with its page_map, split_file_id, and split_file_url; a document merged from several spans carries the union of their pages, sorted ascending, and keeps the position of its first span.
Limits and side effects
- Analysis + schema generation; actual page split follows the schema path.
page_rangescuts directly (no OCR/LLM); each span is capped at 2000 pages andend_pagemust be within the document.- Validation failures can pause the task for review.
Expected errors
- Missing
file_url; missingsplitting_criteriawhenpage_rangesis not used. page_rangeswith a bad span (start_page < 1,end_page < start_page, span too large, orend_pagebeyond the document), adoc_idrepeated without being set explicitly on every one of its entries, overlapping spans that share adoc_id, more than 500 ranges, or combined with any other splitting option (document_id_regex/validation_rules/static_checks/reviewers/exclude_blank_pages/markdown_url).- Locked PDF.
- Unreadable file.