Rework document importer for per-page provenance-carrying extraction master
Rebuild the PDF document importer on the sol-pdf/1 worker: two extract invocations per document (authoritative text pass, then an explicit render pass), a per-page state machine, and honest per-slot failure markers. All model-generated content is blockquote-delimited so a later reader can always separate deterministic extraction from model output; verbatim text-layer pages carry zero model involvement. Rasters for every model-touched page are persisted under pages/ as model-free ground truth, and the derived transcript is written last so a crash never leaves a transcript citing a missing original or raster. Adds ImportResult.hard_failures and document-only --force plumbing: the CLI exits non-zero after a completed batch when the document importer reports hard per-file failures, and re-importing already-imported content is a safe no-op unless --force is given. Deletes the pypdf/ pdf2image extraction path and the document importer's entity seeding. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>