Define the accepted document
In this reference design, I turn permitted uploads into a reviewed representation of text, tables, and metadata. The original stays in restricted storage for a defined period. Each document has an owner, a stable identifier, and a processing history. A parsed file is not automatically approved for distribution.
Start with an explicit set of formats and uses. A searchable passage, a financial table, and a faithful page rendering need different quality checks. State which output the pipeline promises and which document features it may omit.
Bound intake before parsing
Check the caller's authority, permitted types, file size, expanded archive size, and page limits. Compare declared type with detected content, while treating both as fallible. Give stored objects generated identifiers; an uploaded filename must not choose a storage path.
Run parsing in isolated workers with bounded memory and execution time, minimal filesystem access, and restricted networking. Reject unsupported formats and quarantine malformed inputs with a reason. Apache Tika's security model makes an essential distinction: extraction libraries do not make hostile files safe.
Measure what extraction missed
Use native text extraction where it works and route scanned pages through an OCR path. Record language, preprocessing, engine version, and which pages required recognition. Rotation, noise, resolution, and page segmentation affect results; changing these settings creates a new extraction version.
Review representative samples against the original, including headings, reading order, table boundaries, numbers, and footnotes. An OCR confidence value is evidence for triage, not proof of correctness. Mark unreadable regions and truncated pages explicitly instead of silently treating them as empty.
Keep structure tied to its source
Every extracted block retains its document version, page or sheet location, parser version, and relevant settings. Tables preserve cell coordinates and header relationships where available. Keep uncertain raw values alongside proposed normalized values so an inferred date or unit can be challenged.
Store extraction output as an immutable version. Corrections create a new version with a reason and links to the affected blocks. A text offset alone is fragile after reprocessing; downstream consumers need both the source location and the exact representation they used.
Review before publishing derivatives
A publication gate checks required quality and applies sensitivity rules before content enters shared indexes or model prompts. Omit unnecessary fields and review uncertain redactions. Extracted markup, links, and spreadsheet values remain untrusted and need appropriate handling at their destination.
Retrieval is a separate consumer with its own access checks and relevance evaluation. Successful indexing does not validate extraction accuracy. Derived copies retain permissions, version references, and retention rules; withdrawal must reach the active index and cached derivatives under an explicit deletion policy.
Reprocess without duplicate publication
Track partial extraction, quarantine age, worker failures, review backlog, and publication lag separately. Retry transient failures within a budget; repeated parser failures need investigation. Reprocessing stages a new version. Publication atomically records its approved identity and advances the active pointer only if the expected prior version still matches. A stale completion requires review instead of replacing newer content. Repeated completion messages cannot publish that version twice.
This adds storage and review work. A small collection may need only a manual intake checklist and extraction report. Automate as volume grows, while preserving visible limits: no parser handles every document perfectly.