When I build document ingestion for a client, I treat it as a pipeline of stages, each one testable on its own. The language model sits at the end of that pipeline, after every deterministic step has had its turn. I arrived at this design the practical way, by watching where accuracy came from in earlier builds.
The pipeline begins with intake. Every file gets a SHA-256 fingerprint on arrival, and the original is stored untouched. An identical fingerprint means the file is already known, so the run records it as unchanged. A new fingerprint creates a new version, and earlier versions stay in the record.
Next comes classification. Each document is scored against a registry of known types and their format variants, using title text and content markers. A renamed file classifies exactly the same way. A document that matches nothing well goes to a review queue, where a person assigns the type.
Then the system builds a text layer, page by page, so every extracted string carries its page number. A characters-per-page check flags pages that are image only, which tells us where OCR is needed. After that, a page map locates the pages where each group of fields lives. In an earlier build, locating the right pages before reading them produced the single biggest gain in accuracy of anything we tried.
Deterministic parsers handle everything numeric, dated, coded, or tabular. Every value is validated against ranges, allowed values, required fields, and sentinel values like N/A or 0000, and each field receives a status: parsed, missing, or failed.
The language model enters only now, and only for narrative fields such as a scope description or an inspector's note. It works from the located pages, with a JSON schema for its output, a temperature at or near zero, a bounded input, and one document per call. Its results carry the status LLM, which flags them for the same human review as everything else.
That review shows each extracted value beside its source page. A person confirms or corrects it, and every correction lands in a corrections table with the field, the original value, the corrected value, the user, and the time. Only confirmed values flow into the typed tables that later decisions read. Over time, the correction rate per field and document type tells us exactly which parser needs attention, and a regression set of real documents checks every parser change before it ships.
Placing the language model last gives each stage a clear job and a clear test. The parsers handle the structure with full traceability, the language model handles the prose it reads well, and the expert's confirmation turns extracted text into data the business can rely on.
Common questions
How do you extract data from PDFs accurately with AI?
Run deterministic steps first: fingerprint each file, classify it, extract text page by page, locate the pages that hold each field, and parse numbers, dates, and tables with validated rules. Use a language model only for narrative fields, then have a person confirm every value against its source page.
Where does the language model fit in document processing?
At the end of the pipeline, for narrative fields such as a scope description or an inspector's note. It reads only the located pages, returns output in a JSON schema at a temperature near zero, and its results go through the same human review as every other field.
Related reading
- Why Every Managed Intelligence Platform Needs a Decision Science Layer
- Find Your Zone: a free three-minute diagnostic
- Where Open-Weight Models Fit When the Expert's Rules Make the Decision
- Why I Run Daily AI Work on Open-Weight Models
Want AI built for your actual job? Book a discovery call.