How it works
The split is the whole design
A language model reads the documents. Ordinary TypeScript decides what is wrong with them. That boundary is what makes a report worth reading, and it is why the same package always produces the same findings.
The pipeline
Upload, or paste the email
Up to 12 files, or the customer's message pasted straight into the form. A pasted email is analysed exactly like a text file — it is a first-class input, not a fallback.
Parse — no model involved
Files are typed by their magic bytes rather than their extension. PDFs become pages and lines, spreadsheets become sheets and rows, text becomes lines, images become an image segment. Duplicate fingerprints and instruction-like text are noted here, before any model sees the package.
Extract — the model, under constraints
One extraction call per document, with the document delimited as data inside the prompt. The model returns candidate values, each with a citation: which segment, and which literal quote inside it supports the value. Images are transcribed by a vision call.
Verify — the trust boundary
Every candidate is checked. Does the segment exist? Does the quote appear verbatim in it? Is the value actually supported by that quote? Anything that fails is discarded here and never reaches the report. Part numbers are additionally sanity-checked for the shapes a real part number takes.
Check — code, not a model
Required fields, cross-checks, value canonicalisation, conflict clustering, revision grouping, duplicate detection and status derivation are all ordinary functions. A model cannot argue its way past a check, because the model is not the one checking.
Report
Server-rendered HTML, where evidence is a native
<details>element — it works with no JavaScript at all. Exports to PDF, Excel and CSV from the same report object.
Why the split matters when you are wrong
Model-based tools fail in a specific way: asked to find a problem in a document that is actually fine, they tend to oblige. Here, a package with no problems reports READY, and a report says so plainly. The status is derived from the findings, not chosen by anything that was asked to be helpful.
It also means the tool degrades honestly. If no model provider is reachable, the run is marked degraded, field extraction is skipped, and the file-structure findings that were established are still reported — with a visible notice rather than a quietly thinner report.
What the model is not allowed to do
It cannot change the extraction contract, because the contract is not in the document. It cannot invent a citation, because a citation that does not resolve to a real segment is thrown away. It cannot resolve an ambiguity, because ambiguity is reported rather than decided. It cannot see a URL, call a tool, or write to your files.
Document text is untrusted input. A customer who puts instructions in a spec sheet gets them parsed as spec text.