Insights
Data and audit trails
·
7 min read
·
How to Evaluate AI Document Extraction Tools for CRE Documents
Firms evaluating AI document extraction tools almost always grade the wrong output. They read the summary: the paragraph the model writes about the lease, the memo it drafts about the deal. That is the disposable half. The durable half is the set of normalized fields that flows into a model, a database, and a portfolio, and pays out every time someone queries it. A summary is read once and closed. A structured record is queried for the life of the asset. An evaluation that grades prose selects the tool that writes well and rejects the tool that extracts well.
Key Takeaways
Grade the structured output, not the summary. Only one of the two is an asset that survives past the first reading.
Field-level accuracy compounds against you. At 95 percent per field, a 40-field lease abstract is fully correct 12.9 percent of the time. At 99 percent per field, 66.9 percent.
A summary cannot carry a data type, a validation rule, a range check, or a page-level citation. A structured field carries all four.
The hard cases that separate real extraction from a summary dressed as data are tables and schema normalization across inconsistent source formats.
Run the evaluation on your own worst documents: scanned leases, amended rent rolls, and offering memoranda with images of tables. A clean sample set proves nothing.
What Do AI Document Extraction Tools Produce?
AI document extraction tools produce two outputs with different lifespans. One is prose a person reads: a paragraph describing a lease or a deal. The other is typed, labeled fields: base rent as a number, commencement as a date, options as a repeatable record. The first is built for one reading, the second for every system after it.
Consider a single clause: "Tenant shall pay base rent of 32.00 dollars per rentable square foot per annum, escalating three percent annually, commencing on the Rent Commencement Date." A summary restates it in a sentence. A structured extraction returns something a machine can act on.
Field | Value | Type |
|---|---|---|
base_rent_psf | 32.00 | number, USD per SF per year |
escalation_type | fixed | enum |
escalation_rate | 0.03 | number, percent |
escalation_frequency | annual | enum |
rent_start | Rent Commencement Date | date reference |
The sentence and the table carry the same facts. Only the table can be summed across a portfolio, checked against a validation rule, sorted by escalation rate, or fed into a discounted cash flow model without a human retyping it. The summary is a dead end. The table is an input.
How Do You Evaluate AI Document Extraction Tools?
Evaluate AI document extraction tools on six properties of the structured output: is every field typed, are formats normalized across source documents, are values validated against rules, does each field cite its source page, does the output load into a model without retyping, and does the tool report accuracy per field rather than as a single headline number.
Criterion | The question to ask | Why it matters |
|---|---|---|
Typing | Is base rent a number or a string? | Untyped output cannot be summed or sorted |
Normalization | Do ten source formats map to one schema? | A field only compounds if it means the same thing everywhere |
Validation | Can a range or completeness rule fire? | Errors caught at extraction cost nothing later |
Provenance | Does each field cite a page and coordinates? | Investment committees ask where a number came from |
Loadability | Does it export into the model as-is? | Retyping erases the time saved |
Accuracy reporting | Per field, with precision and recall? | A single headline number hides the fields that fail |
The last criterion carries more weight than it appears to. A tool reporting 95 percent accuracy is reporting an average, and averages conceal the distribution. Tenant name may extract at 99 percent and percentage rent breakpoints at 70 percent, and only one of those is safe to be wrong about. The right decomposition is precision and recall per field, not a single number.
Why Does Field-Level Accuracy Compound Against You?
Field-level accuracy compounds because an abstract is only useful if all of its fields are right, and independent per-field error rates multiply across the record. A tool that is highly accurate on any single field can still deliver a document-level abstract that is almost never fully correct, which is the gap between a demo and production.
Work the arithmetic. Take a lease abstract with 40 extracted fields. At 95 percent accuracy per field, the probability that every field is correct is 0.95 raised to the 40th power, or 12.9 percent. Roughly seven of every eight abstracts contain at least one error. Raise per-field accuracy to 99 percent and the document-level rate rises to 66.9 percent, still leaving a third of abstracts with an error. At 99.9 percent per field, 96.1 percent of abstracts are clean.
Per-field accuracy | 40-field abstract fully correct |
|---|---|
95 percent | 12.9 percent |
99 percent | 66.9 percent |
99.9 percent | 96.1 percent |
The conclusion is not that extraction fails. It is that any honest workflow routes uncertain fields to a human, which is why a per-field confidence signal matters more than a headline accuracy claim. A tool that cannot say which fields it is unsure about forces you to check all of them, and a confidence score decides when a human is needed.
What Are the Failure Modes on CRE Documents?
Extraction on commercial real estate documents fails in five predictable places, and none of them show up in a demo built on clean, text-native samples. Test each one against your own documents: a scanned lease with handwritten initials, a rent roll exported from three different systems, and an offering memorandum where the financial summary is an image.
Failure mode | What it looks like | Test |
|---|---|---|
Table collapse | A rent roll grid becomes a paragraph | Extract a 40-line rent roll, check row integrity |
Schema drift | Same field named differently per source | Feed three rent roll formats, compare outputs |
Amendment blindness | Original terms returned, amendment ignored | Supply a lease plus two amendments |
Scan degradation | Accuracy falls on image-based PDFs | Run a scanned original, not a clean export |
Silent confidence | Wrong values returned with no flag | Seed a known error and see if it is caught |
The first two are the hard cases that separate real structured extraction from a summary dressed up as data. A rent roll's grid of tenants and terms has to survive as rows and columns, which is where table extraction most often fails. And a field only compounds if it means the same thing across every document that feeds it, which is the work of normalizing many rent roll formats into one schema. A tool that writes a fluent summary and fails both is optimizing the output you throw away.
Why Does Structured Data Compound While a Summary Depreciates?
Structured data compounds because it is reusable across questions and systems. A summary is written for a single reading and loses value the moment that reading ends. A normalized field enters a database and answers a new query every quarter. Prose answers the one question it was written for and then sits inert in a folder.
Take a firm screening 400 offering memoranda a quarter. Assume each yields roughly 40 fields worth capturing, and that an analyst retypes those fields into a model at a fully loaded 60 dollars an hour, spending 45 minutes per document on entry. That is 300 hours a quarter, about 18,000 dollars, spent moving numbers that already existed from one format to another. If the only output is a set of models, the structured data is discarded at the end of each deal. Next quarter the same fields get retyped from scratch.
Store the same 400 extractions as structured records and the 300 hours still happen once, but the output persists. Rent comps, expense ratios, and tenant rosters accumulate into a dataset the firm queries the next time a deal in that submarket lands. The labor was identical. Keeping the structured output is what converts a recurring cost into a compounding asset.
Frequently Asked Questions
What should you look for in AI document extraction tools?
Typed fields, normalization across source formats, validation rules, page-level provenance, direct loadability into your model, and per-field accuracy reporting with confidence scores. If the primary deliverable is a paragraph a person reads, the tool is selling the summary and treating the structured data as a byproduct.
How accurate do document extraction tools need to be?
Accurate enough that the document-level result is usable, which is a higher bar than per-field accuracy suggests. Across 40 fields, 95 percent per field yields a fully correct abstract 12.9 percent of the time. Per-field confidence scoring matters more than the headline number, because it routes the uncertain fields to a person.
How should you test a document extraction tool?
On your own worst documents. Use a scanned lease with amendments, rent rolls exported from several systems, and an offering memorandum whose financials are images. Seed a known error and check whether the tool flags it. Clean vendor samples measure the sample, not the tool.
Is the AI summary useless?
No. A summary orients a reader quickly and has real value at the moment of reading. The argument is about durability. The summary is consumed once, while the structured data is queried for the life of the asset, so the structured output is what an evaluation should grade.
Conclusion
The summary is the part of an extraction tool that demos well, because it is legible to a human in the first ten seconds. That is exactly why it is the wrong thing to grade. Legibility to one reader at one moment is not the same as value to a firm over the life of an asset.
Run the evaluation on the structured layer. Check that fields are typed, normalized, validated, and cited to a page. Check accuracy per field rather than in aggregate, and do the multiplication across the fields you need correct. Then run it against the ugliest documents in your own archive. A firm that keeps the summary and discards the structured record has kept the receipt and thrown away the purchase.