Insights
AI strategy in CRE
·
7 min read
·
AI Due Diligence Is a Document Problem Before It Is a Model Problem
Teams buy an AI due diligence tool, run it against a live deal, and find it disappointing. The usual conclusion is that the model was not good enough. The usual truth is that the constraint was never the analysis. AI due diligence in real estate is bounded by what arrives in the data room: scans instead of text, twelve rent roll formats instead of one, an amendment chain missing its third link, and half the estoppels landing four days before the contingency expires. No model repairs an input that was never delivered.
Key Takeaways
Diligence AI performs at the level of its worst input, not its average one. One unreadable scan in a chain of amendments invalidates the whole chain.
The first investment is a document standard and a page-level audit trail, not a better model.
Incomplete and late are the two failure modes no model addresses, because both are calendar problems dressed as data problems.
MISMO, a subsidiary of the Mortgage Bankers Association, has published commercial data standards precisely because the industry's constraint is definitional consistency rather than analytical capability.
Measure intake before measuring accuracy: what share of the file arrived machine-readable, complete, and inside the contingency window.
Why does AI due diligence disappoint on a live deal?
Because the tool is evaluated on a demo file and deployed on a real one. Demo files are text-layer PDFs, complete, and consistently formatted. A live data room is a mix of scans, photographs of scans, spreadsheets with merged header rows, and documents that reference other documents nobody uploaded. The gap is intake quality, not model quality.
The disappointment is predictable in its shape. The tool works well on the operating statements, adequately on the rent roll, and poorly on the leases, which is exactly the order of those three document types by formatting consistency. Nobody reads that pattern as a diagnosis. They read it as a verdict on the software.
JLL's Global Real Estate Technology Survey, drawn from more than 1,000 senior decision-makers across 16 markets, found 92 percent of organizations piloting AI in select real estate use cases or planning to, and 5 percent reporting they had achieved all of their program goals. That spread is not a story about models underperforming. It is a story about pilots colliding with document reality on the second deal.
What are the four ways a data room breaks a diligence model?
Data rooms fail in four distinct ways, and each one needs a different fix. Documents arrive as images with no text layer. They arrive in formats that differ from every other version of the same document type. They arrive incomplete, missing the amendment or exhibit that changes the term. And they arrive late, after the analysis was needed.
Failure | What it looks like | The fix |
|---|---|---|
Not machine-readable | Faxed and rescanned leases, photographed signature pages, image-only exhibits | Require a text layer at upload; run recognition and flag pages below a legibility threshold |
Inconsistent format | Twelve rent rolls, twelve column orders, three definitions of "occupied" | A field standard the seller's file is mapped into, not a new template per deal |
Incomplete | Amendment 3 cited in Amendment 4 and absent from the room | A completeness check that reads cross-references and lists what is missing |
Late | Estoppels landing four days before the contingency expires | A sequenced request list issued at signing, ordered by what blocks what |
Only the first is a technology problem. The other three are process problems that a technology purchase will not touch. The sequencing failure in particular is a calendar design issue, which is why the due diligence checklist is a sequencing problem rather than a list.
How much does bad document intake cost on a deal?
It costs the ability to verify, which is the entire point of the exercise. When a share of the file cannot be read, the deal proceeds on the summary rather than on the source, and the risk moves from the document into the underwriting without anyone recording that it moved.
Every figure below derives from the stated inputs. A 180-unit multifamily acquisition, 180 leases plus amendments, 620 total documents in the room. Suppose 82 percent arrive with a usable text layer, 74 percent conform to a recognizable format for their type, 91 percent of referenced documents are present, and 88 percent arrive before the day the analysis is due.
Intake gate | Pass rate | Documents surviving |
|---|---|---|
Machine-readable | 82% | 508 |
Consistent format | 74% | 376 |
Complete, no missing reference | 91% | 342 |
Delivered before the deadline | 88% | 301 |
The four gates compound. Each one looks tolerable in isolation and none is below 74 percent, yet only 301 of 620 documents, 48.5 percent, clear all four. A model with 97 percent field-level accuracy running on that file delivers verified coverage of 47 percent of the room. The model is not the binding constraint. The intake is, by a factor of roughly twenty.
Change one gate. Raise machine-readability to 98 percent by requiring a text layer at upload, and the survivors rise to 360, a 20 percent lift in verified coverage from a rule rather than from a model. No amount of accuracy improvement on a 97 percent model produces that.
Why is the audit trail the first thing to build?
Because the value of an extracted field is not the field. It is the field plus the page it came from. A number with a citation can be checked in ten seconds by a person who did not produce it. A number without one has to be re-derived from scratch, which costs more than reading the document would have.
This is the difference between output a committee can act on and output a committee has to redo. It is also the only defense against a specific failure mode: a value that is plausible, formatted correctly, and not present in the source document. Grounding every field in a cited page is what makes that failure visible instead of invisible, which is the mechanism described in how grounding stops AI hallucination in CRE.
Page-level citation also changes the economics of review. Reviewing 620 uncited fields means reading 620 documents. Reviewing 620 cited fields means opening the cited page for the subset that is flagged low confidence or that contradicts another source. The first is diligence performed twice. The second is diligence performed once and checked.
What should a firm standardize before buying a diligence tool?
Standardize four things: the file format accepted at upload, the field definitions each document type maps into, the completeness rule that flags a missing cross-reference, and the delivery sequence tied to the contingency calendar. All four are written rules. None requires software to define, and none is supplied by a vendor.
MISMO, a subsidiary of the Mortgage Bankers Association, has spent years publishing commercial data standards for exactly this reason: standard business names, definitions, and enumerations so that a term means the same thing across counterparties and transactions. That work exists because definitional inconsistency, not analytical difficulty, is what makes commercial loan data expensive to move. Acquisitions diligence has the same disease and less standardization.
A workable sequence:
Order | Action | Why it comes first |
|---|---|---|
1 | Publish a document request list with required format and deadline per item | Sets the intake contract before the room opens |
2 | Define the field standard per document type | Determines what "extracted" means and what gets compared |
3 | Require page-level citation on every extracted field | Makes review cheap and errors visible |
4 | Measure intake pass rates per deal | Turns document quality into a number the team can improve |
5 | Evaluate models against your own file, not a demo file | Tests the tool on the distribution you receive |
Step 5 is where most firms start. Running it first tests a model against a file that has already been cleaned, which measures the wrong thing. Accuracy claims made on a clean file tell you nothing about performance on yours, which is the argument for measuring extraction accuracy with precision and recall rather than a single number.
Frequently Asked Questions
Is AI due diligence in real estate worth adopting if the data room is messy?
Yes, but adopt it in the order that respects the constraint. Fix format and completeness rules first, then extraction, then analysis. A tool deployed against an unstandardized room produces partial coverage that reads like full coverage, which is worse than no tool.
What is a page-level audit trail and why does it matter for diligence?
It is a citation attached to every extracted value, naming the document and page it came from. It matters because it turns verification from a re-read into a spot check, and because a value with no traceable source cannot be defended to a lender or a committee.
Who owns document intake quality on a deal?
The buy side, in practice, because the buy side bears the cost of a room that cannot be read. Sellers deliver what they have. The request list, the format requirement, and the delivery sequence are the buyer's instruments, and issuing them at signing rather than at week three is the highest-return move in the process.
Conclusion
The question worth asking before any diligence tool purchase is not how accurate the model is. It is what share of your last data room was machine-readable, consistently formatted, complete, and on time. Most firms cannot answer, because nobody has measured it, and the unmeasured number is the one setting the ceiling on everything built above it. A document standard and a page-level audit trail cost a policy memo and a week of enforcement. They also determine whether a diligence model reads half your file or all of it, which is a larger performance difference than any model choice available.