The gap between an AI pilot and production in CRE is treated as a scaling problem. It is not. A pilot tests whether a model can read a lease or screen a deal. Production tests whether the firm can depend on it: who owns the output, who catches the errors, and which system the result lands in. Most CRE pilots are designed to answer the first question and never ask the other three. When the handoff comes, the pilot has proven the wrong thing, and it dies not because the model failed but because nothing around the model was ever built.
Key Takeaways
JLL's 2025 Global Real Estate Technology Survey found 88% of investors, owners and landlords have started piloting AI, while only 5% of corporate real estate teams report achieving all their program goals.
A pilot proves the model can do the task. Production requires proof that the firm can absorb the task: an owner, an exception path, and a system of record.
Most pilots contain an invisible corrector: the champion who quietly fixes errors. When the champion leaves the loop, the accuracy the pilot reported leaves with them.
A pilot that runs on a curated sample measures the easy documents. Production runs on the full distribution, including the amendments and scanned side letters the sample left out.
The fix is to design the pilot as a small production run, not a large demo.
Why do so many AI pilots in CRE never reach production?
Because the pilot is scoped to answer a question production does not ask. A pilot asks whether the model produces a good output on selected documents. Production asks whether the firm can rely on that output every day without the pilot team present. Those are different tests, and passing the first says little about the second.
The industry data shows the pattern at scale. JLL's 2025 Global Real Estate Technology Survey, covering more than 1,500 senior CRE decision-makers across 16 markets, found investors pursuing an average of five use cases at once. JLL's own conclusion: "most initiatives remain experimental with limited scaling." The survey locates the problem in organizational readiness: data quality, infrastructure, and the change-management processes needed to put AI into core workflows.
Outside real estate the pattern is the same. Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs, and unclear business value. None of those four reasons is about whether the model works in a demo.
What does a pilot hide that production exposes?
Three things: a curated sample, an invisible corrector, and a missing system of record. Each makes the pilot look better than the workflow it is meant to predict. Each is invisible during the pilot because the pilot team is standing in for the infrastructure production will need, and the handoff is the moment that stand-in disappears.
What the pilot hides | How it looks in the pilot | What happens in production |
|---|---|---|
Curated sample | 40 clean leases the champion picked | The full portfolio arrives: amendments, scans, side letters, estoppels |
Invisible corrector | Champion fixes errors before anyone sees them | Nobody owns review; errors reach the model and the rent roll |
Missing system of record | Output lives in a spreadsheet or a demo screen | Someone rekeys results into the lease admin or underwriting system, or no one does |
The curated sample is the most common and the least intentional. A champion choosing documents for a pilot picks ones that are available, legible, and representative of the use case as they imagine it. That is a biased sample by construction. We covered the mechanics in why a CRE extraction model that wins the demo fails in production: the tail of the document distribution is where accuracy collapses, and pilots rarely sample the tail.
The invisible corrector is a person doing good work. The champion catches the wrong renewal date, fixes the misread CAM cap, and moves on. The pilot report shows high accuracy. What it measured was the model plus one motivated expert.
The missing system of record is where most pilots die quietly. If extracted fields never flow into the lease administration platform or the underwriting model, someone has to carry them across by hand. That rekeying cost was never in the business case, and within a quarter the team routes around the tool.
How big is the gap between pilot accuracy and production workload?
Larger than the accuracy figure suggests, because accuracy is reported per field and cost is incurred per review. A pilot that reports 94% field accuracy sounds finished. In production, the firm does not know which 6% is wrong, so it must either review everything or accept errors it cannot locate. The worked example below uses stated inputs, not survey figures.
Worked example: a lease abstraction pilot moves to production.
Pilot inputs: 40 leases chosen by the champion, 60 fields abstracted per lease, 94% field accuracy.
Fields: 40 x 60 = 2,400, reviewed at 20 seconds each: 800 minutes
Errors: 6% of 2,400 = 144, corrected at 2 minutes each: 288 minutes
Total: about 18 hours, absorbed by the champion without anyone logging it
Production inputs: 1,500 documents a year across the portfolio, including amendments and scanned originals. Assume field accuracy drops to 88% on the full distribution, a representative assumption for messier documents, not a measured figure.
Fields per year: 1,500 x 60 = 90,000
Expected errors: 12% of 90,000 = 10,800
Review at 20 seconds per field: 500 hours
Correction at 2 minutes per error: 360 hours
Total: about 860 hours a year
The pilot consumed 18 hours that nobody counted. Production requires roughly 860 hours, close to half of one full-time analyst, that somebody must own, schedule, and pay for. The model got six points worse. The workload grew almost fifty times, because the pilot measured the model and production bills the process.
Who should own an AI workflow when it moves to production?
The team whose numbers the output feeds, not the team that ran the pilot. If extracted lease data drives the rent roll, asset management owns it. If screening output drives the pipeline, acquisitions owns it. Innovation or IT can run a pilot. They cannot own a production workflow whose errors land in someone else's model.
The same JLL survey found a considerable portion of corporate real estate teams implementing AI by C-suite mandate rather than by choice. A mandated pilot has a sponsor but no owner. The sponsor wants a result; the owner has to live with the error rate. When those are different people, the handoff has no one to hand off to.
Role | Pilot | Production |
|---|---|---|
Sponsor | Executive who wants AI adoption | Same, but no longer sufficient |
Operator | Champion who ran the test | Line team that uses the output daily |
Reviewer | Champion, informally | Named person with a review quota and an exception queue |
Owner of errors | Nobody, because errors were fixed silently | The team whose numbers the output feeds |
A useful test before any pilot starts: name the person who will be accountable when a wrong field reaches an investment committee memo. If no one can be named, the pilot is a demo.
How do you design a pilot that survives the handoff?
Run it as a small production deployment instead of a large demonstration. Use a random sample of real documents, route output into the actual system of record, assign the future owner as the reviewer, and log every correction. The pilot then measures what production will cost, not what the model can do on a good day.
Four design changes do most of the work:
Sample at random. Draw documents from the full population, including amendments and poor scans. The accuracy figure becomes a forecast instead of a best case.
Log every correction. Each fix the reviewer makes is data: which field, which document type, how long it took. That log is the basis for the review budget and for detecting model drift in production later.
Land output where it will live. If the target is the lease administration system, write to it during the pilot. Integration discovered at handoff is integration that never gets funded.
Put the future owner in the loop. The line team reviews from day one. Their experience of the tool is the experience production will have.
The line worth keeping: a pilot should be judged by the cost of the process it predicts, not by the accuracy of the model it demonstrates.
Firms that design pilots this way compound an advantage. Each pilot either produces a workflow the team already runs or a clear, cheap reason to stop. Firms that run demos accumulate a portfolio of impressive results with no owner, which is the "limited scaling" JLL describes.
Frequently Asked Questions
What is the most common reason an AI pilot fails to reach production in CRE?
The pilot never tested the workflow around the model. Ownership, error review, and integration into the system of record were handled informally by the pilot team, so production inherits a cost and an accountability gap that the business case never included.
How many documents should a CRE AI pilot use?
Enough to include the document types production will see, drawn at random rather than selected. Representative coverage of amendments, scans, and unusual lease structures matters more than volume, because the difficult tail is where production cost concentrates.
Should IT or the business team own an AI deployment?
The business team whose numbers depend on the output should own it. IT can host and secure the system, but the owner must be the team accountable when an extracted field or screening result turns out to be wrong.
Conclusion
AI pilots in CRE do not die at the handoff because the model underperforms. They die because the pilot measured the model and production measures the firm. A curated sample flattered the accuracy, an invisible corrector absorbed the errors, and the output never landed in the system that uses it.
The remedy sits at the start, not the end. Design the pilot as the smallest honest version of production: random documents, a named owner, logged corrections, real integration. The result will look less impressive in the steering meeting. It will also be the only kind of pilot that has something to hand off.