Insights

AI strategy in CRE

·

7 min read

·

AI Underwriting in CRE Replaces Reading, Not Judgment

The industry keeps arguing about whether AI can underwrite a deal. That is the wrong question. Underwriting is two jobs sharing one name: reading, which is the extraction of facts out of documents, and judgment, which is the formation of a view about what those facts are worth. AI underwriting in commercial real estate is competent at the first and has no standing in the second. The distinction is not philosophical. It is a design decision with a cost attached either way, and most firms are paying one of the two prices without having chosen which.

Key Takeaways

  • Underwriting is reading plus judgment. Only the reading half is automatable today, and it consumes the larger share of analyst hours.

  • Extraction is verifiable against a source page. Judgment is not verifiable until the hold period ends, which is why it cannot be graded on the timescale a model needs to learn.

  • Firms that automate the judgment half get confident nonsense delivered faster than a human can check it.

  • Firms that refuse to automate the reading half stay capacity-bound and screen only the deals that arrive on a Monday.

  • The Appraisal Standards Board adopted Advisory Opinion 41 on April 23, 2026, holding that technology can assist analysis but does not transfer responsibility for a credible result away from the professional.

What does AI underwriting in commercial real estate replace?

AI underwriting in commercial real estate replaces document reading: pulling rent rolls, operating statements, leases, and offering memoranda into structured fields an analyst can work with. It does not replace the decision about what those fields mean. Extraction produces inputs. Underwriting is the argument built on top of them.

The confusion comes from the word. When a principal says an analyst "underwrote the deal," the sentence covers three days of typing rent roll lines into a template and twenty minutes of deciding whether the rent mark is credible. Those twenty minutes are the job. The three days are freight. Automation addresses the freight, and the freight is most of the calendar.

Which underwriting steps are reading and which are judgment?

The test is simple: if a step's output can be checked against a document page today, it is reading. If it can only be checked by waiting for the market to resolve, it is judgment. Reading steps have ground truth. Judgment steps have opinions until the exit prints. Sort every step in the process by that test.

Step

Reading or judgment

How it gets verified

Pull unit mix, in-place rents, lease dates from the rent roll

Reading

Field matches the source page

Normalize a T-12 into standard expense categories

Reading

Line items reconcile to the statement

Extract term, options, escalations, and recovery structure from a lease

Reading

Clause and page citation

Identify the loan terms in the assumption package

Reading

Loan document text

Decide whether the submarket supports the rent mark

Judgment

Leases signed 24 to 36 months out

Choose the business plan and the capital to fund it

Judgment

Execution over the hold

Set the exit cap rate

Judgment

The exit trade itself

Decide the price you will not exceed

Judgment

Never, in a single deal

Read the right column. Everything in the reading rows resolves the same day the document arrives. Everything in the judgment rows resolves in years, or in the case of price discipline, never for one deal in isolation. A model learns from feedback. There is no feedback loop on the second group short of a full cycle, which is why the second group stays human. The five-decision structure behind underwriting a commercial real estate deal is a judgment sequence with a reading problem in front of it.

What is the cost of staying capacity-bound on the reading half?

The cost is deal flow that is never screened, and it compounds because the unscreened deals are not random. They are the ones that arrived when the team was busy. A firm that can only read what fits in the week is running a screening process governed by inbox timing rather than by its buy box.

Every figure below derives from the stated inputs. A four-person acquisitions team receives 45 offering memoranda a week. A first-pass read that produces enough structure to compare against the buy box takes 2.5 hours per deal: opening the package, pulling the rent roll, normalizing the trailing statement, and setting up a comparison. Each analyst has 30 hours a week available for screening after meetings, site visits, and live deals under contract.

Input

Value

Inbound offering memoranda per week

45

Manual first-pass hours per deal

2.5

Hours required to screen all inbound

112.5

Screening hours available (4 analysts, 30 hours)

120

Hours available once two live deals consume 60 percent of capacity

48

Deals screenable at 2.5 hours each

19

Deals never opened

26

The team has nominal capacity for the full flow and real capacity for 42 percent of it. Every week the firm is under contract, it discards 26 packages without reading them. Over a year with the team busy half the time, that is roughly 676 deals seen and not screened.

Now change one input. Extraction reduces the first-pass read to 20 minutes of review on a structured output rather than 2.5 hours of manual assembly. The same 48 hours screens 144 deals, which exceeds the inbound. The bottleneck moves off reading and lands where it belongs, on the judgment about which 8 deals deserve a model. That argument does not require believing a machine has a market view. It is the same reason structured data, not the summary, is the real product of any extraction system.

What goes wrong when a firm automates the judgment half?

It produces output with the format of analysis and none of the accountability. A model asked to recommend a price will produce one. It will be specific, formatted, and delivered in seconds, and nothing in the artifact distinguishes a number supported by four verified comparable trades from a number generated to fill a field.

The failure is not that the answer is wrong. Human answers are wrong constantly. The failure is that the error arrives without a confidence signal, without a source, and without an author. A junior analyst who marks rent to $9.00 has a reason and can be asked for it. A generated $9.00 has a shape and no reason behind it, and the committee cannot tell the two apart on the page.

This is why the standards bodies drew the line where they did. The Appraisal Standards Board, in Advisory Opinion 41 adopted April 23, 2026, holds that a computer-assisted valuation tool assists the analysis but does not relieve the professional of responsibility for a credible result, and that an automated valuation output is not by itself an appraisal. The same logic governs acquisitions. The tool can hand you the inputs. It cannot hold the opinion, because it cannot be held to it.

The practical control is a confidence threshold on every extracted field, so that low-confidence values route to a person rather than flowing into a model unchallenged. The mechanics of that routing are covered in when a confidence score means an extraction needs a human.

How should a firm draw the line in practice?

Draw it at verifiability. Automate any step whose output can be checked against a page today, and require a citation for each one. Keep any step whose output can only be checked by waiting for the market. Then measure the two halves separately, because they fail in different ways and improve through different mechanisms.

JLL's Global Real Estate Technology Survey, based on responses from more than 1,000 senior decision-makers across 16 markets, found 92 percent of organizations piloting AI in select real estate use cases or planning to start, up from 61 percent in 2024 and under 5 percent in 2023. In the same survey, only 5 percent reported achieving all of their program goals. The adoption number and the outcome number describe the same mistake: pilots aimed at the judgment half, which cannot be graded, instead of the reading half, which can be graded on the day it runs.

Measure the two halves separately. Extraction is scored on field-level accuracy against source pages and on the share of deals clearing review untouched. Judgment is scored on the pass rate and on the gap between underwritten and realized assumptions at exit. A firm reporting one number for both has already collapsed the distinction that makes the system work.

Frequently Asked Questions

Can AI underwrite a commercial real estate deal on its own?

No. It can produce the structured inputs a deal model runs on, with each field traceable to a source page. The view on rent growth, business plan, capital structure, and price is a claim about the future that has to be owned by a person who can be held to it.

What is the difference between AI underwriting and automated valuation?

AI underwriting assembles verified inputs from deal documents. Automated valuation produces an estimate of value from those inputs and comparable data. Standards bodies treat the second as a tool output rather than a professional opinion, and that distinction is worth carrying into acquisitions.

Where does AI underwriting fail most often in practice?

On tables and on amendments. Rent rolls arrive in inconsistent formats and leases arrive with side letters that modify terms elsewhere in the document. Both are reading problems, both are measurable, and both are fixed by better extraction rather than by a better model of the market.

Conclusion

The useful framing is not whether AI can underwrite. It is which half of underwriting a firm is buying help with. The reading half is high volume, verifiable on the day it runs, and the direct cause of the deals a firm never opens. The judgment half is low volume, unverifiable for years, and the entire reason the firm exists. Automate the first and the second gets more of the week. Automate the second and you get a faster path to a confident number nobody can defend.

Related