Insights

AI strategy in CRE

·

7 min read

·

The AI Acceptance Test: Proving a Deployment Works Before You Depend On It

AI acceptance testing is the step most CRE firms skip, and it is the only step that converts a purchase into a dependency they can defend. A demo shows what a system can do. A pilot shows what it did for a motivated team. An acceptance test shows whether it meets a standard the buyer wrote in advance, on the buyer's documents, with a pass or fail answer. A firm that depends on an AI system it never formally accepted has not adopted a tool. It has adopted an assumption.

Key Takeaways

  • An acceptance test is a pass or fail check, written by the buyer before any results are seen, on a fixed set of the buyer's own documents.

  • Criteria written after the results arrive are not criteria. They are a description of whatever happened.

  • Sample size decides what a clean result proves. Zero errors across 40 instances of a field only bounds the true error rate below about 7% at 95% confidence. Proving below 1% takes about 300 clean instances.

  • Thresholds belong to fields, not to the system. A misread suite number and a misread expiration date should not share a pass mark.

  • Acceptance is not a one-time event. Every model update, prompt change, or new document type reopens the test.

What is an AI acceptance test, and how is it different from a pilot?

An AI acceptance test is a formal, pass or fail evaluation of a deployed system against criteria the buyer fixed before testing began. It differs from a vendor benchmark, which the seller designs, and from a pilot, which measures experience. The acceptance test answers one narrow question: does this system meet our written standard on our documents?


Vendor benchmark

Pilot

Acceptance test

Who designs it

Seller

Champion or innovation team

Buyer, with the future owner

Test set

Seller's documents

Often hand-picked

Fixed, held-out sample of buyer's documents

Pass criteria

Headline accuracy

Informal satisfaction

Written per-field thresholds

When criteria are set

Before you arrive

Rarely at all

Before results are seen

Output

A marketing figure

An impression

Pass, fail, or conditional pass

We covered how to interrogate the first column in how to test a CRE AI vendor accuracy claim, and why the second column misleads in why AI pilots in CRE die at the handoff to production. The acceptance test is the third column, and it is the one that carries accountability.

Why must acceptance criteria be written before the results are seen?

Because criteria written after the results are shaped by the results. Once a team has seen 91% field accuracy, the pass mark quietly becomes 90%. Fixing thresholds, the test set, and the scoring rules in advance is what makes the outcome evidence rather than narrative. Banking regulators built model validation on the same principle.

The Federal Reserve's SR 11-7 Supervisory Guidance on Model Risk Management describes validation as "effective challenge": critical analysis of a model by informed parties with the independence and authority to find its limits. It also calls for outcomes analysis, including backtesting against data not used in development. CRE firms are not bound by SR 11-7, but their lenders are, and the logic transfers intact. A test set the system has seen proves nothing. A threshold chosen after the score proves less.

In practice, pre-registration is a one-page document signed before the test runs. It names the document sample and how it was drawn, the fields in scope, the scoring rule for each field (exact match for dates, tolerance for dollar amounts, clause presence for provisions), the threshold for each field, and what happens on a fail. If the page does not exist before the test, the test did not happen.

How many documents does an AI acceptance test need?

Enough that a clean result bounds the error rate below your threshold. A statistical rule sets the floor: if a field shows zero errors across n instances, the true error rate is below roughly 3 divided by n at 95% confidence. Hanley and Lippman-Hand described this "rule of three" in JAMA in 1983.

The rule changes how a firm reads a clean pilot. Forty leases with zero errors on rent commencement date sounds decisive. It is not.

Worked example: sizing the test for one critical field.

Requirement: the firm wants 95% confidence that the rent commencement date error rate is below 1%. Figures below are exact binomial bounds computed from the stated inputs.

Instances tested

Errors observed

Upper bound on true error rate (95%)

Meets 1% bar?

40

0

about 7.2%

No

100

0

about 3.0%

No

300

0

about 1.0%

Borderline pass

300

1

about 1.6%

No

300

2

about 2.1%

No

A 40-lease pilot cannot prove a 1% standard even if it is perfect; it only shows the system is not worse than about 7%. At 300 instances, a single error fails the test. The firm either accepts a looser threshold for that field, expands the sample, or keeps human review on that field permanently. All three are legitimate. Pretending the 40-lease result met a 1% bar is not.

The sample must also be drawn at random from the population the system will see in production, including amendments, scanned originals, and older lease forms. A clean 300 drawn from recent, typed leases bounds the error rate on recent, typed leases and nothing else.

Which fields deserve the strictest acceptance thresholds?

The fields whose errors move money or trigger deadlines. Base rent, rent steps, commencement and expiration dates, and renewal notice windows feed the rent roll and the valuation directly. Descriptive fields such as suite numbers or notice addresses matter less. Acceptance thresholds should be tiered by consequence, not averaged into one system-wide score.

A single accuracy figure lets a model excel on 50 easy fields while failing on the five that carry the risk. Tiering forces attention onto those five.

Tier

Example fields

Consequence of an error

Illustrative acceptance rule

1: Money and dates

Base rent, rent steps, commencement, expiration, renewal notice deadline

Wrong cash flow, missed option, mispriced asset

Error rate bounded below a strict threshold, or permanent human review

2: Economic provisions

CAM caps, co-tenancy, exclusives, termination rights

Recovery leakage, tenant rights missed in underwriting

Clause presence recall tested separately from value accuracy

3: Descriptive

Suite, notice address, tenant contact

Administrative friction

Looser threshold, spot-check review

The thresholds in the last column are a firm's choice, not an industry standard, and they should be written by the team that owns the output. The asset manager who reconciles the rent roll knows which misreads cost money. A procurement team does not. The same tiering logic underpins a lease abstraction QA process that catches the errors that cost money; acceptance testing applies it once, formally, before the system is trusted.

The line worth keeping: a system is accepted field by field, not as a whole, and the fields that move money should be the hardest to pass.

What happens after an AI deployment passes acceptance?

The acceptance test becomes a regression test. Every model update, prompt change, extraction rule change, or new document type reruns the same fixed sample against the same thresholds. A pass certifies one version of the system on one distribution of documents. When either changes, the certificate lapses until the test is run again.

SR 11-7 frames this as ongoing monitoring: a systematic process to evaluate performance over time, with changes in performance triggering renewed validation. For CRE extraction the triggers are concrete: a new vendor model, an acquired portfolio with a different lease form, a new document type in scope. Each rerun costs little because the sample and scoring rules already exist. Drift that no one tests for shows up later as a reconciliation break, as described in how CRE extraction models degrade in production.

Professional standards are moving the same way. The RICS professional standard on responsible use of AI in surveying practice, in effect since 9 March 2026, applies to AI with a material impact on surveying services and places the surveyor's judgement, not the system's output, at the center of practice. A documented acceptance test is the most direct evidence that judgement was exercised before the output was relied on.

Firms that accept systems this way compound an advantage: each new tool arrives with a known error profile, a named owner, and a test that can be rerun in an afternoon. Firms that skip it accumulate systems whose reliability is a matter of recollection.

Frequently Asked Questions

Who should write the AI acceptance test in a CRE firm?

The team that will own the output, with support from whoever manages the vendor relationship. The owner knows which fields carry financial consequence and will live with the error rate, so the owner sets the thresholds.

Can a vendor's own benchmark replace an acceptance test?

No. A vendor benchmark is run on the vendor's documents against the vendor's definition of correct. An acceptance test is run on the buyer's documents against the buyer's written standard, which is the only result the buyer can defend internally.

What if the system fails the acceptance test on some fields?

Accept it conditionally. Deploy it on the fields that passed, keep human review on the fields that failed, and rerun the test on those fields after the next model update. A partial pass is useful information, not a failed purchase.

Conclusion

An AI deployment is not proven by a demo, a pilot, or a vendor's number. It is proven by a test the buyer wrote before seeing the results, sized to bound the error rate on the fields that move money, and rerun whenever the system or its documents change. That test is the firm's answer when a lender or investment committee asks how it knows the numbers are right. Without it, the firm has a vendor's word and a pilot that went well.