Insights

AI strategy in CRE

·

8 min read

·

What It Means When Your AI Vendor Trains on Your Deal Flow

AI vendor data rights are treated as a privacy question. They are not. When a vendor trains a shared model on your deal flow, the model does not learn your data. It learns your judgment: which deals you screened, which you passed on, how you adjusted the broker's numbers, and where your buy box ends. That judgment is the only asset a CRE firm has that its competitors cannot buy, and a training clause converts it into a product feature sold to everyone, including the firm across the street. The leak is not a document. It is a pattern.

Key Takeaways

  • Training on customer data is a license grant, not a security event. The clause that permits it usually reads "improve" or "enhance the services," language that predates generative AI.

  • A Stanford Law School CodeX analysis of TermScout contract data found 92% of AI contracts claim data usage rights beyond what service delivery requires, against a 63% market average across software agreements.

  • The valuable thing in a CRE firm's deal flow is not the rent roll. It is the decision attached to it. A shared model trained on decisions learns a firm's underwriting posture and reproduces it for other customers.

  • Deidentification does not solve the problem. Strip the address from a passed deal and the pattern that made you pass survives intact.

What does it mean for an AI vendor to train on your data?

It means the vendor uses your inputs, outputs, or both to adjust the weights of a model that serves customers other than you. Training is distinct from processing. A pre-trained model reads your document and returns a result without changing. A model trained on your document changes, permanently, and carries what it learned into every later inference for every later user.

Venable's June 2026 note on contracting for AI model training states the mechanism plainly: training data is "typically 'memorized' by the model, which can then reproduce similar or identical content when prompted." The firm also notes that it is entirely possible to provide AI services without training on customer data. The vendor chooses. The customer's leverage is at signature.

Vendor posture

What happens to your deal flow

Exposure

No training, no retention beyond the session

Processed and discarded

Confidentiality only

Training on a private instance for your firm

Your model learns from your data, nobody else's

Retention and access controls

Training on a shared or foundational model

Your data shapes a model sold to every customer

Your judgment becomes a shared feature

The third row is the one most agreements permit by default.

Why is training on deal flow different from training on documents?

Because a document contains facts, and deal flow contains decisions. A vendor that trains on a thousand offering memoranda learns what an OM looks like. A vendor that trains on a thousand OMs plus the firm's screening outcome on each one learns what the firm buys, what it rejects, and the gap between the broker's pro forma and the number the firm underwrote to.

Consider the fields a screening workflow generates around one deal: asking cap rate, adjusted cap rate, reason code for a pass, the rent growth assumption the analyst overrode, whether the deal advanced. None is confidential in the NDA sense. Together they are the firm's underwriting model expressed as data.

A shared model trained on that data does not need to reproduce any single deal to cause harm. It needs only to have absorbed the posture. When a competitor asks that model to screen a deal in the same asset class, the answer is informed by your history of passes and pursuits. The competitor never sees your data. They receive your judgment as a default recommendation.

The Stanford CodeX analysis of TermScout data found many AI contracts allow vendors to use customer data for competitive intelligence. In CRE, competitive intelligence and deal flow are the same thing.

Where in the contract does the training right hide?

In the clause that lets the vendor use customer data to improve, develop, or enhance the services. That language was standard in software agreements long before generative AI, negotiated as permission to fix bugs and study usage. Read against a model that learns from inputs, the same words authorize training, and most agreements signed in the last three years make no distinction.

Clause language

What it was written for

What it now permits

"to provide and improve the Services"

Bug fixes, usage analytics

Training a shared model on inputs

"aggregated or deidentified data"

Benchmarking reports

Training on decision patterns with names removed

"derivative works" or "insights"

Product analytics

Vendor ownership of patterns learned from your deals

"feedback"

Customer suggestions

Your corrections to model output, the best training data there is

The last row matters most. Every analyst correction to an extracted field or a screening recommendation is the most valuable training signal a vendor can collect, and feedback clauses routinely assign all rights in it to the vendor. A firm that prohibited training on documents and left the feedback clause standard protected the raw material and gave away the refined product.

Juanita DeLoach of Barnes and Thornburg, cited by PYMNTS in June 2026, recommends that agreements define what counts as training data, require notification before model changes, and specify who owns the patterns a model derives from customer data. Those are the three things that matter most for a CRE firm: inputs, model drift, and derived judgment.

Does deidentification protect a CRE firm?

Not for the exposure that matters. Deidentification removes the identifiers: the property address, the seller, the broker, the tenants. What it leaves behind is the structure of the decision, and the structure is the asset. Venable's note says this outright: training data can be "deidentified" and still contain confidential or valuable information.

Work through a concrete case. A firm screens 400 deals in a year and passes on 350. The vendor deidentifies each record before training: address, parties, and exact dates stripped. What survives per record is asset type, submarket tier, size band, asking cap rate, the firm's adjusted cap rate, the outcome, and the reason code.

Across 400 records, that deidentified set encodes the firm's buy box with precision no memo ever achieved: the cap rate spread at which the firm walks, the submarkets it will not enter, the size floor, the rent growth assumption it refuses to accept. None of it names a property. All of it is the firm's strategy, and a model trained on it will hand that strategy to whoever asks next. The firm did not lose 400 documents. It lost what it spent ten years learning from them.

This is the same category error as treating OM confidentiality as a security question rather than a permitted-recipient question. The security review asks whether the data is protected. The right question is whether the vendor is permitted to learn from it.

What should a CRE firm require in the vendor agreement?

A default of no training on customer inputs, outputs, or corrections for any model serving other customers, stated affirmatively rather than left to the absence of a permission. The firm can then decide, deliberately, whether to grant a narrower right in exchange for something it values, rather than discovering it granted a broad one in exchange for nothing.

Term

Minimum position

Why

Training on inputs, outputs, corrections

Prohibited for shared models unless expressly granted

Closes the "improve the services" gap; corrections are the highest-value signal

Deidentified or aggregated use

Defined by method, with decision fields excluded

Names alone are not the asset

Derived insights and patterns

Owned by the customer or prohibited

Covers the judgment layer, not just the records

Model change notice

Advance notice before the production model changes

The model evaluated at procurement may not be the one running in six months

Retention

Deletion on a schedule, verifiable

Retention without training is still a disclosure

Audit right

Ability to confirm the above

An unverifiable promise is a hope

The audit line ties to the broader problem that most vendor accuracy claims cannot be tested: a training prohibition the firm cannot verify is in the same category.

The private instance deserves a direct decision. Venable notes that many enterprise providers offer one, often at a premium, in which the customer's data trains only the customer's own copy of the model. For a firm whose edge is a differentiated buy box, that premium buys improvement without leakage. Not asking is the only wrong answer.

Ownership of derived patterns defaults to the vendor, because most agreements are silent on it and silence favors whoever holds the model. A chain of custody for extracted data tracks which document produced which field. It cannot track which decisions taught a model what, because weights carry no provenance. Once a pattern is in the weights, there is no record to delete. The only control that works is applied before training happens.

Firms that treat this as a contracting problem compound an advantage: their corrections improve a model only they use. Firms that leave the clause standard subsidize a shared model with a decade of judgment, then compete against customers who bought access to it for a monthly fee.

Frequently Asked Questions

Does "we do not train on your data" in a vendor's marketing mean the contract prohibits it?

No. Marketing statements are not contract terms. Figma was sued in November 2025 over allegations that it opted customers into training without informing them, which the company denied. The only statement that binds is the one in the signed agreement.

Can a vendor train on deidentified data without exposing the firm?

Only if the deidentification removes the decision fields, not just the identifiers. The screening outcome, the adjusted assumptions, and the reason codes are the valuable part, and they survive ordinary deidentification untouched.

What is the difference between a private model and a foundational model in a vendor agreement?

A private model trains only on the customer's data and serves only that customer. A foundational or shared model trains across customers and can reproduce patterns from one customer's data in outputs to another. Vendors often offer both, with the private option at a premium.

Conclusion

The instinct is to file AI vendor data rights under security and let the questionnaire handle it. The questionnaire asks whether data is protected, not whether the vendor is permitted to learn from it, and for a CRE firm the learning is the loss. The documents are replaceable. The decisions attached to them are the firm.

Read the "improve the services" clause as a training clause, because it is one. Treat corrections as data. Define deidentification by what it removes, and make sure it removes the judgment. Put ownership of derived patterns in writing. A firm that does this keeps its buy box. A firm that does not has already licensed it, on the vendor's paper, for free.