AI comparable selection is treated as a judgment problem: which sales deserve a place in the set. It is not, or not first. Before anyone judges a comp, a system has to find it, and finding is retrieval. Retrieval has one failure mode that matters more than the rest: it fails without telling you. A comp set that is missing its most relevant sale looks exactly like a comp set that is complete. The analyst reviews what came back, approves it, and never sees what did not.
Key Takeaways
Comp selection is two steps: retrieval (what the search returns) and judgment (what the analyst keeps). Most review effort goes to the second step, while most unseen error enters at the first.
Precision, the share of returned comps that are relevant, can be checked by reading the list. Recall, the share of relevant comps that were returned, cannot be checked by reading the list.
A comp set missing a sale looks exactly like a comp set that is complete. That is why retrieval failures survive review.
On the BEIR benchmark of 18 retrieval datasets, plain keyword ranking (BM25) proved a stronger zero-shot baseline than many dense semantic models.
Recall is audited by testing, not by reading: seed known comps, loosen filters, run lexical and semantic search side by side, and log the query.
Why is comp selection a retrieval problem?
Comp selection is a retrieval problem because every comp set starts as a query against a database: a property type, a submarket, a date window, a size band. The analyst only judges what that query returns. If the query misses a relevant sale, no amount of judgment downstream recovers it, because the sale never reaches the page.
The appraisal standard already frames it this way. USPAP Standards Rule 1-4(a) requires an appraiser to analyze "such comparable sales data as are available." The phrase does the work: the obligation covers the sales that are available, which means the sales that can be found. The sales comparison approach is only as good as the search that feeds it.
Manual comp work hid this. An analyst who has covered a submarket for ten years notices when a known trade is absent. Software replaces that memory with a query, and the query notices nothing. It returns its results with the same confidence whether it found every relevant sale or half of them.
Why does retrieval fail quietly?
Retrieval fails quietly because its two error types are not equally visible. A wrong comp in the set is a precision error, and the analyst can see it and strike it. A right comp absent from the set is a recall error, and nothing on the page shows it. Review catches the first kind and cannot catch the second.
The definitions come from information retrieval, formalized in Manning, Raghavan and Schütze's Introduction to Information Retrieval (Cambridge University Press). Precision is the fraction of retrieved items that are relevant. Recall is the fraction of relevant items that are retrieved. See the precision and recall entry for the full treatment.
Error | What it looks like | Visible on review? | Who catches it |
|---|---|---|---|
Irrelevant comp returned (precision) | A retail strip in an industrial set | Yes | The analyst, by reading |
Relevant comp not returned (recall) | Nothing | No | Only a test built to find it |
Relevant comp ranked below the cutoff | Nothing, if the list is truncated | No | Only a test built to find it |
Relevant comp returned but ignored | Sits deep in a long list | Partly | Rarely, under time pressure |
The asymmetry compounds with the number of comps. As an illustration, assume each truly relevant sale has an independent 20% chance of being missed by the search. With seven relevant sales in the market, the chance that all seven come back is 0.8 to the seventh power, about 21%. Nearly four searches in five miss at least one. The 20% is assumed, but the pattern holds at any rate: small per-sale miss rates become likely set-level misses.
How do keyword filters and semantic search each miss comps?
Keyword filters miss comps that are coded differently from the query, and semantic search misses comps whose descriptions sit far from the query's meaning in the model's vector space. The two fail on different sales. That is the argument for running both, and against trusting either one as the whole retrieval layer.
Failure mode | Keyword or structured filter | Semantic search |
|---|---|---|
Inconsistent property type coding | Misses "flex" when searching "industrial" | Often finds it |
Submarket named differently across sources | Misses it | Sometimes finds it |
Missing field (no size, no date) | Drops the record at the filter | May still return it |
Exact figures and identifiers | Strong | Weak; numbers blur in embeddings |
Domain language the model never trained on | Unaffected | Can misrank badly |
Top-k cutoff | Only if sorted and truncated | Always truncates by rank |
The research record supports running both. The BEIR benchmark (Thakur et al., NeurIPS 2021) tested lexical, sparse, dense, late-interaction and re-ranking models across 18 datasets without domain training. BM25, the classic keyword ranking function, held up as a strong baseline, and several dense semantic models underperformed it once moved outside the data they were trained on. CRE transaction records are exactly that kind of out-of-domain text.
A second failure sits after retrieval. When comps are passed to a language model to summarize, position matters. Lost in the Middle (Liu et al., Stanford, 2023) found model performance is highest when relevant information sits at the start or end of the input and degrades significantly when it sits in the middle. A comp retrieved in position eight of fifteen can be retrieved and still effectively lost. This is the same problem retrieval-augmented generation has to solve for every grounded answer.
What does a missed comp cost?
A missed comp costs the difference between the value indicated by the partial set and the value indicated by the full set, and that difference lands on the buyer's equity. Because comp sets are small, one absent sale moves the average more than intuition suggests, especially when the absent sale traded at a different price.
Worked example, with all inputs stated:
Input | Value |
|---|---|
Subject | 120,000 SF industrial building |
Stabilized NOI | $1,800,000 |
Comps returned by search | 5, average cap rate 6.00% |
Comps missed (coded as "flex" in the source) | 2, at 6.75% and 6.50% |
Full-set average cap rate | (5 × 6.00 + 6.75 + 6.50) ÷ 7 = 6.18% |
Comp set | Cap rate | Indicated value |
|---|---|---|
Partial (5 comps) | 6.00% | $30,000,000 |
Full (7 comps) | 6.18% | about $29,133,000 |
Difference | 18 bps | about $867,000 |
The value gap is about 2.9%. Against a 35% equity check on $30 million, $10.5 million, the gap is about 8.3% of the equity. Nobody made a judgment error. The analyst evaluated every comp correctly. The loss entered at retrieval, and the review process had no way to see it.
Straight averaging is a simplification, and real work adjusts each comp. The point survives adjustment. An adjustment grid can correct a comp it has. It cannot correct for a comp it never received.
How do you audit recall in a comp set?
Recall is audited by testing the search, not by reading its output. The method is to plant sales you know should appear, change the query and watch what moves, and compare two retrieval methods that fail differently. What the tests surface is the set of relevant sales the original query silently dropped.
Test | How it works | What it catches |
|---|---|---|
Known-item seeding | Keep a list of trades the team knows are relevant; confirm each is returned | Coding and field gaps |
Filter loosening | Widen size band, date window and type by one step; review what enters | Hard-filter exclusions |
Hybrid retrieval | Run keyword and semantic search separately; diff the results | Method-specific blind spots |
Rank-depth check | Read past the default cutoff to position 20 or 30 | Top-k truncation |
Query logging | Save the exact query with the comp set | Makes the miss reproducible later |
The last test matters most for accountability. A comp set without its query is a conclusion without its method. When a valuation is challenged a year later, the question is which search produced the comps, and whether the missing sale was reachable by it. Submarket definitions are a frequent culprit, which is why the submarket, not the metro should be the unit the query is built around.
Frequently Asked Questions
Is AI comparable selection more accurate than manual comp pulls?
It can be faster and broader, but accuracy depends on recall, which neither method reports on its own. An AI search that is never tested for recall is not more accurate than a manual pull. It is less visible.
Should comp search use keywords or semantic search?
Use both. Research on the BEIR benchmark found keyword ranking a strong baseline across domains, while semantic models catch differently worded records; the two miss different sales.
How many comps are enough?
There is no fixed number. The question that matters is whether the set contains every relevant sale that was available, which only a recall test can answer.
Can a reviewer catch a missing comp?
Only if the reviewer already knows the sale exists. Otherwise nothing on the page signals its absence, which is why recall must be tested rather than reviewed.
Conclusion
Comps are a retrieval problem first and a judgment problem second. Judgment gets the attention because it is visible. Retrieval gets none, because its failures leave no trace. The operator's job is to make the invisible error testable. Seed known sales, run two methods that fail differently, read past the cutoff, and keep the query with the conclusion. A comp set is only as complete as the search that built it, and the search will never say what it missed.