Insights

Market analysis

·

7 min read

·

Comps Are a Retrieval Problem, and Retrieval Fails Quietly

AI comparable selection is treated as a judgment problem: which sales deserve a place in the set. It is not, or not first. Before anyone judges a comp, a system has to find it, and finding is retrieval. Retrieval has one failure mode that matters more than the rest: it fails without telling you. A comp set that is missing its most relevant sale looks exactly like a comp set that is complete. The analyst reviews what came back, approves it, and never sees what did not.

Key Takeaways

  • Comp selection is two steps: retrieval (what the search returns) and judgment (what the analyst keeps). Most review effort goes to the second step, while most unseen error enters at the first.

  • Precision, the share of returned comps that are relevant, can be checked by reading the list. Recall, the share of relevant comps that were returned, cannot be checked by reading the list.

  • A comp set missing a sale looks exactly like a comp set that is complete. That is why retrieval failures survive review.

  • On the BEIR benchmark of 18 retrieval datasets, plain keyword ranking (BM25) proved a stronger zero-shot baseline than many dense semantic models.

  • Recall is audited by testing, not by reading: seed known comps, loosen filters, run lexical and semantic search side by side, and log the query.

Why is comp selection a retrieval problem?

Comp selection is a retrieval problem because every comp set starts as a query against a database: a property type, a submarket, a date window, a size band. The analyst only judges what that query returns. If the query misses a relevant sale, no amount of judgment downstream recovers it, because the sale never reaches the page.

The appraisal standard already frames it this way. USPAP Standards Rule 1-4(a) requires an appraiser to analyze "such comparable sales data as are available." The phrase does the work: the obligation covers the sales that are available, which means the sales that can be found. The sales comparison approach is only as good as the search that feeds it.

Manual comp work hid this. An analyst who has covered a submarket for ten years notices when a known trade is absent. Software replaces that memory with a query, and the query notices nothing. It returns its results with the same confidence whether it found every relevant sale or half of them.

Why does retrieval fail quietly?

Retrieval fails quietly because its two error types are not equally visible. A wrong comp in the set is a precision error, and the analyst can see it and strike it. A right comp absent from the set is a recall error, and nothing on the page shows it. Review catches the first kind and cannot catch the second.

The definitions come from information retrieval, formalized in Manning, Raghavan and Schütze's Introduction to Information Retrieval (Cambridge University Press). Precision is the fraction of retrieved items that are relevant. Recall is the fraction of relevant items that are retrieved. See the precision and recall entry for the full treatment.

Error

What it looks like

Visible on review?

Who catches it

Irrelevant comp returned (precision)

A retail strip in an industrial set

Yes

The analyst, by reading

Relevant comp not returned (recall)

Nothing

No

Only a test built to find it

Relevant comp ranked below the cutoff

Nothing, if the list is truncated

No

Only a test built to find it

Relevant comp returned but ignored

Sits deep in a long list

Partly

Rarely, under time pressure

The asymmetry compounds with the number of comps. As an illustration, assume each truly relevant sale has an independent 20% chance of being missed by the search. With seven relevant sales in the market, the chance that all seven come back is 0.8 to the seventh power, about 21%. Nearly four searches in five miss at least one. The 20% is assumed, but the pattern holds at any rate: small per-sale miss rates become likely set-level misses.

How do keyword filters and semantic search each miss comps?

Keyword filters miss comps that are coded differently from the query, and semantic search misses comps whose descriptions sit far from the query's meaning in the model's vector space. The two fail on different sales. That is the argument for running both, and against trusting either one as the whole retrieval layer.

Failure mode

Keyword or structured filter

Semantic search

Inconsistent property type coding

Misses "flex" when searching "industrial"

Often finds it

Submarket named differently across sources

Misses it

Sometimes finds it

Missing field (no size, no date)

Drops the record at the filter

May still return it

Exact figures and identifiers

Strong

Weak; numbers blur in embeddings

Domain language the model never trained on

Unaffected

Can misrank badly

Top-k cutoff

Only if sorted and truncated

Always truncates by rank

The research record supports running both. The BEIR benchmark (Thakur et al., NeurIPS 2021) tested lexical, sparse, dense, late-interaction and re-ranking models across 18 datasets without domain training. BM25, the classic keyword ranking function, held up as a strong baseline, and several dense semantic models underperformed it once moved outside the data they were trained on. CRE transaction records are exactly that kind of out-of-domain text.

A second failure sits after retrieval. When comps are passed to a language model to summarize, position matters. Lost in the Middle (Liu et al., Stanford, 2023) found model performance is highest when relevant information sits at the start or end of the input and degrades significantly when it sits in the middle. A comp retrieved in position eight of fifteen can be retrieved and still effectively lost. This is the same problem retrieval-augmented generation has to solve for every grounded answer.

What does a missed comp cost?

A missed comp costs the difference between the value indicated by the partial set and the value indicated by the full set, and that difference lands on the buyer's equity. Because comp sets are small, one absent sale moves the average more than intuition suggests, especially when the absent sale traded at a different price.

Worked example, with all inputs stated:

Input

Value

Subject

120,000 SF industrial building

Stabilized NOI

$1,800,000

Comps returned by search

5, average cap rate 6.00%

Comps missed (coded as "flex" in the source)

2, at 6.75% and 6.50%

Full-set average cap rate

(5 × 6.00 + 6.75 + 6.50) ÷ 7 = 6.18%

Comp set

Cap rate

Indicated value

Partial (5 comps)

6.00%

$30,000,000

Full (7 comps)

6.18%

about $29,133,000

Difference

18 bps

about $867,000

The value gap is about 2.9%. Against a 35% equity check on $30 million, $10.5 million, the gap is about 8.3% of the equity. Nobody made a judgment error. The analyst evaluated every comp correctly. The loss entered at retrieval, and the review process had no way to see it.

Straight averaging is a simplification, and real work adjusts each comp. The point survives adjustment. An adjustment grid can correct a comp it has. It cannot correct for a comp it never received.

How do you audit recall in a comp set?

Recall is audited by testing the search, not by reading its output. The method is to plant sales you know should appear, change the query and watch what moves, and compare two retrieval methods that fail differently. What the tests surface is the set of relevant sales the original query silently dropped.

Test

How it works

What it catches

Known-item seeding

Keep a list of trades the team knows are relevant; confirm each is returned

Coding and field gaps

Filter loosening

Widen size band, date window and type by one step; review what enters

Hard-filter exclusions

Hybrid retrieval

Run keyword and semantic search separately; diff the results

Method-specific blind spots

Rank-depth check

Read past the default cutoff to position 20 or 30

Top-k truncation

Query logging

Save the exact query with the comp set

Makes the miss reproducible later

The last test matters most for accountability. A comp set without its query is a conclusion without its method. When a valuation is challenged a year later, the question is which search produced the comps, and whether the missing sale was reachable by it. Submarket definitions are a frequent culprit, which is why the submarket, not the metro should be the unit the query is built around.

Frequently Asked Questions

Is AI comparable selection more accurate than manual comp pulls?

It can be faster and broader, but accuracy depends on recall, which neither method reports on its own. An AI search that is never tested for recall is not more accurate than a manual pull. It is less visible.

Should comp search use keywords or semantic search?

Use both. Research on the BEIR benchmark found keyword ranking a strong baseline across domains, while semantic models catch differently worded records; the two miss different sales.

How many comps are enough?

There is no fixed number. The question that matters is whether the set contains every relevant sale that was available, which only a recall test can answer.

Can a reviewer catch a missing comp?

Only if the reviewer already knows the sale exists. Otherwise nothing on the page signals its absence, which is why recall must be tested rather than reviewed.

Conclusion

Comps are a retrieval problem first and a judgment problem second. Judgment gets the attention because it is visible. Retrieval gets none, because its failures leave no trace. The operator's job is to make the invisible error testable. Seed known sales, run two methods that fail differently, read past the cutoff, and keep the query with the conclusion. A comp set is only as complete as the search that built it, and the search will never say what it missed.