Glossary
Data Annotation
Data annotation is the process of labeling raw documents so a machine learning model can learn to extract specific fields. In commercial real estate, annotators mark spans in leases and offering memorandums and tag each as base rent, commencement date, or cap rate, turning unstructured pages into the labeled examples a model trains and is measured against.
How Data Annotation Works
Data annotation is a structured pipeline, not free-form highlighting. It starts with a schema: a fixed list of fields to capture and a definition of each, so that "base rent" means the same thing on every document. Labelers then read each page and mark the span that carries each field, and a review pass checks the marks before they enter the training set.
The schema is what makes annotation repeatable. A CRE lease schema might define twenty fields, each with a rule for ambiguous cases: base rent is the initial monthly figure before escalations, commencement date is the rent start date rather than the signing date, and a missing field is tagged absent rather than guessed. Clear rules keep two annotators reading the same lease the same way.
Quality is measured by inter-annotator agreement: the degree to which independent labelers assign the same label to the same span. The standard metric is Cohen's kappa, introduced by Jacob Cohen in 1960, which corrects raw agreement for the rate expected by chance. Landis and Koch (Biometrics, 1977) set the interpretation scale still used today: kappa of 0.61 to 0.80 is substantial agreement and 0.81 to 1.00 is almost perfect. A recent named entity recognition dataset for the Kurdish Sorani language reported kappa of 0.92 (arXiv:2511.22315), landing in the almost-perfect band.
Step | What happens |
|---|---|
Schema | Fields and their definitions are fixed before labeling begins |
Label | Annotators mark the span carrying each field on each document |
Agreement | A subset is double-labeled and scored with Cohen's kappa |
QA | A reviewer adjudicates disagreements and corrects errors |
Release | Approved labels enter the training or evaluation set |
Why Data Annotation Matters
Data annotation matters because label quality sets a ceiling on model accuracy: a model trained on noisy labels learns the noise, and a model measured against wrong labels reports a false score. Research on ground-truth annotation stresses that the accuracy of the labels is an upper bound on the accuracy the model can reach (arXiv:2504.09341). Garbage labels produce a garbage ceiling no amount of model tuning can lift.
For a CRE team building a lease or offering-memorandum extractor, this is the difference between a model that pays off and one that quietly misreads terms. If annotators disagree on where commencement date lives one time in five, the model inherits that confusion, and every downstream number carries the error. Disciplined annotation with measured agreement is the cheapest reliability an extraction system can buy, because it is fixed once at the source rather than caught deal by deal.
Example
Data annotation is easiest to see on a single lease, field by field. A labeler opens a retail lease, works through the schema, and marks the span for each field. Each approved mark becomes ground truth: the answer the model is trained toward and graded against.
Field | Annotated span (marked in the lease) | Becomes ground truth as |
|---|---|---|
Tenant name | "Cedar Park Grocery, LLC" | tenant = Cedar Park Grocery, LLC |
Base rent | "$18,500 per month" | base_rent = 18500 |
Commencement date | "Rent shall commence on March 1, 2026" | commencement = 2026-03-01 |
Term length | "a term of one hundred twenty (120) months" | term_months = 120 |
Renewal option | "two (2) options of five (5) years each" | renewals = 2 x 60 months |
If a second annotator marks the signing date instead of the rent-start date for commencement, the two disagree on that field. That disagreement is caught in the agreement check, adjudicated against the schema rule, and corrected before either label enters the training set. Across a 200-lease batch of roughly 3,000 field marks, that adjudication is what separates a clean ground-truth set from a polluted one.
Variations and Edge Cases
Data annotation is not always full manual labeling of every document. Teams trade label volume against label cost using the variants below, each of which changes where the human effort goes.
Variant | Behavior |
|---|---|
Weak supervision | Labels are generated from rules or patterns, cheaply and noisily, then treated as approximate rather than exact |
Active learning | The model selects the documents it is least sure about and asks a human to label only those |
Pre-labeling | A model proposes labels and annotators correct them, faster than labeling from scratch |
Span vs classification | Field extraction marks text spans; document sorting assigns one label to the whole file |
Absent-field handling | A field truly missing must be labeled absent, not skipped, or the model learns to ignore it |
Data Annotation vs Training Data
Data annotation is often confused with training data, but one produces the other. Data annotation is the process of applying labels to raw documents. Training data is the finished product: the collection of documents plus their approved labels that a model actually learns from. Annotation is the verb, training data is the noun.
The distinction matters when a model underperforms. If the training data is too small, the fix is to annotate more documents. If it is large but the model is inconsistent, the fix is usually annotation quality, not volume: a schema too loose or an agreement score too low to trust. Naming the two separately keeps a team from labeling ten thousand more leases when the first ten thousand were labeled inconsistently.
Frequently Asked Questions
What is data annotation in machine learning? Data annotation is the process of labeling raw data so a model can learn from it. For document extraction, annotators mark the span in each lease or offering memorandum that carries a field, and those approved marks become the training and evaluation examples the model uses.
How is data annotation quality measured? Annotation quality is measured by inter-annotator agreement, the rate at which independent labelers assign the same label to the same item. The standard metric is Cohen's kappa, and on the Landis and Koch scale a kappa of 0.61 to 0.80 is substantial and 0.81 to 1.00 is almost perfect.
Does data annotation limit model accuracy? Yes. The accuracy of the labels is an upper bound on the accuracy the model can reach. A model trained on inconsistent labels learns the inconsistency, and a model measured against wrong labels reports a score that does not hold up on real documents.