Glossary
Confusion Matrix
A confusion matrix is a table that counts a classifier's true positives, false positives, true negatives, and false negatives against known correct labels. In commercial real estate document extraction, it shows exactly where a model that tags lease clauses or field types agrees with ground truth and where it fails.
How a Confusion Matrix Works
A confusion matrix works by sorting every prediction into one of four cells defined by two axes: what the model predicted and what the label actually is. The four cells are true positive (TP), false positive (FP), true negative (TN), and false negative (FN). From these four counts, every common classification metric is derived, so the matrix is the raw material behind precision, recall, and accuracy.
The four cells map cleanly onto a document-extraction task. Suppose a classifier decides whether a given lease clause is a renewal option. A true positive is a renewal option the model correctly flags. A false positive is a non-renewal clause the model wrongly flags. A false negative is a renewal option the model misses. A true negative is a non-renewal clause the model correctly ignores.
Cell | Meaning in clause tagging |
|---|---|
True positive (TP) | Renewal option correctly flagged |
False positive (FP) | Non-renewal clause wrongly flagged |
False negative (FN) | Renewal option missed |
True negative (TN) | Non-renewal clause correctly ignored |
Three metrics fall out of the counts. Precision is TP / (TP + FP), the share of flagged clauses that were real. Recall is TP / (TP + FN), the share of real clauses that were caught. Accuracy is (TP + TN) / total, the share of all predictions that were right. Take a run over 1,000 clauses with TP = 102, FP = 30, FN = 18, and TN = 850. Precision is 102 / 132, or 77.3%. Recall is 102 / 120, or 85.0%. Accuracy is 952 / 1,000, or 95.2%.
Why a Confusion Matrix Matters
A confusion matrix matters because it separates two failure modes that a single accuracy number blends together. A model can miss real renewal options (false negatives) or invent ones that are not there (false positives), and these errors carry different costs. Missing a renewal option can distort a hold-period model, while a false flag only costs a reviewer's glance. The matrix shows which error the model actually makes.
The stakes rise when classes are imbalanced, which is the norm in extraction. Most clauses in a lease are not the specific clause you are hunting for. As the Wikipedia entry on the accuracy paradox notes, if one class appears in 99% of cases, a model that always predicts that class scores 99% accuracy while catching none of the minority class. A confusion matrix exposes that failure instantly: the TP cell reads zero even as accuracy looks excellent. This is the single most useful thing a confusion matrix does, turning a flattering headline number into an honest breakdown.
Example
A confusion matrix is clearest as a filled 2x2 grid. The table below scores a renewal-option classifier over 1,000 lease clauses, of which 120 are true renewal options and 880 are not.
Predicted: renewal | Predicted: not renewal | |
|---|---|---|
Actual: renewal (120) | TP = 102 | FN = 18 |
Actual: not renewal (880) | FP = 30 | TN = 850 |
Deriving the metrics from these cells:
Metric | Formula | Calculation | Result |
|---|---|---|---|
Precision | TP / (TP + FP) | 102 / 132 | 77.3% |
Recall | TP / (TP + FN) | 102 / 120 | 85.0% |
Accuracy | (TP + TN) / 1,000 | 952 / 1,000 | 95.2% |
F1 score | 2 x P x R / (P + R) | 2 x 0.773 x 0.850 / 1.623 | 81.0% |
Accuracy reads 95.2%, but the matrix shows the model still misses 18 of 120 real renewal options and over-flags 30 clauses. Precision and recall, both in the low-to-mid 80s, describe the model far more honestly than the single accuracy figure.
Variations and Edge Cases
A confusion matrix is a 2x2 grid only for binary classification. For a model that assigns each field a type from several categories, the matrix expands to an N x N grid where cell (i, j) counts items whose true class is i but predicted class is j. Correct predictions sit on the diagonal, and every off-diagonal cell names a specific confusion, such as a "commencement date" repeatedly misread as an "expiration date."
Situation | How the matrix behaves |
|---|---|
Binary classification | Standard 2x2 grid; TP, FP, TN, FN |
Multi-class (N classes) | N x N grid; correct predictions on the diagonal, confusions off it |
Class imbalance | Accuracy inflates; read precision and recall per class instead |
Rare positive class | A large TN count can hide a near-empty TP cell |
Class imbalance is the edge case that trips up readers most. When the negative class dominates, the TN cell swamps the grid and pushes accuracy toward 100%, so per-class recall reveals whether the minority class is found at all.
Confusion Matrix vs Accuracy
A confusion matrix is often confused with accuracy, but one contains the other. A confusion matrix is the full four-cell (or N x N) table of correct and incorrect predictions broken out by type. Accuracy is a single number derived from that table: the share of all predictions that were correct, (TP + TN) / total.
The practical difference is resolution. Accuracy collapses every kind of error into one figure and hides the balance between false positives and false negatives. A confusion matrix keeps them separate, which is why a model at 95.2% accuracy can still be missing 15% of the clauses that matter. Accuracy answers "how often is the model right," while a confusion matrix answers "right and wrong about what."
Frequently Asked Questions
What does a confusion matrix show? A confusion matrix shows the counts of a classifier's true positives, false positives, true negatives, and false negatives against known labels. It reveals not just how often the model is wrong but the specific way it is wrong, separating missed items from false flags.
How do you calculate precision and recall from a confusion matrix? Precision is true positives divided by all predicted positives, TP / (TP + FP). Recall is true positives divided by all actual positives, TP / (TP + FN). Both are read directly off the matrix cells with no additional data.
Why is accuracy misleading without a confusion matrix? Accuracy blends every error into one number, so a model can score 99% by always predicting the majority class while catching none of the minority class. A confusion matrix exposes this by showing a true-positive cell of zero despite high accuracy.
Related Terms
Precision and Recall