Glossary

Inference Latency

Inference latency is the time a model takes to produce output for one input, measured from the moment a request arrives to the moment the response completes. In document processing it sets how fast a single lease, rent roll, or offering memorandum returns extracted fields, and it compounds directly into cost at portfolio scale.

How Inference Latency Works

Inference latency is the sum of two stages: prefill and decode. Prefill reads the entire input prompt at once and produces the first output token. Decode then generates the remaining tokens one at a time, each conditioned on the last. Prefill is compute-bound and parallel, while decode is memory-bound and sequential.

Two metrics split this cleanly. Time-to-first-token (TTFT) measures the prefill stage: how long a reader waits before the first character appears. Time-per-output-token (TPOT), the inverse of tokens per second, measures each decode step after the first. Total latency for one document is roughly TTFT plus TPOT multiplied by the output token count.

MLCommons, the consortium that publishes the MLPerf benchmark, set concrete targets in its MLPerf Inference v5.0 release (April 2025). For the Llama 3.1 405B Instruct server scenario it fixed a 99th-percentile TTFT of 6 seconds and a 99th-percentile TPOT of 175 milliseconds. For the smaller Llama 2 70B interactive benchmark it set a 99th-percentile TTFT limit of 450 milliseconds and a floor of 25 tokens per second, citing analysis that a 20 to 50 token-per-second generation rate is the band users perceive as seamless.

Batch size is the lever that turns latency into throughput. Serving many requests in one batch reuses the same weight reads and raises total tokens per second, but it can lengthen any single request's latency, since that request now waits behind others. At a decode rate of 40 tokens per second, a 400-token response takes 10 seconds, so one stream clears roughly 300 short documents per hour before batching multiplies the count.

Why Inference Latency Matters

Inference latency matters because it fixes both the throughput and the unit cost of every document a pipeline touches. A model that decodes at 40 tokens per second and one that decodes at 20 tokens per second differ by 2x in documents cleared per hour on the same hardware, which is a 2x difference in the compute bill for the same workload.

At the scale of a real acquisition pipeline the arithmetic is unforgiving. Screening 5,000 offering memoranda a quarter, each producing 1,500 tokens of extracted output, is 7.5 million output tokens of decode work. Halving TPOT halves the seconds of GPU time, and GPU time is the cost. Latency is not a user-experience nicety here; it is the denominator under cost per document.

Latency also gates what is even feasible interactively. An underwriter who wants a lease abstracted while on a call needs a first token in under a second, not six. The one line worth remembering: inference latency is the tax paid on every token, and at document scale that tax is the budget.

Example

Inference latency is easiest to read across documents of different sizes, where each row traces input tokens, output tokens, latency, single-stream throughput, and cost. The table below uses stated rates as worked-example inputs: prefill at 2,000 tokens per second, decode at 40 tokens per second, and an assumed price of $2.50 per million input tokens plus $10.00 per million output tokens.

Document

Input tokens

Output tokens

TTFT

Decode time

Total latency

Docs/hour (1 stream)

Cost/doc

Rent roll

3,000

400

1.5s

10.0s

11.5s

313

$0.012

Lease

12,000

800

6.0s

20.0s

26.0s

138

$0.038

Offering memorandum

40,000

1,500

20.0s

37.5s

57.5s

63

$0.115

Read the offering memorandum row. TTFT is 40,000 divided by 2,000, or 20 seconds. Decode is 1,500 divided by 40, or 37.5 seconds, for 57.5 seconds total. Cost is 40,000 over one million times $2.50, plus 1,500 over one million times $10.00, which is $0.100 plus $0.015, or $0.115. One stream clears 3,600 divided by 57.5, about 63 of these an hour. The long input and long output make it the slowest and most expensive document in the set.

Variations and Edge Cases

Inference latency shifts with a handful of levers that change the two-stage picture. Context length inflates prefill because the model must read every prompt token before the first output token; a 40,000-token offering memorandum carries a far heavier TTFT than a one-page rent roll. The common variants are below.

Variant

Effect on latency

Longer context

Raises TTFT; prefill reads the full prompt before token one

Larger batch

Raises total throughput, can raise per-request latency

Longer output

Raises total time linearly through the decode stage

Larger model

Raises both TTFT and TPOT for the same tokens

Streaming output

Lowers perceived latency by showing tokens as they decode

Caching matters too. When many documents share a prompt prefix, such as a fixed extraction instruction, reusing the cached prefill states cuts TTFT on every request after the first.

Inference Latency vs Throughput

Inference latency is often confused with throughput, and the two pull against each other. Inference latency is the time to finish one request. Throughput is the number of tokens or documents finished per unit time across all concurrent requests. A system tuned for low latency serves each request fast but may leave hardware idle; a system tuned for high throughput packs large batches and keeps hardware busy but makes any single request wait.

The MLPerf server scenario exists precisely because unconstrained batching inflates throughput while ruining latency. Hold TTFT and TPOT to a limit, as MLCommons does, and batch sizes shrink and throughput falls. For document processing the choice is workload-dependent: interactive abstraction on a call optimizes latency, while an overnight batch of 5,000 memoranda optimizes throughput.

Frequently Asked Questions

What is inference latency? Inference latency is the time a model takes to produce output for a single input, measured from request arrival to completed response. It is the sum of prefill, which produces the first token, and decode, which generates each remaining token in sequence.

What is the difference between TTFT and TPOT? Time-to-first-token (TTFT) measures the prefill stage, the wait before the first output token appears. Time-per-output-token (TPOT), the inverse of tokens per second, measures each decode step after the first. Total latency is roughly TTFT plus TPOT times the number of output tokens.

How does inference latency affect document processing cost? Inference latency sets how much compute time each document consumes, and compute time is the cost. Decoding a 1,500-token extraction at 40 tokens per second takes 37.5 seconds; halving that decode rate doubles the seconds of hardware time and the cost per document.

What is a good inference latency for interactive use? MLCommons found a generation rate of 20 to 50 tokens per second is the band users perceive as seamless, and set a 450-millisecond first-token limit for its interactive benchmark. Interactive document review generally needs a first token in well under one second.

Related Terms