
Benchmarking OpenAI model latency: time to first text and a correct answer
Revised
Published
~ 7 min read
A fast answer is not useful when it is wrong. That makes latency on its own a poor way to choose an OpenAI model or reasoning level. The useful question is: which configuration produces an acceptable result fastest, reliably and at a reasonable cost?
This article uses two separate benchmarks. The first isolates streaming behaviour with a tightly controlled output. The second measures a small, automatically scored reasoning task. Together they distinguish raw serving speed from time to a correct answer.
Model aliases, capabilities and defaults change. Check the current OpenAI model guidance before running the benchmark. Use a dated model snapshot when one is available, and always record the requested and returned model IDs.
What to measure
The Responses API emits typed streaming events. response.created means the request has started producing events; it is
not a text token. For a text interface, time to first text begins with the first non-empty
response.output_text.delta event.
Track these timings separately:
- Time to first event: request start to the first server-sent event. Useful for diagnosing connection and request acceptance latency.
- Time to first text (TTFT): request start to the first non-empty text delta. This is the streaming metric users notice.
- Time to first usable unit: request start to the first complete sentence, JSON object or tool call that the client can act on.
- Total request time: request start to
response.completed. - Generation time: first text delta to completion.
- End-to-end time: total time including retries and backoff.
Do not estimate tokens by counting stream chunks. A chunk is a transport event and may contain part of a token or several
tokens. Read token usage from the completed response instead, including input tokens, cached input tokens, output tokens
and reasoning tokens. The reasoning guide documents the
output_tokens_details.reasoning_tokens field.
Latency still needs an outcome measure. For each run, record whether the response passed the task’s acceptance criteria. That lets you compare pass rate, time to a correct answer and, after applying the prices current on the test date, cost per correct answer.
Use two different tasks
One prompt cannot measure both serving performance and reasoning quality well.
1. Controlled streaming task
Ask for a fixed-size answer, such as a plain-text explanation of exactly 120 words. This keeps output length reasonably stable, making TTFT and generation throughput easier to compare.
The scorer only checks the constraint. It does not claim to measure intelligence.
2. Scored reasoning task
Use a task with one verifiable answer. The example script contains a dependency-scheduling problem and checks the returned JSON automatically. It is deliberately small enough to keep the example understandable.
For a production decision, replace it with representative cases from the application: code repairs checked by tests, structured extraction checked against labelled data, or tool workflows checked for the expected actions and final state. The acceptance test matters more than the benchmark prompt.
Experimental controls
Use the same controls for every model and reasoning level:
- Pin exact model IDs or dated snapshots where available. Record both the requested and returned model IDs.
- Record the UTC timestamp, Node version, API service tier and test location.
- Run a few warm-ups and exclude them from the results.
- Randomise and interleave configurations so one model is not always tested during the same short time window.
- Use one in-flight request for an unloaded latency benchmark.
- Run load tests separately at explicit concurrency levels.
- Keep prompts, output limits, scoring and every supported request parameter fixed across comparable configurations.
- Record failed, incomplete and rate-limited requests rather than dropping them.
Sampling controls are not portable across every model and reasoning setting. Check the selected model’s documentation
before setting temperature, top_p, seed or similar parameters. If the API accepts one, pin and record it. These
controls reduce variation; they do not make model output deterministic.
Thirty successful runs per configuration is a reasonable starting point for a directional median and P90. Use a larger sample if tail latency or small differences will drive a production decision. Prefer p50 and p90 over min, max and mean; a single outlier can dominate the latter statistics.
Prompt caching is another experimental variable. OpenAI’s
prompt caching works on reusable prompt prefixes and
reports cache activity in usage.input_tokens_details. Capture cached_tokens and cache_write_tokens. Label each run
as uncached, cache write, cache hit, or both. Keep failures in the overall requested-configuration summary, then break
completed requests down by cache state separately.
Run the benchmark
The complete dependency-free Node 22 script lives in
scripts/measure-openai-latency.mjs.
It calls the Responses API using Node’s built-in fetch, writes per-run metrics to JSON and prints a compact summary.
The generated openai-latency-*.json reports are ignored by Git.
Start with a cheap smoke test:
RUNS=2 node scripts/measure-openai-latency.mjs
Once the configuration and scorers are correct, collect a more useful sample:
RUNS=30 \
WARMUPS=1 \
MODELS=gpt-5.6-sol,gpt-5.6-terra,gpt-5.6-luna \
EFFORTS=none,low,medium \
node scripts/measure-openai-latency.mjs
The defaults are examples, not a permanent recommended model list. Use a snapshot when the selected model offers one. Otherwise record the requested and returned model IDs with the test date before publishing results.
The report omits model output text by default. That keeps real prompts from accidentally leaving sensitive responses in a benchmark artefact. If a synthetic test needs the text for debugging, opt in explicitly:
INCLUDE_MODEL_OUTPUTS=true RUNS=2 node scripts/measure-openai-latency.mjs
Review the generated file before sharing it.
The important part of the stream handling is the distinction between events:
if (firstEventAt === null) firstEventAt = now;
if (event.type === "response.output_text.delta" && event.delta) {
if (firstTextAt === null) firstTextAt = now;
output += event.delta;
}
if (event.type === "response.completed") {
completedAt = now;
completedResponse = event.response;
}
The script deliberately runs requests sequentially. Increasing the request-start rate while earlier responses are still running changes the question from “how long does one request take?” to “how does this configuration behave under load?” Both are useful, but they should not share a results table.
Each attempt has a two-minute timeout by default. Set REQUEST_TIMEOUT_MS if the chosen model or task legitimately needs
longer, and keep the same value across comparable configurations.
Interpret the results
Start with the scored task rather than the raw latency task.
For each configuration, compare:
- task pass rate;
- p50 and p90 time for passing responses;
- p50 and p90 TTFT;
- input, cached, output and reasoning token usage;
- errors, incomplete responses and retries;
- cost per passing response, using prices recorded on the test date.
Use the overall summary for completion and pass rates because it includes failed requests. The cache-state breakdown only contains completed requests. Read those rows separately for latency and token comparisons; combining cache hits with uncached or cache-write requests can make a model look faster or cheaper for reasons unrelated to its serving performance.
A latency-versus-pass-rate scatter plot is usually more useful than a model leaderboard. Configurations on the Pareto frontier are the interesting ones: no other tested configuration is both faster and more reliable.
Be careful with apparently fast failures. If a model returns in two seconds but only passes 60% of cases, its expected time and cost per successful outcome may be worse than a slower model with a 98% pass rate. Retrying poor answers also adds latency that a successful-request-only chart hides.
Optimise after measuring
Use the lowest reasoning effort that passes the application’s evals. The current
GPT-5.6 model catalog lists none, low, medium, high, xhigh and
max for Sol, Terra and Luna; this script uses a cheaper subset by default. Check the model page again before changing
the list or adding another model.
OpenAI’s latency optimisation guide recommends looking beyond prompt length. Useful changes include generating fewer output tokens, choosing an appropriately sized model, reducing sequential requests, parallelising independent work and streaming useful progress.
Apply those changes to the full workflow:
- Reduce output length without removing information required by the scorer.
- Prefer a smaller model only when it maintains the required pass rate.
- Keep stable prompt content at the front so eligible requests can reuse cached prefixes.
- Measure time to the first valid tool call for tool-using applications.
- Measure the browser, API, tools and data stores together for user-facing latency.
Streaming improves perceived responsiveness, but it does not make the model finish sooner. A background job may care only about total time, while an interactive interface may care more about first usable text. Choose the metric that matches the product.
Limits of the comparison
These results remain specific to the prompts, account, service tier, network path and time of the run. Model behaviour and serving infrastructure can change. Safety checks, tool calls, structured output and long context can also alter the shape of a response.
Client-side timings cannot cleanly separate every part of network and server processing. They are still valuable because they measure what the application and user actually experience.
The benchmark therefore should not end with a universal “fastest model”. It should identify the fastest configuration that meets the quality, reliability and cost requirements of a particular workload.