Skip to content

Results and reports

Every run of a load test ends with a summary: totals, per-request numbers, one chart point per second, the thresholds and a verdict. The app keeps the summaries of recent runs so you can look back, compare, and export them.

When a run ends (it finished, was stopped, or failed with an error), its summary is saved in the app’s data folder:

<app data folder>/load-runs/<workspace key>/<load test id>/<run id>.json

The history belongs to your computer, not the workspace: it is not committed to Git and teammates don’t see it. See Data locations.

  • The newest 30 runs of each load test are kept; older ones are deleted as new ones arrive.
  • Renaming a load test moves its history along, also while it runs.
  • Deleting a load test deletes its history.
  • The results header has a runs menu (for example “12 runs”). Each entry shows PASS, FAIL or ERROR, the date, “stopped” when it was stopped early, and its duration, req/s, p95 and error rate. Pick one to view it; Back to the latest run (or Back to the live run) returns.
  • The trash icon deletes the run on screen from the history. This can’t be undone.

With no run open, the results side shows the latest run of the test, from this session or from the history.

A finished run starts with a card: PASSED, FAILED or ERROR, a line such as “1 of 3 thresholds failed” (or the error message), the start date, the duration, and Stopped early when it was stopped. See Thresholds.

Latency says how long a request took. The Timing panel says where that time went, for the whole test:

PhaseFrom → toMeasured for
ConnectDNS, TCP, the proxy tunnel and TLS of a new connectionOnly requests that opened a new connection. The panel shows “New connection for x % of requests”.
Time to first byteRequest sent → first byte of the responseEvery answered request. This is the server’s time plus one network round trip.
TransferFirst byte → last byte of the responseEvery answered request.
Server-reportedWhat the server says it spent, from its Server-Timing headerOnly responses that had the header. The row is hidden when none had it.

Each phase shows the number of requests measured, p50, p95, p99 and max in milliseconds (the HTML report adds the average). The per-request table adds 1st byte p95 for each request.

With keep-alive on, most requests reuse a connection, so Connect covers only the few that opened one. Turn keep-alive off to measure connection setup on every request (see Options).

Time to first byte always includes the network. To see the server’s own time, have it send a Server-Timing header. Zorvik reads one number per response:

  • the dur of the metric named total (any case), when there is one;
  • otherwise the sum of the dur of every metric.
HeaderServer-reported time
Server-Timing: total;dur=12.312.3 ms
Server-Timing: db;dur=53, app;dur=47.2100.2 ms
Server-Timing: db;dur=5, total;dur=20, app;dur=720 ms
Server-Timing: cache;desc="Hit, fast";dur=2.5, miss2.5 ms
Server-Timing: db;desc="no duration"none (not counted)

Quoted values may contain commas and semicolons; dur may be quoted. Negative, non-numeric and missing durations are ignored.

When time to first byte is high but server-reported time is low, the time goes to the network, the proxy or queueing in front of the application.

Compare in the results header lists the other runs of the same test (runs that ended with an error are left out). Pick one to put a comparison panel above the charts: the run on screen against the earlier one, row by row.

RowBetter when
Requests / sHigher
Error rateLower
p50, p90, p95, p99 latency, Max latencyLower
p95 first byteLower
Request p95, one row per request (matched by request path)Lower

The change is shown in percent against the earlier run (“+12%”, “−3.4%”), green when better and red when worse. Changes under 1 % are noise between runs and have no colour. “new” means the earlier run’s value was 0, and ”–” means one of the runs has no data for that row. Stop comparing in the same menu removes the panel.

Compare runs with the same settings: a different model, stage plan or data file changes the numbers for reasons that have nothing to do with the server.

Export in the results header saves the run on screen:

FormatWhat you get
HTML report…One self-contained page: inline styles, SVG charts and a small script for their tooltips, no external files. Open it in any browser, attach it to a pull request, or archive it as a CI artifact.
JSON…The full summary as JSON, for your own tools.

The suggested file name is the test’s name with the run’s start date and time, for example Checkout smoke 2026-09-28 14-05.html.

zorvik load writes the same files with --html <file> and --json <file>:

Terminal window
zorvik load ./api "Checkout smoke" --html report.html --json summary.json
SectionContents
HeaderTest name, “Load test · Started 2026-09-28 12:05:00 UTC · ran duration”, “stopped early” when stopped, and the verdict: ✓ Passed, ✓ Passed (no thresholds), ✗ Failed (n of m thresholds) or ✗ Failed (error).
TilesRequests and req/s, error rate, p95 and p99 latency, p95 first byte, data in and out, connections, and when they apply: capture misses, dropped requests, peak generator CPU.
ThresholdsEach threshold, its actual value (“no data” when there was none) and ✓ Passed / ✗ Failed.
Over timeCharts of requests per second (completed and failed), latency (p50, p95, p99) and users or requests in flight. Point at a chart, tap it, or focus it with Tab and move with the arrow keys to see the time and every value at that moment. Runs longer than 600 seconds are shown in up to 600 columns: counts and p50 are averaged per column, p95 and p99 take the column’s highest value.
Latencymin, avg, p50, p90, p95, p99, p99.9 and max of all requests.
TimingThe phases above, with a note when no response had Server-Timing.
RequestsPer request: requests, req/s, error rate, p50, p95, p99, max, p95 first byte, capture misses (when any), data in.
Status codesEach status with its count and share.
Network errorsEach kind of network error with its count.

Latencies are in milliseconds, times since the epoch in milliseconds, rates in requests per second.

summary.json (shortened)
{
"startedAt": 1790517900000,
"durationMs": 60412,
"totals": {
"requests": 118230,
"errors": 12,
"errorRate": 0.0101,
"rps": 1957.1,
"bytesIn": 88712004,
"bytesOut": 10403840,
"latency": { "min": 1.2, "avg": 9.8, "p50": 8.1, "p90": 14.2, "p95": 18.9, "p99": 41.0, "p999": 120.3, "max": 311.7 },
"statusCodes": [[200, 118218], [503, 12]],
"errorKinds": [],
"dropped": 0,
"connections": 64,
"timing": {
"connect": { "count": 64, "avg": 3.1, "p50": 2.9, "p95": 5.0, "p99": 6.2, "max": 7.0 },
"ttfb": { "count": 118230, "avg": 9.1, "p50": 7.6, "p95": 17.8, "p99": 39.2, "max": 310.9 },
"transfer": { "count": 118230, "avg": 0.7, "p50": 0.4, "p95": 1.9, "p99": 3.3, "max": 12.0 },
"server": { "count": 0, "avg": 0, "p50": 0, "p95": 0, "p99": 0, "max": 0 }
},
"captureMisses": 0
},
"targets": [
{ "name": "List products", "request": "Products/List products.yaml", "metrics": { "requests": 88672, "errors": 0, "rps": 1467.8 } }
],
"points": [
{ "second": 0, "rps": 180, "errors": 0, "p50": 7.9, "p95": 15.0, "p99": 22.4, "active": 2, "target": 2 }
],
"thresholds": [
{ "label": "p95 < 300 ms", "metric": "p95", "op": "<", "value": 300, "target": null, "actual": 18.9, "passed": true }
],
"passed": true,
"stoppedEarly": false,
"error": null,
"peakCpuPercent": 142.0
}
FieldMeaning
startedAtStart of the run, Unix epoch milliseconds.
durationMsHow long the run really took, including the wait for requests in flight at the end.
totalsAll requests together (fields below).
targets[]The same numbers per request: name, request (path) and metrics.
points[]One entry per full second: second (from 0), rps (requests completed in that second), errors, p50, p95, p99, active (users, or requests in flight, at the end of the second) and target (the stage target: users at the end of the second, or the average rate over the second). The last, partial second is in the totals but not in points.
thresholds[]Each enabled threshold: label, metric, op, value, target, actual (null without data) and passed.
passedEvery threshold passed and there was no error.
stoppedEarlyStopped before the planned end.
errorWhy the run could not continue, or null.
peakCpuPercentHighest CPU used by Zorvik, in percent of one core summed over cores (200 = two full cores), or null when not measured.

totals and each target’s metrics:

FieldMeaning
requestsCompleted requests (answered or failed).
errorsNetwork errors plus HTTP status ≥ 400.
errorRateerrors as a percent of requests (0 to 100).
rpsrequests divided by the elapsed seconds.
bytesIn, bytesOutBytes received and sent over the wire.
latencymin, avg, p50, p90, p95, p99, p999, max in ms (all 0 without data).
statusCodes[status, count] pairs, most frequent first.
errorKinds[kind, count] pairs for network errors, most frequent first (kinds below).
droppedRequest rate: requests not started because maxInFlight was reached.
connectionsNew connections opened.
timingconnect, ttfb, transfer, server: each count, avg, p50, p95, p99, max (count 0 means no data).
captureMissesCaptures that found nothing.
KindMeaning
timeoutThe request timed out.
connectCould not connect (refused, unreachable).
dnsThe host name was not found.
tlsThe TLS handshake failed (for example, an untrusted certificate).
proxyThe proxy refused or failed.
protocolThe server broke the HTTP protocol.
ioThe connection broke while reading or writing.
cancelledCut off by a stop, or by the end of the run after the grace period.
invalidRequestThe request could not be built (for example, a rendered URL that isn’t valid).
notAllowedThe request went to a host the run was not started for (a captured host), or one an AI agent was not allowed to reach.
tooManyRedirectsListed for completeness; load tests don’t follow redirects.

Every run keeps one chart point per second (86,400 for a day). While a run is live, the app keeps up to 7,200 points on screen and merges neighbouring seconds beyond that; the saved summary keeps every second.