One retry is not evidence
A single successful retry is exactly as consistent with "the bug is gone" as it is with "the bug is still there and you got lucky" — there's no way to tell those apart from one data point. The difference between a flaky request and a flaky network covers the deeper question of whether the client or the server is actually the inconsistent side; this article picks up from there, once you're ready to actually run the test rather than reason about it.
What Replay ×N actually does
Replay ×N resends the exact same captured request a number of times you choose, from 1 up to 10, and reports whether the outcome, timing, and body stayed consistent across the attempts. It's worth being precise about the shape of that test: each attempt runs one at a time, in sequence — the next one doesn't fire until the previous one has finished — using its own fresh connection each time. That means it's evidence about whether one request is stable across repeated, isolated attempts. It is not a concurrency test: a backend that only misbehaves when many requests land on it at once will look perfectly consistent here, because Replay ×N never sends more than one at a time.
How it decides "consistent"
The check runs in a specific order, and it stops at the first thing that actually differs rather than reporting everything at once:
- Outcome first — the HTTP status (or a connection error) across all attempts. If these don't all match, that's the finding, and it's the most fundamental one: the request didn't even fail or succeed the same way twice.
- Latency, only if the outcome was consistent — flagged as unstable only when the slowest attempt is at least three times the fastest and the actual gap is at least 100ms, so ordinary jitter doesn't get reported as a finding.
- Body, only if outcome and latency were both stable — an exact comparison of the response body across every attempt.
Only the single most specific issue is ever surfaced — if the outcome itself is inconsistent, you won't also get a separate note about latency, because the outcome difference is the more fundamental problem and the more useful thing to chase first.
The special case: intermittent authentication
When the outcomes split cleanly between success-like and auth-failure-like statuses across the attempts, that gets its own specific conclusion rather than being reported as generic inconsistency — because it points somewhere concrete (a token expiring mid-test, a refresh race) rather than nowhere in particular. If you see this, the next step isn't more repeats; it's checking exactly when the token was issued relative to when the failing attempts ran.
What the result actually proves
A clean "consistent" result across ten sequential attempts is real evidence that this specific request, run this way, isn't flaky in isolation. It isn't proof the endpoint is safe under concurrent load, and it isn't evidence about a different request — different user, different input — even if it looks superficially similar. Once you do have a genuinely inconsistent result, pick one succeeding attempt and one failing attempt and open Compare on that pair — Replay ×N tells you that something varies; Compare is how you find out what.