Two different bugs that look identical from the outside
An intermittent failure has exactly two possible shapes, and they call for opposite fixes: either your app is sending a slightly different request each time (a race condition, uninitialized state, a token that's sometimes stale), or your app is sending the exact same request every time and something downstream is responding inconsistently to it (load balancing across backend instances that don't agree, a flaky cache, rate limiting that only kicks in after enough requests). From the outside — "it worked, then it didn't" — these look like the same bug. They aren't, and fixing the client won't help if the problem is downstream, or vice versa.
The one test that separates them
Capture the exact request as sent, then send that identical request several times in a row and compare the results. If the request itself was byte-for-byte identical each time and the outcome still varied, the problem is downstream of your client — a different backend instance, cache state, or timing-dependent server behavior. If the requests actually differed from each other in a way you didn't intend, the problem is in your client, and you've just found the specific field that's inconsistent.
This is a more reliable diagnosis than reasoning about it, because "the request looks the same" and "the request is the same" are different claims — a header value, a generated request ID, or a timestamp can differ in ways that are easy to miss reading code but obvious once two actual captured requests are compared directly.
Common causes on each side
- Client-side (request varies): a race condition reading state before it's fully initialized, retry logic that fires before the first attempt has actually failed, a cached value that's sometimes stale and sometimes fresh, or a randomly-generated ID that occasionally collides with something the server rejects.
- Server/network-side (request is identical, response varies): a load balancer routing to backend instances running different code or with different cache state, rate limiting that only triggers past a request count, DNS resolving to a different IP between attempts, or a downstream dependency that's itself intermittently failing.
Why "it worked when I tried it once" proves nothing
A single successful retry is exactly as consistent with "the bug is gone" as it is with "the bug is still there and you got lucky." The only way to tell the difference is enough repeated attempts to see whether the outcome is actually stable — which is also why a single side-by-side comparison of one success and one failure is less conclusive than it feels: without knowing the failure reproduces reliably, you don't yet know whether the difference you're looking at explains it, or is just noise.
Replay ×N automates exactly this test, and Compare reads back whether a specific change actually explains a difference in outcome, once you know what to compare. For the exact walkthrough — repeat counts, what "consistent" actually means, and what the result does and doesn't prove — see reproducing a flaky request.