TL;DR: Calling a live payment gateway, auth provider, or any third-party API from an end-to-end test is one of the fastest ways to end up with a flaky suite. Rate limits, sandbox downtime, and plain network latency fail your build for reasons that have nothing to do with your code. Google's own research on its test suite found that 16 percent of its tests show some level of flakiness, and external dependencies are a well documented part of that picture. Playwright's page.route() lets you intercept those requests inside the browser and answer with a response you control instead, so a suite that used to slow down or fail whenever a sandbox API hiccupped now runs the exact same scenario every time.
Quick answers
What is API mocking in end-to-end testing?
API mocking means intercepting a request your app makes to a real service, a payment gateway, an auth provider, any third-party API, and returning a response you control instead of letting the request reach the real service. The application under test cannot tell the difference. It gets a response shaped exactly like the real one, just without the network round trip, the sandbox account balance, or the rate limit attached to it.
How do you intercept API calls in Playwright?
Call page.route() (or context.route() to cover every page in that browser context) with a URL pattern and a handler function. Playwright pauses any matching request before it leaves the browser. Inside the handler you call route.fulfill() to answer with your own mock body and status code, route.continue() to let the request through unchanged, or route.abort() to simulate a dropped connection.
Does mocking third-party APIs make tests less realistic?
For the third-party dependency itself, yes, on purpose. That is the point. You are not testing whether Stripe or Auth0 works, they already test that. You are testing whether your application handles their response correctly, including the responses that are hard to trigger on demand: a declined card, a 429 rate limit, a 500 from their side. Keep a small number of real runs against a sandbox environment for the handful of flows where the live integration itself is what you are validating, and mock everything else.
The Risk of Calling Live Payment Gateways and Auth Providers in Test Runs
Every end-to-end test that calls a real third-party API is borrowing that provider's uptime, rate limits, and response time as part of your own test result. When a payment sandbox is slow, your checkout test is slow. When a shared sandbox account hits its daily rate limit because three other teams are using it too, your auth tests start failing for a reason no one on your team caused.
Payment gateways and auth providers are the two riskiest categories specifically because they sit on the critical path of almost every user journey worth testing, and because triggering their edge cases on demand, an expired card, a locked account, a token that expires mid-session, is often impossible against a real sandbox. You cannot reliably ask a live payment provider to return a decline code on request every single CI run.
Some sandbox environments also rate limit by account rather than by request, so a CI pipeline running dozens of times a day can burn through a shared quota fast. Mocking removes the dependency entirely: no shared quota, no sandbox maintenance window, no test that fails at 3am because a third party pushed a deploy of its own. If you are also evaluating the client-side tools your team uses to explore these same APIs by hand, see our roundup of Postman alternatives; this guide is about what happens once your automated E2E tests talk to those APIs, not the exploratory tooling around them.

Intercepting HTTP Requests in Playwright Using page.route()
Set up the route before you trigger whatever action fires the request, a button click, a form submit, a page load. page.route(urlPattern, handler) takes a glob pattern or full URL to match against and a handler function that receives the intercepted route object. Playwright checks every outgoing request against every registered pattern and hands matching ones to your handler instead of letting them reach the network.
Inside the handler, route.fulfill({ status: 200, contentType: 'application/json', body: JSON.stringify({ success: true, transactionId: 'txn_test_001' }) }) answers the request with your mock payload instead of the real provider's response. The application under test receives that JSON exactly as if the real API had sent it.
Not every request needs a mock. Call route.continue() to forward a request unchanged when you only care about mocking one or two specific calls on the page, and route.abort() when you want to simulate a connection that never completes at all, a dropped network, a DNS failure, a timeout, rather than a response with an error status code.
Simulating Network Delays, Rate Limits (429), and Server Errors (500)
The responses your UI handles worst are usually the ones nobody tests on purpose: a request that hangs for eight seconds, a provider returning 429 because you hit its rate limit, a 500 because its own database is having a bad day. These are exactly the states route.fulfill() makes trivial to reproduce on demand, every single run.
For a rate limit, fulfill the route with status: 429 and a Retry-After header, then assert your UI actually shows a retry message instead of a blank error. For a provider outage, fulfill with status: 500 and confirm your app falls back gracefully instead of leaving the user on a frozen spinner. Both take the same shape: set the status code and body your provider would send on a bad day, then check your own code's reaction to it.
To simulate a slow response rather than a failed one, await a delay inside the handler before calling route.fulfill(), for example await new Promise(resolve => setTimeout(resolve, 4000)) before fulfilling. That reproduces a slow third-party API on a fast local network, so a loading state or timeout handler gets tested even when the real provider happens to answer in under a second that day.

HAR Recording and Replay for Deterministic Test Environments
Hand-writing a route.fulfill() for every endpoint works for a handful of calls, but a checkout flow or a login flow can touch a dozen third-party endpoints in one journey. Playwright's routeFromHAR() records and replays an entire session's worth of network traffic at once instead, and it is covered in full in Playwright's own mocking documentation.
Record once against a real or sandbox environment with context.routeFromHAR('./tests/har/checkout.har', { url: '**/api/**', update: true }). With update: true, Playwright writes every matching request and response into that HAR file as your test runs, instead of serving from it.
Drop update: true for every run after that, and Playwright serves matching requests straight from the recorded file with no network call at all. Add notFound: 'abort' so a request that is not in the recording fails loudly instead of silently reaching the real network, which is exactly the kind of silent flakiness this whole approach exists to remove. One caveat worth knowing upfront: HAR replay matches the URL and HTTP method strictly, and for POST requests it also matches the request body strictly, so re-record the file whenever the request shape itself changes, not just when the response does.
Every mock in this guide is something your team has to write and keep up to date by hand, one endpoint, one status code, one HAR file at a time. ContextQA's API testing is built so that kind of coverage does not turn into its own side project as your API surface grows. See it on a 15-minute demo.
Frequently Asked Questions
Should I ever run end-to-end tests against a real sandbox API?
Yes, for a small number of flows where the live integration itself is what you are validating, a new payment provider integration, a webhook you have never received in production yet. Keep those runs few and clearly separated from your main mocked suite, since they will always be slower and less reliable than the mocked version of the same test.
Does route.fulfill() work for GraphQL requests too?
Yes. Most GraphQL APIs post every query and mutation to a single endpoint, so match on that URL, then inspect route.request().postDataJSON() inside the handler to see which operation is being called and branch your mock response accordingly before calling route.fulfill().
What is the difference between page.route() mocks and HAR replay?
page.route() with hand-written fulfill() calls is the right tool for a small number of endpoints or for edge cases you need exact control over, a specific error code, a specific delay. HAR replay is the right tool when you want an entire realistic session, many endpoints, real-shaped payloads, handled at once without writing a mock for each one by hand.
Bottom line
Mocking third-party APIs is not a workaround for a flaky suite, it is what a deterministic suite is supposed to do with anything outside your own code's control. page.route() handles the small, precise cases: a specific error code, a specific delay, a response you cannot get a real sandbox to produce on demand. routeFromHAR() handles the large ones: an entire checkout or login flow's worth of traffic, recorded once and replayed the same way every run after that. Between the two, there is very rarely a good reason left for an end-to-end test to depend on a real payment gateway or auth provider actually being up. If keeping that mock library maintained by hand is starting to feel like its own project, that is exactly the gap ContextQA's platform is built to close.