Replay: Record-and-Replay Network Tests with Swift Testing
⚠️ Speculative Architecture & Preview: This article discusses future system iterations (e.g., iOS 27, Xcode 27) as conceptual planning and architectural design patterns. Technical details represent previews and proposals rather than finalized APIs.
Replay: Record-and-Replay Network Tests with Swift Testing
How many JSON fixtures in your test bundle are lying to you right now? A hand-written stub compiles forever, survives every refactor, and silently stops matching the server the day someone renames a field. Replay Swift Testing replaces the guesswork with a different contract: record real HTTP traffic once into a HAR file, then replay it byte-for-byte inside Swift Testing.
It is not a new idea. Ruby’s VCR did this in 2010, Python’s VCR.py and pytest-recording followed, and Venmo’s DVR brought a URLProtocol-based version to Swift years ago. What changed in 2026 is the groundwork: HAR is a de facto interchange format that every browser, Charles, Proxyman, and mitmproxy export, and Swift 6.1’s TestScoping protocol makes a per-test recording trait feel native instead of bolted on.
| Approach | Determinism | Fidelity | Maintenance |
|---|---|---|---|
| Live backend in CI | Low — third-party latency, outages, rate limits | Highest | None, until it breaks your pipeline |
| Hand-maintained stubs | High | Decays silently | Grows with every endpoint |
| Record-and-replay (HAR) | High | Captured from real traffic | Fixture review and refresh |
The Verdict
- Record-and-replay is the cheapest route to deterministic network tests. It removes the live service from the test’s dependency graph without asking you to maintain a parallel implementation of your networking layer.
- The fixture is a snapshot, not a source of truth. A green suite proves your client parses the recorded response. It says nothing about what the server sends today.
- Discipline is the whole product. Scrub credentials before recording, review HAR files before committing, version the fixtures, and refresh them whenever the API contract changes — or you are just testing yesterday’s server.
- Loose matching hides real regressions. Dropping to
.methodand.pathto survive volatile query parameters also stops the test from noticing a missing or renamed parameter.
The Fixture That Lies
The failure mode is boring, which is why it survives review. A stub is written by hand from the docs, then the API team adds a field or renames created_at to createdAt. Your decoding still succeeds because the fixture never had the field in the first place. The test is green; production is not.
The alternative most teams reach for — run the tests against the real endpoint — trades that for a different problem set: a p99 that occasionally exceeds the CI timeout, a sandbox that rate-limits parallel jobs, and a suite that turns red because someone else’s service is having a bad day.
// The hand-rolled stub we shipped for two years. It compiled, it was fast,
// and it was wrong: the fixture below predates the pagination contract,
// so the test never exercised the cursor path that broke in production.
final class UserServiceTests: XCTestCase {
func testFetchUser() async throws {
let stubData = #"{"id":42,"name":"Alice"}"#.data(using: .utf8)!
URLProtocol.registerClass(StubURLProtocol.self)
StubURLProtocol.responseBody = stubData
let user = try await UserService().fetchUser(id: 42)
XCTAssertEqual(user.name, "Alice")
}
}
The stub asserted what the author believed the server returned, on the day the author believed it. Nothing in the build forced a re-check.
What Replay Intercepts
Replay hooks the Foundation URL Loading System through URLProtocol, the same interception point DVR and OHHTTPStubs use. That means it works with URLSession.shared, custom sessions, and anything built on top of them, including Alamofire, with no protocol to define and no dependency to inject into production code.
The Swift Testing side is the part that did not exist before. A .replay trait names the HAR archive, and the Swift 6.1 TestScoping protocol handles the setup and teardown around each test:
import Foundation
import Testing
import Replay
struct User: Codable {
let id: Int
let name: String
let email: String
}
@Test(.replay("fetchUser"))
func fetchUser() async throws {
// Production code is untouched: this is the same URLSession it
// always used. The only difference is that the bytes come from
// Replays/fetchUser.har instead of the network stack.
let (data, _) = try await URLSession.shared.data(
from: URL(string: "https://api.example.com/users/42")!
)
let user = try JSONDecoder().decode(User.self, from: data)
#expect(user.id == 42)
}
A HAR file is plain JSON. It is diffable in review, editable by hand when a fixture needs a boundary case, and exportable from a browser’s Network tab, which means an on-call engineer can capture a real failing response and hand it to the test author.
{
"log": {
"version": "1.2",
"entries": [
{
"request": {
"method": "GET",
"url": "https://api.example.com/users/42",
"headers": [{ "name": "Accept", "value": "application/json" }]
},
"response": {
"status": 200,
"content": {
"mimeType": "application/json",
"text": "{\"id\":42,\"name\":\"Alice\"}"
}
}
}
]
}
}
The Deliberate First Failure
Replay refuses to record by accident. The first run of a test whose archive does not exist fails with a “No Matching Entry in Archive” diagnostic rather than silently hitting the network and writing whatever came back. That default is the single best design decision in the library: accidental recording is how credentials and session tokens end up in a committed fixture.
# Defaults are none/strict: never record, and fail if a fixture is missing.
$ swift test
❌ Test fetchUser() recorded an issue at ExampleTests.swift
⚠️ No Matching Entry in Archive
Request: GET https://api.example.com/users/42
# Record exactly once, against the real service, on purpose.
$ REPLAY_RECORD_MODE=once swift test --filter YourSuite.fetchUser
Two environment variables drive everything. REPLAY_RECORD_MODE is none, once, or rewrite. REPLAY_PLAYBACK_MODE is strict (require fixtures), passthrough (use fixtures when present, fall through otherwise), or live (ignore fixtures entirely). The live mode is what makes the pattern operable: the same test that runs deterministically in CI can be pointed at the real API for a contract check without rewriting a line.
One structural rule matters more than it looks: use one archive per test, and do not stack .replay traits on a single @Test. A HAR holds many request/response entries; a test that makes three calls records all three into one file. Stacking traits is ambiguous about which store serves which request.
Matching Is a Contract
The default matcher requires the HTTP method and the full URL — scheme, host, port, path, query, and fragment — to match exactly. That is strict in the useful way: a query parameter that disappears from your request is a test failure, not a silent fallback.
Real APIs make strict matching painful. Pagination cursors, cache-busters, and timestamps change on every run. Replay lets you weaken the matcher deliberately:
@Test(
.replay(
"listTransactions",
// We chose .path over the default full URL on purpose, and wrote
// down why: `cursor` is a server-generated opaque token, so exact
// matching would re-record the fixture every single run. The cost
// is that a removed query parameter no longer fails this test —
// the contract job below is what catches that.
matching: [.method, .path]
)
)
func listTransactions() async throws { /* ... */ }
Available matchers are .method, .url, .host, .path, .query, .headers, .body, and .custom for arbitrary logic. They compose with AND semantics, so a matcher set is the precise statement of which parts of a request the test considers load-bearing. Choosing it is an editorial act, not a configuration detail.
There is a real trade-off here that the ergonomics can hide. Every matcher you drop buys stability with loss of detectability. A team that reaches for .method on everything ends up with a suite that cannot notice it stopped sending an auth header.
Parallel Tests and the Scope Question
By default, Replay registers a global URLProtocol with serialized access, so tests using .replay run one at a time even under Swift Testing’s parallel scheduler. That is a conservative default, and it is honest about the shared state involved.
If the suite is large enough that serialization matters, scope: .test isolates each test’s playback store and restores parallelism. The price is explicit: the test must use Replay.session instead of URLSession.shared, because the per-test store is routed through a custom header that only Replay.session emits.
@Suite(.playbackIsolated(replaysFrom: Bundle.module))
struct ParallelizableAPITests {
@Test(.replay("fetchUser", matching: [.method, .path], scope: .test))
func fetchUser() async throws {
// This client takes a session by injection precisely so the test
// can hand it the isolated store. If it hard-coded .shared, the
// requests would fall into global playback and cross-talk.
let client = ExampleAPIClient(session: Replay.session)
_ = try await client.fetchUser(id: 42)
}
}
The dependency-injection requirement is worth adopting regardless. A networking client that accepts a URLSession is testable under any strategy; one that reaches for .shared internally is testable only by interception. The pattern rewards the interface you should have had anyway.
One boundary the URLProtocol route cannot cross: AsyncHTTPClient uses SwiftNIO rather than Foundation’s loader, so Replay ships a separate HTTPClientProtocol path for it. The lesson generalizes — if your networking stack does not go through URLSession, record-and-replay needs a seam, and you should confirm that seam exists before planning the migration.
The Discipline the Library Cannot Enforce
Replay ships filters that strip sensitive data during recording, which is the correct place to do it:
@Test(
.replay(
"fetchUser",
filters: [
.headers(removing: ["Authorization", "Cookie"]),
.queryParameters(removing: ["token", "api_key"])
]
)
)
func fetchUser() async throws { /* ... */ }
Filters are opt-in. The library cannot know which of your headers carries a bearer token, and the first HAR you generate will faithfully archive whatever the server echoed back. The Okta breach of 2023 is the canonical reminder: attackers obtained HAR files containing live session tokens, and Cloudflare responded by shipping a HAR sanitizer. Treat a freshly recorded archive as untrusted until it has been read by a human, and treat “review before committing” as a hard step, not a courtesy.
Three further pieces of discipline decide whether the pattern pays off or quietly rots:
- Version the fixtures in the same commit as the test. A HAR is a source artifact, not a build output. It belongs in review next to the assertion it supports, so a contract change shows up as a diff someone can read.
- Refresh on contract change, and detect it automatically. Add a scheduled job that runs the suite with
REPLAY_PLAYBACK_MODE=liveagainst a sandbox. When it diverges from the committed fixtures, you have found drift before a customer does. The plugin’sswift package replay statussurfaces stale and orphaned archives so the fixture set does not silently outlive the tests. - Cap fixture size. HAR captures response bodies verbatim, and a single list endpoint can add megabytes to the repository. Trim bodies to the fields the test asserts on, or the fixtures become a storage problem that reviewers start skipping.
The hidden cost is not the library. It is the ownership. A record-and-replay suite without a named owner and a refresh trigger becomes the most confident-looking stale test in the codebase: every assertion passes, every fixture is fiction.
When Record-and-Replay Is the Wrong Tool
The pattern earns its keep on integration tests that exercise decoding, pagination, error mapping, and retry logic against a realistic contract. It is a poor fit in a few familiar places:
- Tests that must prove live behavior — auth negotiation, server-side feature flags, rate-limit responses — need the real service. Point those at
REPLAY_PLAYBACK_MODE=liveand accept that they are integration-gated. - Streaming and server-push transports that do not traverse the
URLSessionloader need their own seam; a HAR models request/response pairs, not a long-lived socket. - Fixtures for a single status code or edge case are often clearer as inline stubs. Replay supports them directly, and reaching for a HAR to assert a 500 is more ceremony than signal.
The verdict stands, with its condition attached: Replay removes an entire class of flakiness for the cost of a fixture-review habit. Adopt it where determinism is the point, wire the live contract check on day one, and name the owner — because the tool that makes tests fast is also the tool that makes them confidently wrong when nobody is watching.
Internal Links
- Swift Testing in Production: Incremental Migration with XCTest Interoperability — the interop groundwork that lets a
.replaytest live beside existingXCTestCasesuites during migration. - Custom Traits in Swift Testing: Decorating Test Execution — the trait and
TestScopingmechanics behind.replay, and how to build your own scoped setup. - withTaskCancellationShield: Structured Cleanup Under Cancellation — the cancellation semantics a network-backed test must respect when a
strictplayback miss should fail fast rather than hang. - iOS 27.1 for Developers: SDK, Concurrency, and Liquid Glass Fixes — the checkpoint discipline of diffing a known-good baseline, applied here to fixtures instead of diagnostics.
External Links
- Apple Developer Documentation — Testing — the Swift Testing framework that hosts
@Test,#expect, and the trait system Replay builds on. - Apple Developer Documentation — Testing Traits — how traits like
.replaycompose and how a trait declares the scope around a test. - Apple Developer Documentation — URLProtocol — the URL Loading System extension point Replay intercepts, and the contract a custom protocol must honor.
- Swift.org — Documentation — canonical reference for the Swift 6 concurrency model these async tests compile against.
- mattt/Replay on GitHub — the library, the SPM command plugin (
status,record,inspect,validate,filter), and the matching and filter API surface. - NSHipster — Replay — the write-up that frames Replay against fifteen years of VCR-style recording and the Okta HAR exposure.