Replay: Record-and-Replay Network Tests with Swift Testing

Replay: Record-and-Replay Network Tests with Swift Testing

⚠️ Speculative Architecture & Preview: This article discusses future system iterations (e.g., iOS 27, Xcode 27) as conceptual planning and architectural design patterns. Technical details represent previews and proposals rather than finalized APIs.

Replay: Record-and-Replay Network Tests with Swift Testing

How many JSON fixtures in your test bundle are lying to you right now? A hand-written stub compiles forever, survives every refactor, and silently stops matching the server the day someone renames a field. Replay Swift Testing replaces the guesswork with a different contract: record real HTTP traffic once into a HAR file, then replay it byte-for-byte inside Swift Testing.

It is not a new idea. Ruby’s VCR did this in 2010, Python’s VCR.py and pytest-recording followed, and Venmo’s DVR brought a URLProtocol-based version to Swift years ago. What changed in 2026 is the groundwork: HAR is a de facto interchange format that every browser, Charles, Proxyman, and mitmproxy export, and Swift 6.1’s TestScoping protocol makes a per-test recording trait feel native instead of bolted on.

ApproachDeterminismFidelityMaintenance
Live backend in CILow — third-party latency, outages, rate limitsHighestNone, until it breaks your pipeline
Hand-maintained stubsHighDecays silentlyGrows with every endpoint
Record-and-replay (HAR)HighCaptured from real trafficFixture review and refresh

The Verdict

  • Record-and-replay is the cheapest route to deterministic network tests. It removes the live service from the test’s dependency graph without asking you to maintain a parallel implementation of your networking layer.
  • The fixture is a snapshot, not a source of truth. A green suite proves your client parses the recorded response. It says nothing about what the server sends today.
  • Discipline is the whole product. Scrub credentials before recording, review HAR files before committing, version the fixtures, and refresh them whenever the API contract changes — or you are just testing yesterday’s server.
  • Loose matching hides real regressions. Dropping to .method and .path to survive volatile query parameters also stops the test from noticing a missing or renamed parameter.

The Fixture That Lies

The failure mode is boring, which is why it survives review. A stub is written by hand from the docs, then the API team adds a field or renames created_at to createdAt. Your decoding still succeeds because the fixture never had the field in the first place. The test is green; production is not.

The alternative most teams reach for — run the tests against the real endpoint — trades that for a different problem set: a p99 that occasionally exceeds the CI timeout, a sandbox that rate-limits parallel jobs, and a suite that turns red because someone else’s service is having a bad day.

// The hand-rolled stub we shipped for two years. It compiled, it was fast,
// and it was wrong: the fixture below predates the pagination contract,
// so the test never exercised the cursor path that broke in production.
final class UserServiceTests: XCTestCase {
    func testFetchUser() async throws {
        let stubData = #"{"id":42,"name":"Alice"}"#.data(using: .utf8)!

        URLProtocol.registerClass(StubURLProtocol.self)
        StubURLProtocol.responseBody = stubData

        let user = try await UserService().fetchUser(id: 42)
        XCTAssertEqual(user.name, "Alice")
    }
}

The stub asserted what the author believed the server returned, on the day the author believed it. Nothing in the build forced a re-check.

What Replay Intercepts

Replay hooks the Foundation URL Loading System through URLProtocol, the same interception point DVR and OHHTTPStubs use. That means it works with URLSession.shared, custom sessions, and anything built on top of them, including Alamofire, with no protocol to define and no dependency to inject into production code.

The Swift Testing side is the part that did not exist before. A .replay trait names the HAR archive, and the Swift 6.1 TestScoping protocol handles the setup and teardown around each test:

import Foundation
import Testing
import Replay

struct User: Codable {
    let id: Int
    let name: String
    let email: String
}

@Test(.replay("fetchUser"))
func fetchUser() async throws {
    // Production code is untouched: this is the same URLSession it
    // always used. The only difference is that the bytes come from
    // Replays/fetchUser.har instead of the network stack.
    let (data, _) = try await URLSession.shared.data(
        from: URL(string: "https://api.example.com/users/42")!
    )
    let user = try JSONDecoder().decode(User.self, from: data)
    #expect(user.id == 42)
}

A HAR file is plain JSON. It is diffable in review, editable by hand when a fixture needs a boundary case, and exportable from a browser’s Network tab, which means an on-call engineer can capture a real failing response and hand it to the test author.

{
  "log": {
    "version": "1.2",
    "entries": [
      {
        "request": {
          "method": "GET",
          "url": "https://api.example.com/users/42",
          "headers": [{ "name": "Accept", "value": "application/json" }]
        },
        "response": {
          "status": 200,
          "content": {
            "mimeType": "application/json",
            "text": "{\"id\":42,\"name\":\"Alice\"}"
          }
        }
      }
    ]
  }
}

The Deliberate First Failure

Replay refuses to record by accident. The first run of a test whose archive does not exist fails with a “No Matching Entry in Archive” diagnostic rather than silently hitting the network and writing whatever came back. That default is the single best design decision in the library: accidental recording is how credentials and session tokens end up in a committed fixture.

# Defaults are none/strict: never record, and fail if a fixture is missing.
$ swift test
❌  Test fetchUser() recorded an issue at ExampleTests.swift
⚠️  No Matching Entry in Archive
    Request: GET https://api.example.com/users/42

# Record exactly once, against the real service, on purpose.
$ REPLAY_RECORD_MODE=once swift test --filter YourSuite.fetchUser

Two environment variables drive everything. REPLAY_RECORD_MODE is none, once, or rewrite. REPLAY_PLAYBACK_MODE is strict (require fixtures), passthrough (use fixtures when present, fall through otherwise), or live (ignore fixtures entirely). The live mode is what makes the pattern operable: the same test that runs deterministically in CI can be pointed at the real API for a contract check without rewriting a line.

One structural rule matters more than it looks: use one archive per test, and do not stack .replay traits on a single @Test. A HAR holds many request/response entries; a test that makes three calls records all three into one file. Stacking traits is ambiguous about which store serves which request.

Matching Is a Contract

The default matcher requires the HTTP method and the full URL — scheme, host, port, path, query, and fragment — to match exactly. That is strict in the useful way: a query parameter that disappears from your request is a test failure, not a silent fallback.

Real APIs make strict matching painful. Pagination cursors, cache-busters, and timestamps change on every run. Replay lets you weaken the matcher deliberately:

@Test(
    .replay(
        "listTransactions",
        // We chose .path over the default full URL on purpose, and wrote
        // down why: `cursor` is a server-generated opaque token, so exact
        // matching would re-record the fixture every single run. The cost
        // is that a removed query parameter no longer fails this test —
        // the contract job below is what catches that.
        matching: [.method, .path]
    )
)
func listTransactions() async throws { /* ... */ }

Available matchers are .method, .url, .host, .path, .query, .headers, .body, and .custom for arbitrary logic. They compose with AND semantics, so a matcher set is the precise statement of which parts of a request the test considers load-bearing. Choosing it is an editorial act, not a configuration detail.

There is a real trade-off here that the ergonomics can hide. Every matcher you drop buys stability with loss of detectability. A team that reaches for .method on everything ends up with a suite that cannot notice it stopped sending an auth header.

Parallel Tests and the Scope Question

By default, Replay registers a global URLProtocol with serialized access, so tests using .replay run one at a time even under Swift Testing’s parallel scheduler. That is a conservative default, and it is honest about the shared state involved.

If the suite is large enough that serialization matters, scope: .test isolates each test’s playback store and restores parallelism. The price is explicit: the test must use Replay.session instead of URLSession.shared, because the per-test store is routed through a custom header that only Replay.session emits.

@Suite(.playbackIsolated(replaysFrom: Bundle.module))
struct ParallelizableAPITests {
    @Test(.replay("fetchUser", matching: [.method, .path], scope: .test))
    func fetchUser() async throws {
        // This client takes a session by injection precisely so the test
        // can hand it the isolated store. If it hard-coded .shared, the
        // requests would fall into global playback and cross-talk.
        let client = ExampleAPIClient(session: Replay.session)
        _ = try await client.fetchUser(id: 42)
    }
}

The dependency-injection requirement is worth adopting regardless. A networking client that accepts a URLSession is testable under any strategy; one that reaches for .shared internally is testable only by interception. The pattern rewards the interface you should have had anyway.

One boundary the URLProtocol route cannot cross: AsyncHTTPClient uses SwiftNIO rather than Foundation’s loader, so Replay ships a separate HTTPClientProtocol path for it. The lesson generalizes — if your networking stack does not go through URLSession, record-and-replay needs a seam, and you should confirm that seam exists before planning the migration.

The Discipline the Library Cannot Enforce

Replay ships filters that strip sensitive data during recording, which is the correct place to do it:

@Test(
    .replay(
        "fetchUser",
        filters: [
            .headers(removing: ["Authorization", "Cookie"]),
            .queryParameters(removing: ["token", "api_key"])
        ]
    )
)
func fetchUser() async throws { /* ... */ }

Filters are opt-in. The library cannot know which of your headers carries a bearer token, and the first HAR you generate will faithfully archive whatever the server echoed back. The Okta breach of 2023 is the canonical reminder: attackers obtained HAR files containing live session tokens, and Cloudflare responded by shipping a HAR sanitizer. Treat a freshly recorded archive as untrusted until it has been read by a human, and treat “review before committing” as a hard step, not a courtesy.

Three further pieces of discipline decide whether the pattern pays off or quietly rots:

  1. Version the fixtures in the same commit as the test. A HAR is a source artifact, not a build output. It belongs in review next to the assertion it supports, so a contract change shows up as a diff someone can read.
  2. Refresh on contract change, and detect it automatically. Add a scheduled job that runs the suite with REPLAY_PLAYBACK_MODE=live against a sandbox. When it diverges from the committed fixtures, you have found drift before a customer does. The plugin’s swift package replay status surfaces stale and orphaned archives so the fixture set does not silently outlive the tests.
  3. Cap fixture size. HAR captures response bodies verbatim, and a single list endpoint can add megabytes to the repository. Trim bodies to the fields the test asserts on, or the fixtures become a storage problem that reviewers start skipping.

The hidden cost is not the library. It is the ownership. A record-and-replay suite without a named owner and a refresh trigger becomes the most confident-looking stale test in the codebase: every assertion passes, every fixture is fiction.

When Record-and-Replay Is the Wrong Tool

The pattern earns its keep on integration tests that exercise decoding, pagination, error mapping, and retry logic against a realistic contract. It is a poor fit in a few familiar places:

  • Tests that must prove live behavior — auth negotiation, server-side feature flags, rate-limit responses — need the real service. Point those at REPLAY_PLAYBACK_MODE=live and accept that they are integration-gated.
  • Streaming and server-push transports that do not traverse the URLSession loader need their own seam; a HAR models request/response pairs, not a long-lived socket.
  • Fixtures for a single status code or edge case are often clearer as inline stubs. Replay supports them directly, and reaching for a HAR to assert a 500 is more ceremony than signal.

The verdict stands, with its condition attached: Replay removes an entire class of flakiness for the cost of a fixture-review habit. Adopt it where determinism is the point, wire the live contract check on day one, and name the owner — because the tool that makes tests fast is also the tool that makes them confidently wrong when nobody is watching.

Ready for more depth?

Master these concepts with our structured technical roadmap.

View Roadmap