`Span` in Hot Paths: Retiring `UnsafeBufferPointer`

`Span` in Hot Paths: Retiring `UnsafeBufferPointer`

Span in Hot Paths: Retiring UnsafeBufferPointer

Run the same parser under Debug and Release and watch where the failure moves: once the debug preconditions in an UnsafeBufferPointer client compile out, an “index out of range” that crashed on CI becomes a silent out-of-bounds read that only surfaces in production telemetry. Swift Span closes exactly that gap. It is the bounds-checked, non-owning, non-escaping view over contiguous memory — accepted as SE-0447, ship-ready since Swift 6.2 — and it is the reason hot-path parsers and buffer transforms should stop importing pointers just to read a byte. This is the safe buffer view Swift always wanted, now that borrow-checking makes it possible.

The Verdict

  • Migrate pure-Swift parsers and buffer transforms first. Tokenizers, frame walkers, and in-place byte passes are the workloads where Span and MutableSpan remove the unsafe pointer wholesale while keeping bounds checks on in every build.
  • The win is memory-safety by construction, not raw throughput. A Span is a pointer-and-count value with the same performance class as UnsafeBufferPointer; the gains come from removing redundant work (double bounds-checking, slice copies, closure-inlining walls), not from the type itself being faster.
  • Keep unsafe pointers only where you must form a pointer across an FFI boundary. The trampoline into a C or C++ symbol is the one place a raw (pointer, count) pair legitimately exists — and even there, C++20 std::span annotation gets you safe overloads.
  • Bounds checks are on in Release. The single labeled opt-out, span[unchecked: i], carries the same precondition as the checked subscript — you swear the index is valid — which is a different contract than UnsafeBufferPointer, where unchecked access is the default and the pointer is unmanaged.
  • Budget the deployment floor now. The vending properties (Array.span, String.utf8.span, Data.span) gate on Apple OS 26; the types themselves back-deploy, but the ergonomic surface does not.

Why UnsafeBufferPointer Is the Wrong Default in Parsers

The SE-0447 motivation documents what every team discovers after the third withUnsafeBytes nesting level: the UnsafeBufferPointer handed to a closure-style API is unsafe in ways that accumulate. The pointer itself is unmanaged, the subscript is bounds-checked only in debug builds of client code, and it can escape the closure’s duration. Even when the closure body is disciplined, every function drawn inside it is now expressed in terms of unsafe types. The “safety” is a review habit, not a language guarantee.

For byte-pushing code the escape risk is the worst of the three because it is invisible. Returning a view of Data or Array storage from withUnsafeBytes produces a pointer that outlives its allocation; the compiler accepts it, the runtime is silent, and the crash lands on whichever thread happens to touch the memory later. That is not an edge case in a parser — it is the default outcome the moment someone refactors “compute header, then decode body” into two functions.

// v1 — every parser entry point started the same way: withUnsafeBytes to
// get a raw view, then a hand-rolled slice built from baseAddress + count.
// Two scalars. They can drift. The closure fences every read, so any
// helper needs the raw buffer passed down as an explicit parameter.
private func lex(_ data: Data) -> [Token] {
    data.withUnsafeBytes { raw -> [Token] in
        var tokens: [Token] = []
        var cursor = 0
        while cursor < raw.count {
            guard raw[cursor] == 0x24, cursor + 8 <= raw.count else {
                cursor += 1
                continue
            }
            // Manual LE32 because there is no try on the pointer family;
            // masking math is also the way we dodged alignment preconditions.
            let length = UInt32(raw[cursor + 4])
                | (UInt32(raw[cursor + 5]) << 8)
                | (UInt32(raw[cursor + 6]) << 16)
                | (UInt32(raw[cursor + 7]) << 24)
            guard cursor + 8 + Int(length) <= raw.count else { break }

            // Copy-out: the caller needs an owning payload, but we cannot
            // return a view of `raw`. Repeat for every sub-read.
            let payload = raw.baseAddress!
                .advanced(by: cursor + 8)
                .assumingMemoryBound(to: UInt8.self)
            tokens.append(
                .blob(Array(UnsafeBufferPointer(start: payload, count: Int(length))))
            )
            cursor += 8 + Int(length)
        }
        return tokens
    }
}

The shape of the problem: the view exists only inside a closure, so all reads are fenced by nesting or by threading UnsafeRawBufferPointer through signatures; and anything you want to hand to the next layer is copied out. Both costs are structural, not accidental.

Reading Storage Without the Escape Clause

Span changes the contract at the call site: instead of “here is a pointer, don’t store it,” the type itself cannot escape. A Span value (and a RawSpan for untyped reads) borrows from the container that vends it, so outliving the storage is a compile error, not a runtime bet. The vending properties — Array.span, String.utf8.span, Substring.utf8.span, ContiguousArray.span, Data.span — replace the closure-taking accessors in the common path.

// v2 — precondition: Apple OS 26 deployment floor, Swift 6.2+ toolchain.
// The view is now a value we can thread through helpers, and a sub-range
// is a first-class Span instead of two scalar temporaries.
private func lex(_ data: Data) -> [Token] {
    let bytes = data.span
    var tokens: [Token] = []
    var cursor = 0
    while cursor < bytes.count {
        guard bytes[cursor] == 0x24, cursor + 8 <= bytes.count else {
            cursor += 1
            continue
        }

        // extracting(Range) validates the window and rebases indices to 0,
        // which removed a whole class of "off by cursor" mistakes from v1.
        let head = bytes.extracting(cursor..<(cursor + 8))
        let length = head.withUnsafeBytes {
            $0.loadUnaligned(as: UInt32.self)  // little-endian on Apple
        }

        // If length is junk, extracting disallows the window and the guard
        // below trips; no pointer arithmetic was ever formed for it.
        let payload = bytes.extracting(cursor + 8..<(cursor + 8 + Int(length)))
        guard payload.count == Int(length) else { break }

        let payloadArray: [UInt8] = payload.withUnsafeBufferPointer { Array($0) }
        tokens.append(.blob(payloadArray))
        cursor += 8 + Int(length)
    }
    return tokens
}

Two details make this more than a syntax swap. First, the extraction is checked: an invalid Range is trapped the moment it is formed, before a single element is touched. Second, the only remaining pointer formation is the deliberate copy-out into an owning Array — the exact place a pointer should exist. Reading never needs one.

When the transform is mutating rather than reading, MutableSpan (SE-0467) carries the same story into in-place passes. For an inout [Float] buffer with unique ownership, the lifetime machinery also removes the per-write COW uniqueness probes that dominate the allocator profile:

// v1 — Instruments showed _ArrayBuffer.beginCOWMutation at the top of the
// call tree for this loop. The array is provably unique, but Swift re-runs
// the exclusivity probe on every write because sets of writes are harder
// to prove than one big borrow.
private func applyGain(_ samples: inout [Float]) {
    for i in samples.indices {
        samples[i] = samples[i] * 1.4 + 0.02
    }
}

// v2 — one exclusive borrow for the whole pass. The compiler also rejects
// a sibling Span read of the same array while this borrow is live, which
// the v1 loop could not express at all.
private func applyGain(_ samples: inout [Float]) {
    var view = samples.mutableSpan
    for i in view.indices {
        view[i] = view[i] * 1.4 + 0.02
    }
}

The Measured Picture

Community benchmark runs on recent hardware put the perf expectations in order. Span-based searches track Array-based ones within a couple of percent at every input size, from 1,000 to 10,000,000 elements — and allocate zero heap blocks at every percentile. The sharper number is split: a span loop that visits pieces through a non-escaping closure performs the same split with 0 mallocs where Array.split(separator:) materializes 7, 14, and 20 result slices at 1K, 100K, and 10M elements. The gap is not “Span is optimized better”; it is that the span version never builds the [ArraySlice] at all.

Published parser rewrites tell the same story. A TOML decoder’s ~800% wall-clock improvement came from making the algorithm lazier and eliminating redundant bound checks when accessing substrings — the same two removals Span makes idiomatic — and a matrix-multiplication rewrite traced its largest single win to replacing Array writes that re-ran COW uniqueness checks with a MutableSpan over the same storage.

Two honest caveats:

  • A hand-tuned UnsafeBufferPointer loop that already checked every access will not get faster. It gains checks in Release and statically enforced lifetimes; throughput is flat. If a profile shows the bounds check itself on top, that is the case for the labeled unchecked accessors — losing the preconditioned skip is a bug, not a speedup.
  • Per-access checks are elidable, not stupid. Span’s subscript carries fixed_storage.check_index semantics so the optimizer drops checks it can prove redundant; the escape hatch is small, explicit, and conditional.

The Hidden Cost

Deployment floor. Array.span and the other vending properties are not back-deployable to Apple OS releases older than 26 — the bridge to contiguous storage changed underneath them. The Span types themselves ride the CompatibilitySpan module for older targets, but hard-requiring the ergonomic surface effectively raises the deployment floor by one release. For a dependency that must support iOS 24/25, this is the blocker, not the language version.

Non-escapability is a structural constraint. A Span cannot be stored in a class property, captured in an escaping closure, or returned across an async boundary. Code that wanted to hold a view now copies once into an owning container instead. The compile-time wall is the point, but it lands on existing code as a restructure: lazy caches of “the bytes for this frame” become Array copies, and concurrency boundaries stop viewing and start copying. Budget for that churn in any fluent-API or AsyncSequence surface you own.

No Collection conformance. Because Span is non-escapable, none of the stdlib’s generic algorithms come for free. map, filter, split, firstIndex(of:) over a span are re-implemented per project or pulled from third-party packages — and there is no [Span] to return, so “chunk it into a collection” patterns must become visitation closures. For teams that lean on the Collection battery, this is a genuine ergonomics regression, not a footnote.

Bridged storage materializes. For pure-Swift Array/ContiguousArray/String, .span is O(1). For storage that arrived over the bridge (an NSString-backed string, an NSArray slice), the first span access can allocate a contiguous buffer — which the container then reuses, giving amortized O(1). That is a deliberate improvement over withUnsafeBufferPointer, which on some bridged instances allocates a temporary and discards it per call — an accidental quadratic that the span properties were designed to remove. But the first-access spike still exists, so cache the access instead of re-reading .span in a loop.

The unchecked door exists. subscript(unchecked:), extracting(unchecked:), and swapAt(unchecked:) skip validation under the same parameter preconditions as their checked counterparts. These are the @unchecked Sendable of the span world: callers promise what the compiler would otherwise prove. The design intention is that audits and code review can grep for unchecked: — a property no one could grep out of a sea of UnsafeBufferPointer arithmetic.

std::span, and Where Pointers Genuinely Stay

The second half of the thesis is the C++ bridge. Where your binary format or ML pipeline hands buffers to C++20 code, annotated std::span parameters import into Swift as Span/MutableSpan overloads, so the safe surface extends past module boundaries:

// frame_reader.h — spell the intent out once, at the seam. Without these
// attributes the importer falls back to a raw-pointer pair and every
// Swift caller is back to manual (pointer, count) bookkeeping.
#include <span>

using ByteSpan = std::span<const uint8_t>;

// __noescape promises the call does not retain `data`; __lifetimebound
// links the result to the argument so the returned view cannot dangle.
ByteSpan read_frame(ByteSpan data __noescape) __lifetimebound;
// Swift 6.4: the safe overload is synthesized directly from the std::span
// signature — the earlier "must add a header-side type alias" friction was
// removed in the 6.4 toolchain. The call shrinks to one line.
let frame = reader.read_frame(data.span)   // -> Span<UInt8>

For C pointer-plus-length signatures, __counted_by(len) maps the pair to a Span, and __sized_by to a RawSpan, provided the API is annotated __noescape — otherwise Swift offers a bounds-safe UnsafeBufferPointer overload that is not lifetime-safe. Headers that are not annotated at all are the residual case where you keep raw pointers: the 200-line C library you did not write, or the void* shuffle inside a shim. The rule of thumb is that the pointer lives only inside the trampoline function that makes the actual foreign call, and span.withUnsafeBufferPointer { ... } is the sanctioned hatch for the duration of that call.

Retiring Order

  1. Pure-Swift parsers and decoders reading Array, Data, or String.utf8 — highest safety return, no FFI involved. Convert the read paths first.
  2. Byte transforms and normalization passes via mutableSpan — removes the COW probe traffic and enforces read/write exclusivity in the same edit.
  3. Internal signatures currently typed UnsafeBufferPointer — flip to borrowing Span<Element>; callers that own storage satisfy it with .span, and the shakeout finds every escaped buffer.
  4. The real FFI surface last. Annotate the headers so the importer synthesizes safe overloads, then delete everything downstream of the trampoline that was only there to touch the pointer.

The residual unsafe code is then small enough to count: pointer formation at foreign-call sites, and the unchecked spans you audited and named as such. That is a retirement, not a migration.


Ready for more depth?

Master these concepts with our structured technical roadmap.

View Roadmap