`withTaskCancellationShield`: Structured Cleanup Under Cancellation
withTaskCancellationShield: Structured Cleanup Under Cancellation
Between any two awaits on a cancelled task, Task.checkCancellation() can throw. Which suspension point observes it first is unspecified, so cleanup that must not be interrupted — committing a write, flushing a buffer, releasing a handle — has no structured home. withTaskCancellationShield (SE-0504, shipped in Swift 6.4 cancellation work) gives one: a scope in which the task’s cancelled state is temporarily unobservable. It is also the easiest new API in the release to misuse, and the misuse is silent.
The Verdict
- Use
withTaskCancellationShieldonly for short, bounded critical sections: committing a write, flushing a buffer, releasing a resource. The shield and the critical section should have the same duration. - A shield that spans a network call or an unbounded loop is a bug, not a convenience. You are re-enabling work the user explicitly cancelled, and every
checkCancellation()inside it now returns normally. - The shield does not cancel or un-cancel the task. It prevents the code inside from observing cancellation:
Task.isCancelledreadsfalse,Task.checkCancellation()does not throw, and cancellation handlers do not fire. When the scope exits, the real state is visible again. - It also cuts automatic propagation: child tasks created inside the shielded scope are not auto-cancelled with the parent. Explicit cancellation (a handle,
group.cancelAll()) still works. Task.hasActiveCancellationShieldis a diagnostic, not control flow. If a function needs it to behave correctly, the design is wrong.
Why Cleanup Has No Structured Home
Cancellation in Swift is cooperative and best-effort. Nothing stops a task; the runtime only flips a bit that your code is expected to consult. The consequence is a timing ambiguity:
func upload(_ chunk: Chunk) async throws {
try Task.checkCancellation()
try await client.append(chunk) // suspension point 1
// cancellation can land here, or here...
try await client.sync() // suspension point 2
// ...or here. The ordering is unspecified.
try await session.close() // close() checks Task.isCancelled internally
}
The interesting failure isn’t the early throw. It’s the late one. session.close() may internally do guard !Task.isCancelled else { return } — perfectly reasonable library code — and on a cancelled task it bails out having released nothing. The handle stays open, the buffer stays unsynced, the transaction stays unrolled. The task’s own cancellation is the exact moment the cleanup is supposed to run, and it is precisely the moment the cleanup can be skipped.
You can try to dodge it by checking cancellation yourself at every level. Two suspension points later, some third-party teardown you do not control declares itself done without doing the work. The cleanup problem is not that your checks are wrong; it is that nothing guarantees the continuation.
What a Shield Actually Does
SE-0504 adds a scope where the task behaves as if it were not cancelled:
defer {
// close() internally short-circuits on a cancelled task; without the
// shield this defer would "succeed" while leaking the session handle.
// We tried moving the guard earlier — the gap between suspension
// points is exactly what bit us in production.
await withTaskCancellationShield {
await session.close()
}
}
While the block runs, three observables change and exactly one. Task.isCancelled reads false; Task.checkCancellation() does not throw; a withTaskCancellationHandler installed inside the scope does not fire. The task itself is still flagged cancelled — withUnsafeCurrentTask and a held Task handle keep reporting the truth, which is deliberate. Instance methods are not scope-relative and must not flip-flop based on what the task happens to be executing right now.
The static methods are the shield’s contract, and the shield is their only boundary. There is no other hook, no timeout, no budget. The scope is the critical section, definitionally. Anything that outlives a bounded wall-clock time or an indefinite iteration inside it is a promise the API cannot keep.
The Two Implicit Commit Points
Two behaviors turn “temporarily cannot observe cancellation” into something developers read as “make this code finish”:
- Observability returns at scope exit. The task that was cancelled before the shield is immediately cancelled again the line after it. Nothing inside the scope can detect how late it is.
- Propagation into children stops. An
async letor task-group child created inside the shield does not receive the parent’s cancelled state automatically:
await withTaskCancellationShield {
// This child observes Task.isCancelled == false even though the
// parent was cancelled before we entered the scope. Intentional.
let result = await fetchProfile()
...
}
The common mistake is shielding the addTask call instead of the child’s body — cheap, correct-looking, and useless, because the child runs after the scope that shields it has exited. Shield inside the child, or not at all.
This is where a “convenience” can quietly become a resource leak in reverse: children silently start as uncancelled in a cancelled parent, and they are the only ones who know.
Where the Shield Belongs
The legitimate surface is narrow and you can enumerate it. Each of these is bounded by a finite, small number of operations, not by network or user time:
func flushAndClose(_ handle: FileHandle) {
withTaskCancellationShield {
// synchronize() may consult Task.isCancelled; an unsynced buffer
// on a cancelled write is the failure mode we are paying for here.
try? handle.synchronize()
try? handle.close()
}
}
let txn = try await store.begin()
do {
try await store.apply(updates, in: txn)
await withTaskCancellationShield {
// An interrupted commit is the expensive error — partial write,
// unknown state, journal recovery on next launch.
try? await store.commit(txn)
}
} catch {
await withTaskCancellationShield {
try? await store.rollback(txn)
}
throw error
}
That second example is the archetype. The business logic — applying updates — stays fully cancellable and throws normally. Only the commit, the bounded operation that must not be interrupted, is shielded. When you read the code, the shield marks exactly the section with an invariant on the other side.
The Hidden Cost: The Shield Is the Bug
The trade-off is not subtle and it is the reason the verdict is narrow. You are replacing one correctness problem with another, and the failure is loud in the first case and silent in the second only because you chose the second.
A shield around a network call
Wrap an upload in a shield and the user’s cancel is ignored until the request finishes or times out. If the cleanup itself needs the network — telling a peer you are disconnecting cleanly, for example — bound the cleanup from inside rather than drowning the whole call in a shield:
withTaskCancellationShield {
// The message we must send is itself network I/O. Scope a hard
// deadline so the shield cannot become a stuck task in disguise.
try? await withThrowingTaskGroup(of: Void.self) { group in
group.addTask { try await peer.send(.cleanDisconnect) }
group.addTask { try await Task.sleep(for: .seconds(5)) }
group.cancelAll()
}
}
A shield around an unbounded loop
Inside a shield, the loop’s own checkCancellation() never fires, so the loop has no exit other than its data. If the loop is bounded by work count, fine. If it is bounded by wall time or by input that stalled, it is now a task your user cannot stop. This is the one misuse that resembles a hang and is diagnosed as a bug in the app, not in the pattern.
The propagation costs
Shielded children are not auto-cancelled, so a work-fan-out inside a shield continues after the user left the screen — the opposite of structured concurrency’s default. And handlers do not fire: if the only thing that unblocks a waiter is a stored withTaskCancellationHandler, deferring it past the shield pushes the two-suspension-point ambiguity one scope further out, now invisible to the code that introduced it.
When investigating any of these, Task.hasActiveCancellationShield tells you whether a shield owns the current scope. Add it to the log line during the post-mortem, then remove it. It exists for diagnosis, and reaching for it in normal control flow means the shield is doing work it should not have been asked to do.
Operating Rule
Scope the shield to the critical section and the critical section to the shield. If they have different sizes, you have picked the wrong one. Cancellation should stop business logic promptly and still let the two lines of teardown behind it run — commit, flush, close. Everything past those two lines belongs outside the shield, where the user’s cancel can reach it.
Internal Links
- Swift 6.4 Concurrency: Async Defer, ~Sendable, and Borrow Accessors — Async
defer(SE-0493) is the companion to shields: it gives cleanup a guaranteed place to run, and shields give it a guaranteed cancelled-state context. - The Swift 6 Concurrency Manifesto: Isolation over Synchronization — The cooperative-cancellation and structured-concurrency model that SE-0504 builds on.
- Task-Local Logging and Diagnostic Context in Swift — Sibling “task-local context” surface; where
Task.hasActiveCancellationShieldfits when diagnosing cancellation behavior.
External Links
- Apple Developer Documentation — Task — The
Taskreference, including the “Shielding Tasks from Cancellation” group:withTaskCancellationShield(operation:)andhasActiveCancellationShield. - Apple Developer Documentation —
Task.checkCancellation()— The per-suspension-point check that a shield renders unobservable. - SE-0504 — Task Cancellation Shields — The accepted proposal on Swift.org’s Swift Evolution, including the child-task and handler semantics.
- Swift.org — Swift 6.4 released — Release notes for the toolchain in which the API ships.