Asyncio Gather: One Task Fails, Others Keep Running Anyway

Python · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Explains why asyncio.gather lets sibling tasks keep executing after one fails, and how Python 3.11's TaskGroup fixes the silent partial-failure trap.

What Changed

Python 3.11 added asyncio.TaskGroup, a structured-concurrency construct that changes what happens when one task in a group fails. If you've ever wondered why asyncio.gather lets other coroutines keep running for seconds after an exception was already logged and handled, TaskGroup is the fix: it cancels every sibling task the instant one of them raises, and it never silently drops exceptions.

The Old Way vs The New Way

Old way -- asyncio.gather, where a failure doesn't stop anything else:

async def old_way(urls):
    try:
        results = await asyncio.gather(*(fetch(u) for u in urls))
    except Exception as e:
        # by the time we get here, the other fetches
        # may still be running in the background
        log.error("one fetch failed: %s", e)
        results = None
    return results

New way -- asyncio.TaskGroup, where failure cancels the group:

async def new_way(urls):
    results = []
    try:
        async with asyncio.TaskGroup() as tg:
            tasks = [tg.create_task(fetch(u)) for u in urls]
    except* Exception as eg:
        for exc in eg.exceptions:
            log.error("fetch failed: %s", exc)
        return None
    return [t.result() for t in tasks]

Why It Was Added

Developers routinely assumed gather behaved like a transaction -- one failure, everything stops. It doesn't. Unless you pass return_exceptions=True and manually inspect every result, a failing task's siblings keep executing, keep holding connections open, and keep writing partial state, even though your except block already "handled" the error. This produced a whole class of bugs: half-written files, duplicate API calls, orphaned database rows, and tasks whose exceptions were never even observed (Python would just log a warning at garbage-collection time, long after it mattered). TaskGroup was added to give Python real structured concurrency: a parent scope that guarantees no child outlives a sibling's failure, borrowed directly from patterns proven in Kotlin coroutines and Trio's nurseries (Trio pioneered this in the Python ecosystem years earlier).

How It Works Underneath

A TaskGroup is an async context manager that tracks every task created via tg.create_task(). When any task raises (other than CancelledError), the group immediately cancels all remaining unfinished tasks, waits for them to actually unwind, then re-raises every collected exception together as an ExceptionGroup (or BaseExceptionGroup). You catch specific exception types inside that group with the except* syntax introduced alongside it.

sequenceDiagram
    participant TG as TaskGroup
    participant A as Task A (fetch)
    participant B as Task B (fetch)
    participant C as Task C (fetch)
    TG->>A: create_task()
    TG->>B: create_task()
    TG->>C: create_task()
    A-->>TG: raises ValueError
    TG->>B: cancel()
    TG->>C: cancel()
    B-->>TG: CancelledError (suppressed)
    C-->>TG: CancelledError (suppressed)
    TG-->>Caller: raise ExceptionGroup([ValueError])

The key mechanical difference from gather: cancellation is pushed into siblings proactively by the group itself, rather than left as an exercise for the caller.

Should You Adopt It Yet

Yes, if you're on Python 3.11+: TaskGroup is stable, well-tested, and the mechanism CPython itself now uses internally for related APIs like asyncio.timeout(). The cost is mostly syntactic -- except* is new syntax that some linters and older tooling (pre-3.11 type checkers, some IDEs) don't fully understand yet, and mixing except* with regular except in the same block isn't allowed. If you support Python 3.9 or 3.10 you can't use it directly; a backport (taskgroup on PyPI) exists but pulls in exceptiongroup as a dependency too. Anywhere you currently call gather() and treat a single failure as fatal for the whole batch, this is a strict improvement.

Migration Notes

Grep your codebase for asyncio.gather( and check each call site: if any exception should abort the whole batch, it's a TaskGroup candidate. If you're intentionally using return_exceptions=True to collect all results including failures (e.g., a health-check fan-out), leave those as gather -- TaskGroup's cancel-on-first-failure behavior is the wrong tool there. When migrating, watch for two gotchas: (1) except* requires the entire enclosing function to not mix it with plain except in the same try block, and (2) results are no longer returned by the async with block -- you must keep references to each Task object returned by create_task() and call .result() on them afterward.

Key takeaway: Use asyncio.TaskGroup instead of gather whenever a failure in one task should cancel its siblings, not just get reported alongside them.

Real-world challenge

Your batch job launches five asyncio.gather'd coroutines that each write results to a shared file. One coroutine throws partway through due to a malformed record. Logs show the exception was raised and handled, but the output file still contains partial writes from two other coroutines that kept running for another 30 seconds after the exception was logged, corrupting the batch. How do you diagnose and fix this?

Diagnosis: asyncio.gather() raises the first exception as soon as it happens but does not cancel the remaining tasks -- they keep running until they finish naturally. That's why writes continued after the error was already logged and handled by your try/except.

Fix: Replace gather with asyncio.TaskGroup (Python 3.11+), which cancels all sibling tasks the moment one raises, and aggregates failures into an ExceptionGroup:

async def run_batch(records):
    try:
        async with asyncio.TaskGroup() as tg:
            for r in records:
                tg.create_task(write_record(r))
    except* ValueError as eg:
        for exc in eg.exceptions:
            log.error("bad record: %s", exc)

If you're stuck pre-3.11, manually track tasks and cancel them in an except block around gather, or pass return_exceptions=True and check results yourself before letting any writes commit.