Django Advanced

Async Django and ASGI in Production: Async Views, the Async ORM, and When It Actually Helps

Async Django done right: the ASGI model, async views and the async ORM, safely mixing sync and async with sync_to_async, serving under Uvicorn, and the decision of when concurrency actually pays off.

KC Ramo · DjangoZen Team Jul 11, 2026 19 min read 344 views

Async support in Django is real and production-ready, but it is not a free speedup you sprinkle on a slow view. Used well, it lets a single worker handle many concurrent I/O-bound requests — external API calls, slow databases, streaming responses — without a thread per request. Used carelessly, it makes code slower, deadlocks under load, or corrupts data. This tutorial covers the ASGI model, async views and the async ORM, safely mixing sync and async, and the decision of when async is worth it at all.

WSGI, ASGI, and what async actually buys you

The classic Django deployment runs under WSGI: a synchronous protocol where each request occupies one worker (a process or thread) for its entire lifetime. If that request spends 300ms waiting on an external API, the worker is idle but unavailable — you scale concurrency by adding workers, and each worker costs memory. ASGI is the asynchronous successor: a single event-loop worker can suspend a request that is waiting on I/O and serve others in the meantime. The win is concurrency under I/O wait, not raw CPU speed. Async does nothing for CPU-bound work — a view that computes for 200ms blocks the whole event loop for 200ms, starving every other request on that worker.

The practical rule: async pays off when your views spend most of their time waiting — calling third-party APIs, hitting slow upstream services, holding open long-lived connections (SSE, WebSockets). If your views are cheap and DB-bound with a fast local database, plain WSGI with a sensible worker count is simpler and usually just as fast.

Async views

An async view is just a coroutine. Django detects the async def and runs it on the event loop under an ASGI server; under WSGI it still works but Django wraps it in its own event loop per request, losing the benefit.

import asyncio

import httpx
from django.http import JsonResponse

async def dashboard(request):
    async with httpx.AsyncClient() as client:
        # These two upstream calls run truly concurrently
        prices, weather = await asyncio.gather(
            client.get("https://api.example.com/prices"),
            client.get("https://api.example.com/weather"),
        )
    return JsonResponse({"prices": prices.json(), "weather": weather.json()})

The asyncio.gather here is the whole point: two 200ms calls complete in ~200ms total instead of 400ms. Fan-out to independent I/O is where async shines, and it is impossible to express cleanly in synchronous code without threads.

A production-shaped aggregation view

The snippet above is right about the concurrency and wrong about almost everything else a real endpoint has to handle. It builds a new AsyncClient per request, so every request pays for fresh TCP and TLS handshakes and throws away the connection pool. It has no deadline, so one hung upstream holds the request open indefinitely. And if either call fails, asyncio.gather raises the first exception and the user gets a 500 even though the other half of the data was fine. Here is the same idea written the way it should run behind real traffic:

import asyncio

import httpx
from django.http import JsonResponse

SOURCES = {
    "prices": "https://api.example.com/prices",
    "rates": "https://api.example.com/rates",
    "news": "https://api.example.com/news",
}

_client: httpx.AsyncClient | None = None
_upstream_slots = asyncio.Semaphore(20)


def get_client() -> httpx.AsyncClient:
    # One pooled client per worker process, created lazily inside the loop
    global _client
    if _client is None:
        _client = httpx.AsyncClient(
            timeout=httpx.Timeout(3.0, connect=1.0),
            limits=httpx.Limits(max_connections=100, max_keepalive_connections=20),
        )
    return _client


async def fetch_json(url: str) -> dict:
    async with _upstream_slots:
        resp = await get_client().get(url)
        resp.raise_for_status()
        return resp.json()


async def portfolio(request):
    names = list(SOURCES)
    try:
        async with asyncio.timeout(2.5):
            outcomes = await asyncio.gather(
                *(fetch_json(SOURCES[name]) for name in names),
                return_exceptions=True,
            )
    except TimeoutError:
        return JsonResponse({"error": "upstream timeout"}, status=504)

    data, failed = {}, []
    for name, outcome in zip(names, outcomes):
        if isinstance(outcome, Exception):
            failed.append(name)
        else:
            data[name] = outcome
    return JsonResponse({"data": data, "failed": failed}, status=200 if data else 502)

Several decisions in that code are worth making consciously rather than copying:

  • The client is shared per process. Django's ASGI handler does not implement the ASGI lifespan protocol, so there is no framework-level startup hook to open the client and a shutdown hook to close it. Lazy creation on first use is the pragmatic answer; the process exit closes the sockets. The client and the semaphore both attach to the event loop that first uses them, which is fine in production (one loop per worker) but will fail with "bound to a different event loop" in test suites that create a fresh loop per test — reset _client in a fixture if that bites you.
  • The semaphore is a process-wide budget, not a per-request one. It caps how many upstream calls this worker has in flight across all concurrent requests, which is what protects the upstream's rate limit. A per-request semaphore only bounds a single fan-out.
  • There are two timeout layers. The httpx timeout bounds each individual connect and read; asyncio.timeout bounds the whole request's upstream phase. Without the outer one, three sequential retries of 3 seconds each can still add up to a 9-second response.
  • return_exceptions=True buys partial results. Each failure comes back as a value instead of propagating, so the view can decide which sources are optional. When the outer timeout fires, gather is cancelled and cancels its children, so nothing keeps running after the 504.

Whether to use asyncio.gather or asyncio.TaskGroup here is a semantic choice, not a style one. A task group is all-or-nothing: the first child exception cancels the siblings and is re-raised as an ExceptionGroup. That is exactly right when the response is useless without every piece (building an invoice from three services), and exactly wrong for a dashboard where two out of three panels is a perfectly good page.

The async ORM

Since Django 4.1 the ORM exposes async query methods with an a prefix, and querysets support async for. You must use these inside async views — calling a synchronous ORM method directly from a coroutine raises SynchronousOnlyOperation.

async def article_list(request):
    articles = []
    async for article in Article.objects.filter(published=True):
        articles.append(article.title)

    count = await Article.objects.filter(published=True).acount()
    latest = await Article.objects.afirst()
    obj, created = await Article.objects.aget_or_create(slug="intro")
    return JsonResponse({"titles": articles, "count": count})

The async methods (acount, aget, afirst, acreate, aget_or_create, aupdate, adelete) do not make the database itself asynchronous — under the hood Django still uses a synchronous database driver run in a thread pool. What you gain is that the event loop is not blocked while that query runs, so other requests proceed. Full async database drivers are on the roadmap but the thread-pool bridge is what ships today.

The parts of the ORM that stay synchronous

The a-prefixed methods cover query execution, but three everyday ORM behaviours have no async form, and each one surfaces as a SynchronousOnlyOperation the first time a code path touches it.

Lazy relation access. Reading article.author on an instance whose author was not loaded triggers a query from attribute access, and attribute access cannot be awaited. The fix is to load what you need up front, which is better practice anyway because it also removes N+1 queries:

async def article_detail(request, slug):
    article = await (
        Article.objects.select_related("author")
        .prefetch_related("tags")
        .aget(slug=slug)
    )
    return JsonResponse({
        "title": article.title,
        "author": article.author.name,           # already loaded, no query
        "tags": [t.name for t in article.tags.all()],  # served from prefetch cache
    })

Model instances have asave(), adelete() and arefresh_from_db(), and related managers have aadd(), aremove(), aset() and aclear(), so writes through relations are covered. It is the implicit reads that bite.

Transactions. transaction.atomic() does not work as an async context manager, and an atomic block cannot span await points, because each awaited ORM call may execute on a thread holding a different connection. Anything that needs a transaction — especially select_for_update() — should be written as an ordinary synchronous function and called across the boundary as a single unit:

from asgiref.sync import sync_to_async
from django.db import transaction
from django.db.models import F


@sync_to_async
def transfer(src_id: int, dst_id: int, amount: int) -> None:
    with transaction.atomic():
        src = Account.objects.select_for_update().get(pk=src_id)
        if src.balance < amount:
            raise InsufficientFunds(src_id)
        Account.objects.filter(pk=src_id).update(balance=F("balance") - amount)
        Account.objects.filter(pk=dst_id).update(balance=F("balance") + amount)


async def transfer_view(request):
    await transfer(request.user.account_id, int(request.POST["to"]), int(request.POST["amount"]))
    return JsonResponse({"ok": True})

This is not a workaround to feel bad about. Transactional business logic is naturally sequential, gains nothing from the event loop, and is easier to reason about as one synchronous block with a clear start and commit.

Templates. The template engine is synchronous. Calling render() from an async view is allowed, but if the context contains an unevaluated queryset, the {% for %} loop inside the template executes the query synchronously and raises. Evaluate querysets in the view first — articles = [a async for a in qs] — and pass plain lists to the template.

The same principle applies to authentication: in an async view use await request.auser() (Django 5.0+) instead of touching request.user, which is a lazy object that hits the session and user tables on first access.

Mixing sync and async safely

Real code is rarely all-async. You will call synchronous libraries from async views and occasionally async code from sync context. Django gives you two adapters: sync_to_async to call blocking code from a coroutine, and async_to_sync for the reverse.

from asgiref.sync import sync_to_async

async def profile(request):
    # A third-party sync SDK, or ORM code you can't easily convert
    result = await sync_to_async(legacy_sdk.fetch, thread_sensitive=True)(request.user.id)
    return JsonResponse(result)

The thread_sensitive=True default matters: it forces the wrapped code to run in a single shared thread so that libraries relying on thread-local state — including Django's database connections and transactions — behave correctly. Setting thread_sensitive=False runs the call in a fresh thread-pool thread, which is faster for genuinely independent work but dangerous for anything touching the ORM inside a transaction. When in doubt, keep it True.

Async middleware and the sync/async boundary

Middleware can be sync or async, and Django adapts between them automatically — but every adaptation has a cost. If you run async views behind a stack of synchronous middleware, Django wraps each boundary in a thread, eroding the benefit. Mark middleware async-capable so the request flows through the event loop end to end.

import time


class TimingMiddleware:
    async_capable = True
    sync_capable = False

    def __init__(self, get_response):
        self.get_response = get_response

    async def __call__(self, request):
        start = time.monotonic()
        response = await self.get_response(request)
        response["X-Elapsed"] = f"{time.monotonic() - start:.3f}"
        return response

Serving ASGI in production

You need an ASGI server. The common production setup is Gunicorn managing Uvicorn workers, which gives you Gunicorn's process management with Uvicorn's fast event loop.

# pip install uvicorn-worker  (the Gunicorn worker class now lives in its own package)
gunicorn myproject.asgi:application \
    -k uvicorn_worker.UvicornWorker \
    --workers 4 --bind 0.0.0.0:8000

Note that async concurrency and worker count are now two different dials. Each Uvicorn worker runs one event loop that handles many concurrent requests, so you need far fewer workers than the WSGI rule of "2×CPU + 1". But you still want several workers to use multiple cores and to survive a worker restart without dropping all traffic.

Pitfalls that bite in production

The most common failure is a hidden blocking call in an async view — a synchronous requests.get, a time.sleep, a CPU-heavy loop, or a sync library call not wrapped in sync_to_async. Any of these blocks the entire event loop, so one slow request stalls every other request on that worker. Under load this looks like sudden latency cliffs that are hard to trace. Audit async views for anything that does not await.

The second is database connections. Django opens a connection per thread; with async and thread pools you can create far more connections than you expect, exhausting Postgres. Put PgBouncer in front (or use Django 5.1+'s built-in psycopg connection pool) and keep CONN_MAX_AGE = 0 under ASGI, as Django's documentation recommends, because persistent connections are not reliably reused across async requests. The third is unbounded concurrency: asyncio.gather over a large list fires every call at once and can hammer an upstream into rate limits — bound it with an asyncio.Semaphore.

Streaming and long-lived connections

Async's other natural home is responses that stay open. Server-Sent Events, long polling, and streaming AI tokens all hold a connection while data trickles out, and under WSGI each one pins a whole worker for the duration. Under ASGI an async generator streams without blocking the loop, so one worker can hold thousands of open streams.

from django.http import StreamingHttpResponse

async def events(request):
    async def stream():
        async for item in watch_updates():
            yield f"data: {item}\n\n"
    return StreamingHttpResponse(stream(), content_type="text/event-stream")

Full bidirectional WebSockets go a step further and live in Django Channels, which is built on the same ASGI foundation. If you find yourself reaching for streaming and sockets, ASGI is no longer optional — it is the only model that serves them efficiently.

Async views are not a task queue

A frequent misuse is treating an async view as a place to run background work — kicking off a long job and hoping it finishes after the response is sent. It will not reliably: once the response is returned the request scope can be torn down, and any unawaited work is at the mercy of the event loop's lifecycle. Async concurrency is for work the request genuinely waits on and returns. Fire-and-forget work that outlives the request belongs in Celery or a similar queue, where it has its own worker, retries, and durability. Keep the two ideas separate and you avoid a whole class of "the job sometimes just doesn't run" bugs.

Databases, connections, and pooling

Async magnifies a problem WSGI mostly hid: connection count. Django opens one database connection per thread, and the async stack multiplies threads through its pools, so a handful of Uvicorn workers can open far more Postgres connections than you planned, hitting max_connections and refusing service. The fix is a connection pooler in front of the database — PgBouncer in transaction mode — so hundreds of app-side connections multiplex onto a small pool of real ones. Keep CONN_MAX_AGE = 0 under ASGI and load-test the connection count, because it fails suddenly and only under real concurrency.

Testing async code

Async views need async tests. Django's test client has an async variant, and pytest-asyncio lets you write coroutine tests directly. The important discipline is to test the concurrency you rely on — assert that two upstream calls actually overlap, and that a slow dependency does not serialize requests — because a blocking call sneaks back in easily and silently undoes the entire benefit.

import pytest

@pytest.mark.asyncio
async def test_dashboard_concurrent(async_client):
    resp = await async_client.get("/dashboard/")
    assert resp.status_code == 200

Treat "no accidental blocking call" as an invariant worth a test, not a thing you check once by hand.

Timeouts and cancellation

An async view that awaits an upstream service must bound the wait — without a timeout, one hung dependency holds a connection open until the client or server gives up, and enough of them exhaust the loop. Wrap external calls in asyncio.timeout and decide the fallback explicitly: a cached value, a partial response, or a clean error.

try:
    async with asyncio.timeout(2.0):
        data = await client.get(url)
except TimeoutError:
    data = cached_fallback()

Cancellation is the flip side. When a client disconnects, the request's task is cancelled and any child tasks you spawned should be too — leaking orphaned tasks is a slow memory leak that only shows under sustained traffic. Structure concurrent work with task groups so cancellation propagates cleanly rather than leaving work running with nowhere to return.

Observability across await points

Debugging async is harder because a single request hops threads and suspends at every await, so naive thread-local logging loses the request's identity. Use contextvars to carry a request id across await boundaries, and instrument with OpenTelemetry, which understands async context and stitches the spans of concurrent upstream calls into one coherent trace. Without that, a slow async request is a puzzle; with it, you can see exactly which awaited call cost the latency.

Diagnosing a blocked event loop

The signature of a blocking call in production is distinctive once you know it: latency rises for every endpoint served by a worker at the same moment, including trivial ones like a health check, while CPU on the machine stays modest and the database looks idle. Under WSGI a slow request only hurts itself; under ASGI it hurts its neighbours. When you see that pattern, work through it in this order.

Turn on asyncio debug mode in staging. Running with PYTHONASYNCIODEBUG=1 (or python -X dev) makes the loop log a warning whenever a single step of a task runs longer than loop.slow_callback_duration, which defaults to 100ms. The log line names the task and the duration, which usually points straight at the offending view. Debug mode adds overhead, so use it in staging and load tests rather than permanently in production.

Measure loop lag continuously in production. A cheap, always-on alternative is a background probe that schedules itself at a fixed interval and records how late it wakes up. On a healthy loop the lag stays in the low milliseconds; spikes line up precisely with blocking calls.

import asyncio
import logging
import time

logger = logging.getLogger("loop_lag")


async def monitor_loop_lag(interval: float = 0.5, threshold: float = 0.1) -> None:
    while True:
        start = time.monotonic()
        await asyncio.sleep(interval)
        lag = time.monotonic() - start - interval
        if lag > threshold:
            logger.warning("event loop lag %.3fs", lag)

Start it once per worker (for example, lazily from the first request, guarded by a module-level flag) and export the lag as a metric rather than just a log line. The p99 of loop lag is one of the most useful single numbers for an async service.

Inspect a stuck worker directly. If a worker is wedged right now, py-spy dump --pid <pid> prints the current Python stack of every thread without stopping the process. A main thread sitting inside ssl.read, socket.recv or time.sleep from an async view is your blocking call, caught in the act.

A few symptoms and their usual causes:

SymptomLikely causeFix
All endpoints on a worker slow down togetherBlocking call on the loopFind it with debug mode or loop-lag metric; await it or wrap it in sync_to_async
Workers killed by the arbiter with timeout messagesLoop blocked longer than --timeoutSame as above; do not just raise the timeout
SynchronousOnlyOperation in logsLazy relation, template iterating a queryset, or sync ORM call in async codeUse select_related/prefetch_related, evaluate querysets in the view, use a-methods
"too many clients already" from PostgreSQLConnections scale with concurrent requestsPgBouncer or Django's pool; bound concurrency
Memory grows slowly under steady loadOrphaned tasks or unbounded in-process queuesTask groups, timeouts, worker recycling with --max-requests
Upstream returns 429 during traffic peaksUnbounded fan-outProcess-wide semaphore on outbound calls

When to reach for async

Reach for async when a view is dominated by concurrent, independent I/O — aggregating several external APIs, proxying slow upstreams, streaming server-sent events, or holding many long-lived connections. Stay on WSGI when your workload is CPU-bound, when it is simple DB-bound CRUD against a fast local database, or when the team is not ready to reason about event loops and the sync/async boundary. Async is a powerful tool with a real complexity tax; adopt it where the concurrency win is concrete, not because it sounds modern.