Part II · Retrieval
Chapter 08
Scaling with Async Queues
Generation takes seconds to minutes; web request handling is designed for milliseconds — almost every scaling failure in an AI product comes from ignoring that mismatch one release too long.
Deliverable: An ARQ worker fleet behind FastAPI with SSE delivery, idempotency, dead letters and age-based scaling.
What's inside
11 topics
- 8.1Sync versus Async in RAG Architectures
- 8.2Queue System Design for Async Setup
- 8.3Setting Up Redis and Valkey with Docker
- 8.4Python ARQ and Distributed Queues
- 8.5Worker Orchestration
- 8.6FastAPI Endpoints for the Chat Queue
- 8.7Asynchronous Enqueueing and Streaming Back
- 8.8Polling, SSE and Result Delivery
- 8.9Classes of Service and Reserved Capacity
- 8.10Idempotency, Visibility Timeouts and Dead Letters
- 8.11Scaling Worker Nodes on the Right Signal
Preparing PDF viewer…
A note on this content
The book and its chapters are my personal learning notes — compiled from online research and hands-on practice, with most of the content AI-generated from that research and learning. It is not a peer-reviewed publication, and I make no claim that it is 100% error-free. If you spot a mistake, I'd genuinely appreciate hearing about it — contact me.