Flashduty AI SRE is an AI agent for root cause analysis and everyday SRE work. Sign up for a free trial. This is the first post in Inside Flashduty AI SRE, a series about the designs, mistakes, and tradeoffs behind our agent engineering. Getting it ready for production has given us plenty to talk about.
We made our AI agent distributed: it runs around the clock, serves everyone in an organization, and spreads tasks across servers. Compared with running Codex or Cursor on your own machine, that brings a whole set of extra problems.
This first post covers message handling. Within each agent session, messages must be processed strictly in order, server restarts must not lose messages or unfinished work, and idle sessions must not consume resources. After quite a bit of trial and error, we landed on a design that comes down to this:
Give each session a dedicated "butler," and use a Redis lock that expires after 30 seconds to make sure only one butler in the entire system serves that session at a time.
Everything else follows from that: how messages queue up, what happens when a server dies, and why we didn't use an off-the-shelf message queue.
The problems we had to solve
A user asks the agent, "Help me figure out what happened with last night's alerts." The agent may spend several minutes checking metrics, reading logs, running commands, and calling tools.
That is very different from an ordinary web request. A normal endpoint responds in tens of milliseconds. If something goes wrong, the frontend can retry. An agent's investigation makes a dozen or more model calls, spending real money on tokens; its commands may change a real environment; and the user is watching progress arrive line by line. A failure halfway through costs much more. And during those few minutes, something will go wrong. Thanks, Murphy.
- Users send follow-up messages. "Check the database's slow queries while you're at it." That message cannot interrupt the investigation already running. The model is waiting for the last tool's result; inserting a new message would break the tool-call sequence and leave the turn's context unusable. Nor can we drop it: the user would see their message get no response. It has to queue up and wait for the current turn to finish.
- Servers restart. Every release rolls instances out of service. Imagine an investigation has run for four minutes and needs one more to reach a conclusion when its process receives a shutdown signal. If no other process takes over, the user gets a loading spinner that never stops. The tokens spent during those four minutes? Wasted.
- We run many servers, across multiple pods. A user's next message can land on any instance through the load balancer, but only one process may work on a given session at a time. Without coordination, two servers could claim the same session and run the same investigation twice, doubling the bill. They could create the same ticket twice or execute the same restart command twice. Their output would interleave in one conversation history, leaving the agent treating things it never said as its own context.
Three design decisions
One lock, one butler
Each session gets a "butler." Internally we call it an actor; it's just a piece of code dedicated to serving that session. Before starting work, it tries to acquire a Redis lock:
- Create the lock only if it does not exist, with a 30-second expiry: Redis
SET NX EX. The process that acquires it becomes the sole owner. Those that fail to acquire it exit immediately. - Renew the lock every 10 seconds while working.
- After the session has been idle for five minutes, release the lock and exit.
What if the owning process crashes? The lock expires within 30 seconds, another process acquires it, and takes over the session. No manual intervention and no leader election.
Producers deliver messages and leave
Every entry point that sends something to a session, whether it's a web message, a sub-agent reporting back, or an automation firing, does just two things: append the message to the tail of that session's Redis queue, then publish a Redis Pub/Sub notification. After that, it returns. It never waits for processing.
The butler runs a loop: take messages from the queue one at a time and process them; remove each after successful processing; when the queue is empty, wait for the next notification.
Separating delivery from processing gives us something useful: the producers' number, speed, and location are all independent of the butler. Any server or process can enqueue a message. The Redis queue determines the order, and only the butler processes it.
Idle butlers must not tie up resources
This is easy to overlook, and it has the biggest effect on cost.
On an agent platform, most sessions are idle most of the time. A user asks a question and leaves; the next message might arrive half an hour later. If every idle butler holds a Redis connection open while waiting on a blocking read, the connection count grows linearly with idle sessions. Ten thousand open sessions mean ten thousand connections. Establishing and maintaining those connections is a cost of its own.
Instead, once the butler drains the queue, it holds no Redis connection. It waits in its process's memory until a notification from the same server or another server wakes it. In case a notification gets lost, there's a fallback: every five seconds, it checks the queue, processes anything waiting, and otherwise keeps waiting. A five-second delay is unnoticeable to a user returning minutes later. Redis connections no longer grow with the number of idle sessions; they depend only on how many sessions are processing messages right now.

Several entry points deliver messages; one butler processes them in order. The lower sketch shows that same butler when idle. The ledger retains events; broadcasts provide immediate notifications.
When the butler's process crashes mid-message
The flow above has a gap. The butler takes a message from the queue, starts processing it, and then its process crashes. The message has already left the queue. Where did it go?
We solve this by keeping records in two steps. Taking a message doesn't simply remove it: the butler moves it from the pending queue to another Redis list, the processing list (Redis LMOVE). Only successful processing removes it from the processing list.
The first thing a new butler does is check the processing list. Did the previous owner leave anything unfinished? If so, move it back to the head of the pending queue (RPOPLPUSH), preserving its order. Messages aren't lost or reordered. The cost is that the same message may be processed twice, so we design each downstream processing step to be safe to repeat.
There is one more case: the previous butler didn't crash, it just stalled (say, one step took far too long to return) and missed its renewals, so the lock expired and a new butler took over. When the old one resumes and finishes its message, clearing the processing list as usual would delete the message the new butler is working on. So each butler receives an owner token when it takes over, and checks it before clearing the processing list. If the token is no longer its own, it leaves the list alone. An expired lock only shows that the holder hasn't renewed for 30 seconds, not that it has exited.
This pattern is old enough to appear in every message queue tutorial. But it deserves a mention in an agent system. A single message may represent an investigation that has already run for several minutes and spent a lot of tokens. Losing it doesn't mean "one less email sent." It means a user watching their investigation stop, with the token budget already spent.
What we deliberately left out
No Redis Streams, and no Kafka. We evaluated pretty much every conventional distributed queue option and chose the simplest combination: Redis lists and Pub/Sub. The reason is straightforward: MySQL holds the complete record of every session. Every message, tool call, and model reply is written to the database. From day one, we treated everything in Redis as a cache that could be rebuilt. Lost the queue? The messages are in the database. Lost a live notification? The client reconnects and catches up from its cursor. With the database covering recovery, there's no reason to add another streaming store to operate. That would store the same data twice and leave us reconciling the two copies for good.
No automatic session migration or load balancing. Which butler owns a session depends on who acquires the lock. There's no mechanism to move sessions from a busy server to an idle one. Most of an agent session's workload is waiting for model APIs, rather than using our process's CPU. We don't need to migrate sessions to balance that process load.
Leaving those two things out makes the system much simpler. The entire runtime depends on just Redis and MySQL. An engineer familiar with ordinary backend services can understand the whole mechanism in an afternoon.
Where this design fits, and where it doesn't
This design has a clear prerequisite: processing within one session is strictly sequential, and that session's throughput is capped at what one butler can handle. If a single session needs to process hundreds of messages per second, this model doesn't fit from day one.
Agent workloads tend to be the opposite: huge numbers of sessions, sparse messages within each session, and processing times measured in minutes. Under that load, butlers are idle most of the time, so the cost comes from the resources they hold while idle, not from how many butlers exist. Get idle resource use right, and two ordinary pieces of infrastructure, Redis and MySQL, can support a multi-tenant agent platform.
Ordering is only half of the session runtime problem, though. The other half is trickier: users can press "Stop" at any time, including before the current turn has even started. During a rollout, lock acquisition and stop instructions can race. Which arrives first? We had a real incident here: a user clicked Stop, and the agent still ran the whole turn. Internally, we called it "the stop that stopped nothing." The next post explains how we fixed it with a tombstone.


