A customer was waiting on a billing correction for 22 hours. Our internal SLA promised four. Nobody on the team knew the ticket existed until the customer copied their account manager on a follow-up email, and the account manager forwarded it to me directly with a single-word subject line: "Felix?"
That was when I discovered we had been quietly missing SLAs for weeks, not constantly, but consistently, on a specific class of ticket that our process had a gap around. The customers who sent that second email showed up in the escalation queue. The customers who did not, did not show up anywhere at all.
Here is the full autopsy: what we thought was protecting us, what actually broke, and the three concrete changes I made afterward.
What We Thought Was Protecting Us
Before this happened, I genuinely believed we had a working SLA system. We had three things in place:
A timer on every open ticket showing how long it had been since the last reply. At a glance, the queue looked handled.
A weekly report showing average first response time. Ours was consistently well inside our four-hour target. It looked healthy.
A team that genuinely cared about the work. Nobody was ignoring tickets on purpose.
All three of these things were real. None of them were actually protecting us. Here is why.
The Three Ways the System Was Lying to Us
The timer was measuring the wrong event. In our helpdesk at the time, the clock on each ticket showed time since the agent's last reply, not time since the customer's last message. So when an agent had asked a clarifying question six hours ago and the customer had replied two hours ago, the ticket showed as "responded." The customer's message was sitting unread, and the timer was counting backward from the wrong moment. We were watching the wrong clock.
Average response time hides the slow tickets. A metric that averages across all tickets rewards fast first replies but tells you nothing about what happens in the middle of a conversation. If your team replies fast to new tickets but slow on follow-up messages, your average looks fine while your actual SLA performance is degrading. The billing ticket that triggered this autopsy was not a new ticket. It was a reply the customer had sent to a thread an agent had opened 12 hours earlier. First response time: excellent. Actual time the customer waited for resolution: 22 hours.
We had no breach alert. There was no mechanism to flag a ticket as it approached the SLA window. If you miss the moment, you only find out when the customer tells you. And many customers, especially in B2B relationships where they value the account and do not want to be difficult, will not tell you for a long time. They will just get quietly frustrated, quietly trust you less, and quietly evaluate alternatives.
The customers who escalated were not the ones most at risk. They were just the ones willing to push back. The silent ones were more dangerous.
The Three Fixes
I want to be specific here, because this is where most postmortems get vague. "We improved our processes" is not useful. Here is exactly what changed.
Fix 1: Change what the timer measures. The helpdesk we use now shows time since the customer's last message, not the agent's. If a customer has replied and no agent has responded since, the ticket shows as waiting for us regardless of what any agent wrote before the customer's message. This sounds like a small config change. It changed how the entire queue felt to work in. Suddenly tickets that "looked handled" started looking like what they were: tickets where a customer was waiting.
If your current helpdesk does not support this view natively, you can usually approximate it with a custom view filtered to tickets where the most recent reply is from the customer. Both Zendesk and Help Scout support this. Check your view configuration before assuming you cannot do it.
Fix 2: Replace average response time with breach rate. I no longer care about average first response time as a headline metric. I care about what percentage of tickets breached the SLA window at any point in the conversation, not just on first reply. This is a harder metric to feel good about, which is exactly why it is more honest. A team that is fast on new tickets but slow on follow-ups will look great on average response time and bad on breach rate. The second number is the real one.
In Zendesk this is built into SLA reporting if you configure the policies correctly. In Help Scout it requires a bit more manual work, but you can build a saved report filtered by response time on each message, not just the first one. Build the report before you need it, not after the customer emails the account manager.
Fix 3: Set up a breach warning before it happens. We now have a trigger that fires when a ticket is approaching the SLA window, posting a message to a dedicated Slack channel with a direct link to the ticket. Not after the breach. Before. The person monitoring that channel does not need to go looking for at-risk tickets. The at-risk tickets come to them.
This is not a sophisticated automation. In Zendesk you can configure this with a time-based trigger and a webhook. In Help Scout you can build a similar alert with their workflow engine plus a notification integration. The important thing is that the alert goes somewhere that a person actually monitors in real time. An alert that fires into the helpdesk queue is only useful if someone is watching that queue. Route it to Slack, Teams, or wherever your team's eyes actually are during the day.
What the Autopsy Actually Revealed
When I went back and pulled the data after fixing the monitoring, what I found was not that agents were dropping tickets. It was that a specific type of ticket, follow-up messages on threads that agents believed were "in progress," was falling through a gap in how the queue was prioritized. The queue was sorted by ticket creation date, not by most recent customer activity. So a new ticket from today sat at the top. A three-day-old thread where the customer had replied this morning sat near the bottom, labeled as an old ticket, effectively invisible unless you scrolled to it.
That is a queue design problem, not a people problem. And queue design problems do not show up when you are only measuring averages. They show up when you look at the tails: the oldest unanswered customer replies, the tickets closest to breach, the ones that would be embarrassing if an account manager saw them.
Now I sort by customer's last activity, not by ticket age. The oldest unread customer message is always at the top of the queue. The result was not a dramatic overnight improvement. It was a quieter one: a steady drop in breach rate over the following weeks, and a noticeable drop in the kind of surprised, apologetic conversations where you are explaining to a customer why they had to follow up.
What I Would Tell Someone Setting Up SLA Monitoring for the First Time
Three things, in the order that matters:
Get the view right first. Before you set up any alerts or reports, make sure you can see which tickets have unread customer replies. If you cannot answer "how many tickets in my queue are waiting on me right now?" with a single look, fix the view before anything else.
Measure breach rate, not average. Set up a report that shows how often you exceed your SLA window, by conversation turn, not just by first response. Look at it weekly. It will be uncomfortable at first. That discomfort is information.
Alert before the breach, not after. A notification that fires when you miss the SLA is a record of failure. A notification that fires 30 minutes before is a chance to catch it. Build the warning, not the report.
The customer who sent us that billing ticket stayed. We fixed the issue, apologized directly, and did not try to explain our process to them. They did not need to know about our queue problem. They needed the correction done and a clear acknowledgment that we had made them wait too long. What the incident did change was how we run the desk every day since.
The silent customers, the ones who did not send that second email, are the ones I think about when I look at the monitoring setup now. The whole point of the alerts and the queue view is not to avoid escalations. It is to serve the people who will never escalate but are still forming an opinion about whether to renew.
If you want to talk through your current SLA setup and where the gaps might be, drop me a message. I am happy to look at what you have and point out where the blind spots usually hide.
Comments