Insights

Surviving a Support Volume Spike Without Burning the Team

Spike handling is a queue-shaping problem before it is a staffing problem

A product launch, a pricing change, a shipping delay or an outage all produce the same shape of problem: conversation volume multiplies within hours, and the team responds by working harder. Working harder is the one lever that does not scale, and it is the lever most teams reach for first.

The reason is that spikes are usually treated as a staffing problem — how many people can we get online — when they are actually a decision-latency problem. During a spike the queue grows faster than any team can answer it, so the question shifts from how do we answer everything to what do we refuse to answer, and how do we tell the customer that. Teams that make that shift early come out fine. Teams that do not end up with a backlog they are still clearing a month later.

Spike handling is a queue-shaping problem before it is a staffing problem

The three levers, in the order you should pull them

Deflect. During a spike, a large fraction of inbound messages are the same two or three questions. Intercom’s guidance on supporting launches makes the point that pre-empting the question is worth more than answering it quickly, and that means publishing the answer where the customer already is — a status note, a pinned message, an updated help article linked from the product itself. Every question answered before it is asked is a question that never enters the queue.

The important detail is timing. Deflection content written during the spike arrives after the peak. It has to be drafted before launch, using the questions the last launch generated, which is why keeping a post-incident list of top questions is worth the ten minutes it takes.

Defer. Not everything in the queue is equally urgent, and during a spike the difference becomes enormous. A customer asking about a delivery date is not in the same category as a customer whose payment failed. The work is to write down, in advance, which intents may be answered with a holding message and which may not.

A holding message that says we are experiencing higher than normal volume, here is what we know, here is when we will update you is a legitimate response. A holding message that says nothing is a complaint generator. The difference between the two is whether it contains information the customer did not have before.

Divide. Only after the first two levers have been pulled does it make sense to add capacity, and by then the added capacity is doing something specific rather than absorbing an undifferentiated flood. Crisp’s writing on workload management frames the objective precisely: balancing the queue without burning the team out, which is a different target from clearing the queue.

Burning out the team is easy to do accidentally during a spike, and the cost is paid later: the people who handled the peak take leave, and the following month is understaffed for reasons nobody connects back to the launch.

Split the queue before you split the people

The most useful structural move during a spike is to stop running one queue. Split by intent so that the high-consequence conversations stay with experienced people and the repetitive ones go to whoever can process them fastest. This is a variant of the language-based routing described in the queue routing piece, and it works for the same reason: routing by what the conversation needs beats routing by arrival time.

A second split that is usually missing is by channel. During a spike, chat and email behave completely differently. Chat punishes delay in seconds; email punishes delay in days. Treating them as one backlog means the chat queue is always the loser, because its messages are shorter and look easier.

Write the plan before you need it

The deliverable that separates teams that handle launches well from teams that do not is a one-page document that exists before the launch. It needs five things.

Named owners for the spike, including who can edit the help centre and who can post a public status update — because during an incident the bottleneck is often permission rather than effort.

The threshold that triggers the plan. A specific number, agreed in advance, at which the team stops answering everything and starts triaging. Deciding this in the moment is what produces the worst outcomes, because a queue that is already behind makes everyone reluctant to spend time on process.

The holding message library, written in each supported language rather than translated during the incident. A holding message that reads like machine translation during an outage does more damage than no message, which is the failure mode discussed in the tone piece.

The deferral list — which intents get a holding response and which are never deferred.

The escalation rule for the spike itself: who decides that the plan is not working and switches to answering only the highest-consequence queue.

Restore, do not just recover

When volume returns to normal, the backlog does not disappear; it changes form. Messages answered late generate follow-ups, and follow-ups during a recovery week look like a second, smaller spike.

The practical response is to schedule recovery explicitly: a fixed block in the days after the spike where the team clears the backlog and, more importantly, writes down the top ten questions the spike produced. That list is the source material for the next launch’s deflection content, which is how the next spike costs less than this one.

If you want the wider context — how these spikes interact with quality measurement and with platform choices — the arguments in the quality metrics piece and the platform evaluation piece are written for exactly this kind of planning.

Share:fXintgwa
Telegram客服TG频道双向客服WhatsApp返回顶部