Skip to content
Communications7 min read

Scaling a contact centre

Doubling call volume does not require double the agents, and the reason it does not is the same reason a small centre feels chaotic while a large one feels calm. Staffing does not scale linearly, and planning as though it does is how centres end up simultaneously overstaffed and missing service levels.

Erlang CAsteriskVICIdialSIPWFM

A contact centre planning to grow asks a reasonable-sounding question: we are handling four hundred calls a day with twelve agents, so what do we need for eight hundred?

The intuitive answer is twenty-four. The correct answer is fewer, and understanding why changes how the whole growth is planned — because the same effect explains why a small team feels permanently under pressure while a larger one handling proportionally more feels comfortable.

Why staffing is not linear

Calls do not arrive evenly. They arrive randomly, which means clusters and gaps. A team must be staffed for the clusters, and in the gaps that staffing is idle.

The larger the team, the less this hurts. With three agents, two simultaneous calls plus one more means a queue. With thirty agents, the random clustering averages out across a much larger pool, so the same proportional surge is absorbed without anyone waiting.

This is the pooling effect, and it is why doubling volume requires substantially less than double the staff. It is formalised in the Erlang C model, which is the basis of essentially all contact centre workforce planning, and the specific numbers depend on your call volume, average handling time and target service level.

Occupancy: the trade-off nobody states

Occupancy is the proportion of an agent's logged-in time spent handling calls. Low occupancy looks like waste and every operation under cost pressure tries to raise it.

The trade-off is that queue time rises non-linearly as occupancy approaches full utilisation. Moving from moderate to high occupancy costs some waiting; moving from high to very high costs disproportionately more, because there is no slack left to absorb the random clustering described above.

Two consequences follow, and the second is the expensive one.

  • A service level target and an occupancy target are in direct tension. Both can be stated; both cannot be maximised. Which one yields should be an explicit decision rather than an emergent one.
  • Sustained very high occupancy is a staff retention problem. Back-to-back calls with no recovery time produces burnout, and attrition in a contact centre is expensive — recruitment, training, and a period of lower productivity while a replacement learns.

That second point is why cost-driven occupancy targets frequently fail to save money. The saving is visible in the staffing model and the cost appears in the attrition rate, where nobody attributes it back.

Shrinkage: the number that makes forecasts wrong

Agents are not available for the whole of their paid time. Breaks, training, meetings, coaching, system problems, holiday and sickness all remove availability, and the total is large.

Forecasts built on headcount rather than on available agent hours are wrong by whatever that total is, consistently, in the same direction. This is the most common reason a centre that is staffed on paper misses its service level in practice.

  1. 1

    Measure your own shrinkage rather than using a benchmark

    It varies enormously by operation, and an industry figure applied to your centre is a guess wearing a number.

  2. 2

    Split it into planned and unplanned

    Training and holiday are schedulable. Sickness and system outages are not. They need different treatment in the model.

  3. 3

    Forecast in available agent hours

    Not headcount, not full-time equivalents. The unit that matters is hours actually on the queue.

  4. 4

    Track forecast against actual, per interval

    A forecast nobody reconciles never improves, and reconciling at the daily level hides the intervals where the failures happen.

What breaks first on the infrastructure side

Growth exposes infrastructure constraints in a fairly predictable order, and knowing the order lets a centre address them before they bind rather than during a hiring push.

Constraints in the order they typically appear
ConstraintSymptom when it bindsUsually appears at
Carrier channel capacityCalls fail to connect at peakFirst — trunk capacity is bought, not automatic
Internet bandwidthAudio quality degrades at busy timesEarly, especially with remote agents
Database write throughputAgent screens slow, reports time outAs attempt volume rises
Recording storageDisk fills, platform stopsPredictably, and it is a hard stop
PBX or dialler CPUAudio problems, dropped callsLater than expected
Supervisor span of controlQuality and adherence driftAround every 12-15 agents

The last row is not infrastructure and is included deliberately, because it binds sooner than most of the technical constraints and is almost never planned for. Growing from twelve agents to thirty without adding supervision produces a quality problem that gets attributed to hiring standards.

Carrier capacity heading the list is worth acting on directly. Channel capacity is purchased, and a centre that grows headcount without increasing it discovers the limit at the busiest moment of the busiest day, which is exactly when the failure is most visible to customers.

Where the real gains usually are

Before adding agents, three changes routinely deliver more than a proportional headcount increase, and all three are configuration or process rather than capacity.

  • Pool the queues. Separate small queues per team or product each suffer the small-team inefficiency independently. One pooled queue with skills-based routing captures the pooling effect and needs fewer agents for the same service level.
  • Reduce handling time by removing work from the call. Screen pop, click-to-call and write-back remove seconds of manual lookup and typing from every single call, and at volume that compounds into a headcount equivalent.
  • Remove the calls that should not be calls. A meaningful share of contact volume in most centres is status enquiries that a notification or a self-service view would have answered. Those are the cheapest calls to eliminate because the customer prefers it too.

The third is the largest in most operations and the least often examined, because contact volume is treated as demand to be met rather than as a signal about something upstream. A recurring enquiry type is usually evidence of a process or communication gap, and fixing that is cheaper than staffing it forever.

Does doubling call volume require double the agents?

No, it requires meaningfully fewer. Calls arrive randomly rather than evenly, so a team must be staffed for the clusters and is idle in the gaps. Larger teams absorb that randomness across a bigger pool, so the same proportional surge causes less queuing. This pooling effect is formalised in the Erlang C model used for contact centre workforce planning, and the specific figures depend on call volume, average handling time and target service level.

Why does a small contact centre feel busier than a large one?

Because small teams are inherently less efficient and there is nothing to fix operationally. With three agents, a third simultaneous call means a queue; with thirty, the same proportional surge is absorbed. Small teams will have both queues and idle time simultaneously because the arithmetic allows no other outcome. Merging small separate queues into one pooled queue with skills-based routing is often the largest available efficiency gain and costs only configuration.

What is shrinkage and why does it matter?

Shrinkage is the proportion of paid agent time not available to handle calls — breaks, training, meetings, coaching, system problems, holiday and sickness. It is substantial, and forecasts built on headcount rather than available agent hours are consistently wrong in the same direction by that amount. It is the most common reason a centre staffed adequately on paper misses its service level in practice, and it should be measured for your own operation rather than taken from a benchmark.

Is high agent occupancy a good target?

Only up to a point, and it conflicts directly with service level. Queue time rises non-linearly as occupancy approaches full utilisation because no slack remains to absorb random call clustering. Sustained very high occupancy is also a retention problem — back-to-back calls without recovery time produces burnout, and attrition costs recruitment, training and a period of reduced productivity. Cost-driven occupancy targets frequently fail to save money because the saving is visible in the staffing model while the cost appears in attrition.

What infrastructure constraint binds first when a contact centre grows?

Carrier channel capacity, in most cases, because it is purchased rather than automatic and a centre that grows headcount without increasing it discovers the limit at the busiest moment of the busiest day. Internet bandwidth follows closely, particularly with remote agents. Less obviously, supervisor span of control binds around twelve to fifteen agents and is almost never planned for, producing a quality problem that gets misattributed to hiring standards.

Sources and further reading

Services This Relates To

Written by KYCONNECTS Engineering. Client names are withheld under confidentiality.

Talk Through Your Requirements

We typically respond within 4–8 business hours.