Home · Solutions · Customer service

Solution · Customer service

Every conversation scored on evidence, and the agent sees the result first

Service quality measured on every call, not on five a month

Call transcripts are labelled for reason, resolution, sentiment and required phrases; scores arrive with the passages behind them, and a team leader confirms every one that counts.

DepartmentalMicrosoft TeamsHuman in the loopAI where it earns its place
46,000calls a month reach this illustrative contact centre. Four hundred and fifty of them are ever scored, and that sample decides who gets coached.

Executive summary

Challenge

Your agents are rated on five calls out of five hundred, and the other 495 are never opened.

What changes

The first deliverable is not a model.

Business value

Feedback reaches the agent the next working day, rests on every transcribed call rather than five, and points at passages rather than impressions.

Systems involved

Power BI quality dashboards; the coaching register in Microsoft Lists; the quality record on SharePoint under Purview retention

Business problem

Service quality

Quality management exists because the call is the part of the service a customer actually experiences. Everything else can be read from a system; what was said, in what tone, and whether the customer got an answer can only be heard.

The flaw is known to everyone in the centre and fixed by nobody. The sample is chosen by the person judged on it, late in the month, under pressure: recent calls, short calls, calls that come to mind. Five out of five hundred explain nothing.

The consequences land on four groups. Agents get feedback five weeks after a conversation neither party remembers, attached to a number that shapes their appraisal. Team leaders spend three days a month listening instead of coaching. Compliance verifies required phrases on one percent of the traffic, and operations reads process faults as attitude.

Scale makes it worse. Thirty more agents mean more listening days and lower coverage, two leaders read the same behaviour differently, and calibration, the control that would hold the scores together, is cancelled first in a peak.

How it works today

This is quality management in most contact centres before anything is automated.

  1. PersonLate in the month a team leader opens the recording platform and picks five calls per agent, usually recent and short
  2. PersonEach call is played, often twice, and marked against an eighteen-criterion form, one tab per agent
  3. WaitingScores wait for the monthly one-to-one, so a call from the first week is discussed four or five weeks later
  4. PersonThe leader writes a coaching note in the same tab, often carried over from last month
  5. Risk of errorTwo leaders read the same behaviour differently and nothing reconciles them, because calibration is cancelled whenever volumes rise
  6. SystemThe tab totals feed the quarterly appraisal and, in many centres, the bonus calculation
  7. Risk of errorRecording notices, identity checks and required disclosures are verified only on the sampled calls
PersonWaitingRisk of errorSystem

Why the current process costs more than it appears

Nobody planned this work; it accumulated.

  • Listening and coaching compete for the same hours and listening wins, because it has a deadline. Three days spent marking calls are three days not spent coaching.
  • Five calls cannot separate an agent from a queue. A score pulled down by a confusing invoice line is coached against a person, month after month, while the cause stays where it is.
  • Scores that trace back to nothing are contested quietly rather than formally. Agents stop reading them, and the centre keeps paying for a measurement nobody believes.
  • Compliance is sampled at the same one percent. A disclosure that stopped being spoken in March surfaces in an audit, in a complaint, or not at all.

Cost of inaction

Twelve months of five calls per agent≈ €58,500
The same sample through three appraisal cycles≈ €175,500
Once the new queue adds fifteen agents (per year)≈ €68,250

A one percent sample produces a number every month, and a number is all anyone downstream asks for. That is why the method survives every review: the appraisal has a score, the quality report has a trend, and neither shows that both came from five calls chosen in the last week.

What grows meanwhile is not in the table. Agents who cannot see why they scored 78 stop treating the figure as feedback, and where attrition runs into tens of percent a year that costs more than the scoring does. Process faults keep arriving as coaching points, so the invoice line generating four hundred calls a month is explained to agents instead of rewritten.

Illustrative scenario

A plausible organisation with realistic proportions. The figures are there to be recalculated on your data; they are not a client result.

Organisation

The service centre of a European energy retailer: 90 agents on two sites serving household and small-business customers, on Microsoft 365 E3, with calls handled in Dynamics 365 Contact Center and Teams Phone for voice.

Volume

46,000 calls a month, about 511 per agent, in Polish and English. Five per agent are scored by hand, 450 a month, against a form with eighteen criteria.

Current process

Team leaders choose the calls, play them, complete the workbook and hold a monthly one-to-one. Quality reporting is the average of those tabs; calibration happens when the month allows.

Bottleneck

Roughly 25 minutes per scored call and about one percent coverage. Nobody can say whether a team's score describes the team or the sample, and the other 45,550 conversations have never been read.

Solution

Transcripts from queues covered by the recording notice are labelled by UiPath Communications Mining for reason, resolution, sentiment movement, required phrases and repeat contact. A draft scorecard follows, with the passage behind each criterion. The agent sees their result in Microsoft Teams first; a leader confirms or changes every score that counts.

Potential outcome

Coverage moves from one percent to every transcribed conversation, feedback reaches the agent the next working day, and marking days convert into coaching time. All of it is arithmetic on the assumptions below, not observed at a client.

Proposed solution

The first deliverable is not a model. It is a written answer to five questions: which queues may be analysed, what the customer hears when the call starts, how long a transcript is kept, who may open one, and who decides a score. The Polish labour code requires monitoring of employee communications to be set out in the work regulations, a collective agreement or an announcement, with employees informed at least two weeks beforehand. The data protection regulation adds the rule that matters most here: a decision with a significant effect on a person cannot rest on automated processing alone.

Then the model. Twelve months of your own transcripts train a taxonomy your quality specialists define and annotate: why the customer called, whether it was resolved, how sentiment moved, which required phrases were spoken, and whether the call repeats an earlier one. This is supervised classification with measured coverage, not a prompt over a recording. The scorecard above it is dull arithmetic: criteria, weights and thresholds, versioned and owned by your quality function.

Where the result lands decides whether any of it is accepted. The agent sees their own scorecard in Microsoft Teams the next working day, with the passages behind each criterion and a button that disputes it. Only then does the leader see a shortlist: disputes, compliance flags, low-confidence calls and a control sample. Coaching actions go into a Microsoft List with an owner and a date, the weekly review runs on one Power BI model, and Microsoft Purview retention labels remove transcripts on schedule.

Native capabilities used

UiPath Communications Mining in UiPath IXP (taxonomy trained by annotation, sentiment, extraction fields); UiPath Orchestrator queues, schedules and audit; UiPath Integration Service connectors for Microsoft Teams and Microsoft Dynamics 365 CRM; Microsoft Graph call record and transcript endpoints; Microsoft Purview retention labels and audit log; Power BI; Microsoft Teams and Microsoft Lists

What we build

The coverage rules deciding which calls are ingested; the taxonomy and annotation programme; the scorecard, weights and required phrases; the agent-first release in Teams and the dispute route; the confirmation queue and shortlist rules; the coaching register, calibration report and Power BI model

Custom integration

Mapping Teams Phone transcripts and Dynamics 365 Contact Center records into Communications Mining datasets, keeping speaker attribution and routing by language; writing the confirmed score and coaching action into the quality record on SharePoint

How the automated process works

  1. AutomationA call ends on a covered queue; the transcript and call record are collected through Microsoft Graph and matched to the Dynamics 365 Contact Center conversation
  2. AutomationCommunications Mining labels the transcript: reason, resolution, sentiment movement, required phrases, and whether it repeats an earlier call
  3. SystemOrchestrator calculates the draft scorecard from the labels and weights and stores the passage behind every criterion
  4. PersonThe agent opens their own scorecard in Microsoft Teams the next working day and either accepts it or disputes it on the same card
  5. AutomationThe leader's queue is a shortlist: disputes, compliance flags, low-confidence calls and a random control sample
  6. PersonThe leader confirms or changes each score and records the coaching action in a Microsoft List; no rating reaches an appraisal before this step
  7. SystemPower BI publishes quality by team, criterion and cause for the weekly review; Purview retention labels remove transcripts on schedule
AutomationSystemPerson

Human-in-the-loop model

Automation handles

  • Collecting transcripts only for the queues and periods the coverage rules allow
  • Labelling every conversation for reason, resolution, sentiment, required phrases and repeat contact
  • Calculating the draft scorecard and keeping the passage behind each criterion
  • Building the leader's shortlist, publishing the dashboards, running the retention schedule

People decide

  • Every score that counts: a leader confirms or changes the draft, and no rating reaches an appraisal, a bonus or a disciplinary step without that decision
  • What quality means: criteria, weights and required phrases are owned by your quality function, versioned like any policy
  • The coaching conversation, the development plan, and whether a pattern belongs to a person or a process
  • Disputes: a challenge is read by a person and the outcome recorded

Before and after

BeforeAfter
Conversations examined each month450 of 46,000every transcribed call labelled
Time from the call to the feedbackfour to five weeksthe next working day
Basis of a coaching conversationrecollection of five callsnamed passages from the agent's own calls
Required phrases verifiedon the sampled one percenton every transcript
Who sees a score firstthe leader, then the appraisal filethe agent, with a route to dispute it

Systems and integrations

The stack is deliberately short: one engine, one execution layer, one place where a person decides.

Inputs

  • Teams Phone transcripts and call records through Microsoft Graph
  • Dynamics 365 Contact Center conversation records
  • the scorecard definition on SharePoint
  • per-queue notice status

Automation layer

  • UiPath Communications Mining (IXP)
  • UiPath Integration Service
  • UiPath Orchestrator
  • UiPath Robots

Target systems

  • Power BI quality dashboards
  • the coaching register in Microsoft Lists
  • the quality record on SharePoint under Purview retention

Human touchpoints: the agent's scorecard in Microsoft Teams; the dispute thread; the leader's confirmation queue; the weekly review channel

Teams Phone transcriptsUiPath Communications MiningUiPath Integration ServicePower BI quality dashboardsthe agent's scorecard in Microsoft Teams

Technologies used

UiPath Communications Mining (IXP)

labels every transcript for reason, resolution, sentiment and required phrases

A
UiPath Robots + Orchestrator

schedule ingestion, calculate the draft scorecard, keep queues, retries and audit

A
UiPath Integration Service (Microsoft Teams and Microsoft Dynamics 365 CRM connectors)

posts scorecards and coaching tasks, reads Dynamics 365 Contact Center conversation records

A
Microsoft Graph (Teams Phone call records and call transcripts)

retrieves transcripts under named permissions an administrator can switch off

A
Microsoft Teams

the agent's scorecard, the dispute thread, the confirmation queue and the weekly review

A
Power BI

quality by team, criterion and cause, plus calibration variance between leaders

A
Microsoft Purview

retention labels on transcripts and evidence; the audit log of who opened what

A
Averified product capability (vendor documentation)

Illustrative economic model

Numbers you can check against your own data.

Illustrative model
450 calls scored by hand × 25 minutes each= 188 h / month
188 h × €26 fully loaded hourly cost= €4,875 / month
× 12 months≈ €58,500 / year
Annual leader capacity released (illustrative)≈ €58,500

Five calls per agent is where the model starts, so 90 agents give 450 scored conversations a month. The 25 minutes cover one of them end to end: picking the call, playing it once and usually twice, completing an eighteen-criterion form and writing the note. €26 is a fully loaded hourly cost for a team leader in a Central European contact centre. Coaching time is deliberately outside the table; what the arithmetic prices is the listening that must happen before coaching can start. Every figure is illustrative.

Run the numbers on your data

hours released per month
of annual capacity released

An illustrative estimate from your own inputs. It models released capacity; it is not a promise of savings.

Business benefits

  • Feedback reaches the agent the next working day, rests on every transcribed call rather than five, and points at passages rather than impressions
  • The leader's month moves from marking to coaching: the same hours produce conversations instead of scores
  • Required phrases are checked on every transcript, so a disclosure that stopped being spoken is found in days
  • Repeat contacts and transfers separate from agent behaviour, so a confusing invoice line is rewritten once instead of coached ninety times
  • New agents are coached from their own calls in the weeks when most of them decide whether to stay

The management view

  • Quality becomes a population measure, so a rise in transfers is investigated where it starts instead of coached where it lands
  • Calibration is visible: where two leaders read the same behaviour differently, the gap is a number rather than a suspicion
  • Coaching capacity rises without recruiting, because the constraint was never coaching skill, it was listening time

Board-level KPIs

share of conversations labelledfirst-contact resolutionrepeat-contact rateadherence to required phrasescalibration variance between leaderscoaching actions closed on time

Security and governance

Control is not an add-on.

  • Recording and transcription are announced, not discovered: customers hear the notice when the call starts, employees are informed before monitoring begins, and its purpose, scope and manner sit in the work regulations or an announcement as the Polish labour code requires. An uncovered queue is not ingested.
  • Retention is fixed before the first transcript is read: Purview retention labels delete transcripts and labels on the agreed schedule, and the confirmed score and coaching note outlive the recording.
  • Access is narrow and logged: an agent sees their own results, a leader their own team, quality and compliance the aggregates and the conversations they may open. Microsoft Entra ID groups decide it; the Purview audit log records every retrieval.
  • No performance or disciplinary decision is taken by the system. The model produces a draft with the evidence attached, a named person confirms or changes it, and the agent can dispute it first. Article 22 of the data protection regulation makes that a design rule rather than a courtesy.
  • Ingestion runs under a dedicated application identity with Graph permissions scoped to the transcripts it may read; no secret sits in a workflow, and an administrator can withdraw transcript access centrally.
  • Transcripts, labels and scores stay in your Microsoft 365 tenant and the UiPath Automation Cloud EU region; the model classifies your own conversations and drafts nothing.

Why now

01

Transcripts are created and deleted on a retention schedule, so the twelve months that could train a scorecard this quarter are not the twelve you will hold next year, while €4,875 a month of leader time keeps buying one percent coverage

02

Employee monitoring is regulated rather than forbidden, and notice periods, work regulations and the ban on decisions taken by automated processing alone are cheaper to build in now than to retrofit later

03

Communications Mining is a supervised classifier trained on your own calls, and Teams Phone transcripts are retrievable through Microsoft Graph under permissions an administrator controls

Relevant executive roles

Customer Service Director

Quality stops being a sample defended in a meeting and becomes a measure of the whole operation

CHRO

Appraisal and pay rest on evidence an employee can see and challenge, announced before the monitoring starts

COO

Repeat contacts and transfers become visible as process faults, so fixes land where the calls are generated

CIO

One supervised model over data the tenant already holds, with EU residency and an audit trail, not a separate speech-analytics platform

Common questions and objections

Are we allowed to analyse recordings of our own employees?

In Poland monitoring of employee communications is lawful when its purpose, scope and manner are set out in the work regulations, a collective agreement or an announcement and employees are informed beforehand. What is not allowed is a decision with a significant effect on a person taken by automated processing alone, which is why a leader confirms every score.

Automatic transcription is not good enough to judge anyone on.

It does not have to be. The model reads patterns across a whole conversation rather than grading a sentence, every criterion carries the passage that produced it, and whoever confirms a score can open the recording. Queues where transcript quality falls short stay on human scoring.

Our quality team already scores calls, so why change what works?

The team stays; what changes is what they work from. Instead of picking five calls they confirm a shortlist that carries the evidence, and the released hours go into coaching. Their scores also become comparable between leaders.

When this is not the right solution

  • Centres small enough that a leader can hear a meaningful share of the traffic, roughly under twenty agents
  • Calls that are not recorded or transcribed, or where the notice, the retention rule and the employee information are not agreed; then the first project is the governance
  • No agreement on what a good call is: if quality, compliance and operations cannot state the criteria, an automated scorecard only makes the disagreement arrive faster

A question for the next management meeting

Every agent in this company was rated last month on five conversations out of roughly five hundred: would this board accept a financial figure built on the same sample?

Implementation approach

A scope without ambiguity, before anything is signed.

We deliver

  • The governance groundwork first: queues in scope, the customer notice, the retention period, the access model and the employee information your legal and HR teams sign off
  • The label taxonomy and annotation programme run by your own quality specialists, with coverage targets that decide when it may be published
  • The scorecard: criteria, weights, required phrases, shortlist rules and the control sample, owned by your quality function
  • The agent-first release in Microsoft Teams, the dispute route, the confirmation queue and the coaching register in Microsoft Lists

We need from you

  • Twelve months of transcripts, or the ability to start collecting them, with queue, agent and language identifiers
  • The current scoring form, its weights, any regulator-required phrases, and the person who owns them
  • A decision from HR and legal on monitoring, retention and how quality results may be used in appraisal and pay
  • Two to four quality specialists to annotate during training and calibrate briefly each month

Stages

Governance

Notice, coverage, retention, access and the rule that no score is automatic, agreed with HR, legal and employee representatives

Discovery

Queues, languages, transcript quality, volumes and what the current form measures

Taxonomy and training

Labels defined with quality, annotated by your specialists, measured for coverage and consistency

Build

Ingestion, scorecard, the Teams release and dispute route, the coaching register, the Power BI model

Parallel run

Machine draft beside human scoring on the same calls, calibrated before any result counts

Go-live

Agent briefing, a first month under hypercare, then monthly calibration and taxonomy maintenance

Departmental. Effort is driven by the number of queues and languages, the quality of the transcripts your platform produces, and how much agreement there is about what may be measured.