Insights

Using AI in complaints handling: workflow design and human oversight

Where AI can assist complaints teams, which decisions should remain controlled and how to measure quality, fairness and customer outcomes.
Insurance AI 7 min read

Complaints handling is an attractive AI use case because the work is document-heavy, time-sensitive and rich in repeatable tasks. It is also a poor place for vague automation. A complaint can involve vulnerability, disputed facts, redress, regulatory duties and a customer’s right to challenge the outcome.

The most useful design question is not “Can AI handle complaints?” It is “Which parts of the workflow can AI support, under what controls, and how will the business know that customer outcomes remain sound?”

This article is operational guidance, not legal advice. Firms should involve their compliance, legal, privacy and security specialists for their circumstances.

Break the process into distinct tasks

Treating complaints handling as one use case hides important differences in consequence and complexity. Map the process as separate steps:

  1. Receipt and acknowledgement.
  2. Classification, routing and prioritisation.
  3. Evidence gathering.
  4. Chronology and issue summarisation.
  5. Investigation and policy analysis.
  6. Outcome and redress decision.
  7. Communication, recording and quality assurance.

AI may have a different role at each step. It could extract key dates, find missing documents and draft a chronology with source references. That does not imply it should decide whether the complaint is upheld or calculate redress without review.

The FCA-commissioned Mills Review describes a spectrum from preparing information and recommending actions to executing low-risk tasks and escalating complex cases. That is a useful way to design bounded assistance rather than an all-or-nothing automation programme.

Good early use cases

Evidence retrieval and case preparation

AI can help locate relevant interactions, policy documents and case records, then organise them for a handler. The output should link back to the source so the user can verify it.

Measure whether the system finds material evidence, not just whether the summary reads well. A missing telephone note or policy version can change the case.

Chronology and issue summarisation

A structured draft can reduce the time spent reconstructing a long history. Define required fields such as customer concern, product, relevant dates, prior responses, requested resolution, vulnerability indicators and unresolved questions.

The user should be able to see uncertainty and conflicting evidence. The system should not turn an allegation into an established fact through confident wording.

Routing and prioritisation

Classification can direct cases to the right team and flag potential urgency, but the classes need operational meaning. Test false negatives closely for vulnerability, deadlines and severe harm.

Provide a route for users to override the category and record the reason. Those corrections become valuable monitoring data.

Drafting controlled communications

AI can prepare acknowledgements, requests for information and sections of a final response using approved content. Customer-specific reasoning, commitments and outcome language require stricter control.

Use current templates and policy sources. Check tone, clarity, required information and unsupported statements. The employee sending the communication remains responsible for it.

Quality-assurance support

AI can compare a completed case against a checklist, identify missing evidence and select cases for human QA. It can broaden the coverage of review, but should not become the only control over its own outputs.

Keep high-consequence judgement controlled

The final outcome often depends on context, evidence weighting, regulation, policy and judgement. AI can help structure this work, but the authority to decide should be explicit.

Use a risk-tiered model:

Assist: retrieve, extract, summarise and draft. A person performs the substantive assessment.

Recommend: suggest a route or outcome with evidence and uncertainty. An authorised reviewer makes and records the decision.

Execute a bounded step: complete a low-impact, reversible action within defined rules, with monitoring and exception handling.

Do not automate: actions with unacceptable impact, weak detectability or no meaningful route to correction.

The boundary should be enforced through permissions and workflow states. A drafting component should not have the credentials to issue payment or send a final response.

Design meaningful human oversight

Human review fails when it becomes a rubber stamp. Reviewers need enough time, context, authority and competence to challenge the output.

The interface should make verification easy:

  • Show source references beside key claims.
  • Highlight missing or conflicting information.
  • Separate facts from generated recommendations.
  • Make editing and rejection straightforward.
  • Record material changes and override reasons.
  • Route high-risk cases to appropriately skilled staff.

Monitor the review process itself. A falling override rate is not automatically good news. It may indicate better output, or it may show over-reliance.

The ICO human review audit framework highlights reviewer competence, authority, testing, logs and fallback. The ICO notes that parts of its guidance are under review following changes in UK data law. Firms should check the current requirements for automated decision-making and profiling.

Build the evaluation around complaint risk

A general language-quality score is inadequate. Test the properties that protect the customer and the investigation.

Evidence accuracy

Are dates, amounts, events and quotations correctly tied to the record? Does the system distinguish the customer’s statement from the firm’s evidence?

Completeness

Does it identify every complaint point, requested resolution, relevant interaction and unresolved question? Use a required-facts checklist.

Policy and process compliance

Does the output follow the current procedure and required communication content? Are deadlines and escalation triggers identified correctly?

Fairness and vulnerability

Does performance differ for channels, products, communication styles, language needs or other relevant groups? Can the system recognise indicators that require adapted support without making unjustified inferences?

Grounded reasoning

Are suggested conclusions supported by authorised evidence and policy? Penalise invented facts and overconfident conclusions heavily.

Workflow usability

Does assistance reduce preparation and investigation time after checking and correction? Does it improve QA scores and reduce rework?

Use experienced complaint handlers to create and review the evaluation set. Include upheld and not-upheld cases, ambiguous evidence, long case histories, missing records, vulnerable customers and uncommon products.

Measure customer and operational outcomes together

Track a balanced set of measures:

Quality: material omission rate, severe factual error rate, QA score, rework.

Timeliness: preparation time, end-to-end resolution time, ageing and deadline risk.

Customer: repeat contact, reopen rate, escalation, complaint about complaint handling and clear-communication measures.

People: appropriate adoption, correction effort, confidence, workload and over-reliance indicators.

Control: override reasons, cases routed outside scope, incidents, access failures and model or policy changes.

The Financial Ombudsman Service’s discussion of AI and consumer complaints stresses that firms remain accountable for fair outcomes when they use AI. That accountability should be visible in the measures and review process.

Protect data across the workflow

Complaint files can contain identity data, health information, financial details and third-party correspondence. Map what enters the system, what is retained, where it is processed and who can access it.

Apply data minimisation, role-based access, appropriate retention and redaction where needed. Review supplier terms and sub-processors. Ensure logs help investigate a result without creating a second uncontrolled copy of the case file.

Test the risk of instructions embedded in documents or messages. Retrieved content should be treated as data, not as trusted system instruction.

Launch in controlled stages

A sensible sequence is:

  1. Offline evaluation: test on a representative historical set.
  2. Shadow mode: generate outputs without showing or using them, then compare with completed work.
  3. User trial: allow a trained group to use the tool on defined cases with enhanced review.
  4. Limited release: expand volume within approved eligibility and monitor daily.
  5. Routine operation: move to a stable review cadence with sampled quality checks and incident management.

Define pause criteria in advance. Examples include a severe unsupported conclusion, data exposure, a critical routing miss or performance below a segment threshold.

Questions for a design review

  • Which exact tasks are assisted, recommended or executed?
  • Which users and case types are eligible?
  • What evidence is shown to the reviewer?
  • What errors would be difficult for a user to detect?
  • Who can override, pause and change the system?
  • How are customers told about AI where transparency is required or helpful?
  • How can a customer challenge an outcome and reach a person?
  • What happens when the system or a dependency is unavailable?
  • Which production evidence will trigger retraining, redesign or withdrawal?

The practical opportunity

The opportunity is not to remove judgement from complaints handling. It is to reduce the mechanical burden around that judgement, improve access to evidence and make quality controls more consistent.

Start with a bounded step, build evaluations around real complaint risk and make the reviewer more capable rather than less responsible.

For a related pattern, see customer contact summarisation. To assess a complaints workflow in your organisation, talk to Sorsana.

  • Complaints handling
  • Human oversight
  • Insurance AI