Skip to content

Case study / AI compliance workbench / 2025–2026

Veriscope AI-assisted AML investigations

I redesigned evidence review for AML investigations. The share of investigations returned for rework fell by 20%, while claim-level AI explanations helped users detect more AI errors in a controlled study.

The product name and branding elements have been changed to maintain confidentiality. The original product solutions and the results of their implementation are shown.

Role
Product Designer
Team
2 designers / 3 devs / 1 PM
Date
Dec 2025 – Jun 2026
Platform
Desktop web
Veriscope investigation queue beside an evidence graph

01

Overview

Project info

AML investigators turn fragmented alerts and records into decisions that must withstand quality review. I designed the queue, evidence comparison and claim-level AI explanations so investigators could resolve conflicts and request missing records before submitting a case. After implementation, fewer investigations were returned for rework, and users detected more AI errors in a controlled study.

−20% Investigations returned for rework: 30% → 24% Relative change in the share of investigations returned for rework
+12.3% AI errors detected by users: 65% → 73% Controlled study of users' detection of AI errors
90% Correct recommendations recognised, before and after Recognition of correct AI recommendations before and after the change

02

The challenge

A risk score alone did not show which evidence mattered, where sources disagreed or when automation should stop. The product had to reduce evidence assembly while keeping every outcome with the investigator.

Traceability
Every decisive AI claim needed a visible path to a transaction, KYC record, registry entry or document.
Decision authority
AI could rank cases, summarise evidence and flag conflicts. Only an investigator could clear or escalate a case.
Safe failure
Low confidence, missing data and conflicting sources had to block finalisation and offer a specific recovery action.

03

Why AI

The investigator must compare many records, connect entities and explain a suspicious pattern. AI is useful for preparing that review when its claims remain inspectable.

Cross-source synthesis

The model links transactions, entities, KYC data, registry entries and documents around one case.

Pattern detection

It surfaces transaction paths and ownership inconsistencies that deserve investigator attention.

Review preparation

It drafts sourced claims and identifies missing evidence before the human decision.

04

AI system model

The system converts fragmented records into reviewable claims and records the human outcome.

  1. 01

    Collect

    Bring the alert, transactions, KYC data, registry entries and documents into one case.

  2. 02

    Resolve

    Match entities and infer relevant risk patterns while retaining source identity.

  3. 03

    Explain

    Produce claims with confidence, provenance and missing-evidence status.

  4. 04

    Review

    Let the investigator compare sources and resolve ambiguity before deciding.

  5. 05

    Record

    Write the confirmed decision and its evidence history to an immutable log.

05

Key decisions

  1. 01 Evidence

    Define the evidence contract

    I defined the minimum source information for a claim: origin, timestamp, controller, confidence and missing evidence. This contract shaped both the interface and the evaluation.

  2. 02 Authority

    Keep judgement with the investigator

    I kept clear and escalate as human actions. The model prepares a recommendation, while unresolved conflicts pause it until the investigator compares the records and addresses the missing evidence.

  3. 03 Evaluation

    Evaluate investigation quality

    The evaluation covers investigations returned for rework and users' assessment of AI recommendations. A controlled study measured detection of AI errors; recognition of correct recommendations was tracked alongside it.

Reviewers need a source for every claim that changes the decision.

06

The solution

Evidence-led investigation queue

Risk, owner and SLA appear in the first viewport. Filters support rapid triage, while the active case and its workflow remain visible in a separate navigation rail.

Observation
Analysts needed to find urgent cases before opening each alert.
Decision
Expose risk drivers, ownership and time pressure in the queue.
Effect
Investigators can review priority, ownership and deadlines directly from the queue.
Investigation queue with visible risk, owner and SLA.
Veriscope queue prioritised by evidence urgency and SLA

Investigation queue with visible risk, owner and SLA.

Transaction path with decisive evidence

The evidence graph keeps entities and transactions beside the records that support the suspicious path. Each source shows its origin, date and conflict state.

Observation
A detached score hid the route that made the activity suspicious.
Decision
Place the transaction path and decisive source rail on the same screen.
Effect
Investigators can inspect the path without reconstructing it across tools.
Evidence graph with a linked source rail.
Veriscope evidence graph showing the suspicious transaction path and linked sources

Evidence graph with a linked source rail.

Claim-level AI rationale

The recommendation is split into claims. Confidence, source coverage and missing evidence sit next to each conclusion, so the investigator can review the model one statement at a time.

Observation
A single confidence score did not reveal which conclusion needed attention.
Decision
Break the rationale into sourced claims with local confidence and coverage.
Effect
In a controlled study, users detected 73% of AI errors, up from 65%, a 12.3% relative increase. Recognition of correct recommendations remained at 90%.
AI rationale reviewed claim by claim.
Veriscope AI rationale split into claims with confidence and linked sources

AI rationale reviewed claim by claim.

Conflict recovery before decision

When ownership sources disagree, Veriscope pauses the recommendation. The investigator compares both records, sees the missing tie-breaker and requests the exact evidence needed to continue.

Observation
Conflicting records and missing sources could remain unresolved when an investigation reached quality review.
Decision
Compare the conflicting records, identify the missing source and offer a targeted evidence request before quality review.
Effect
After implementation, the share of investigations returned for rework fell from 30% to 24%, a 20% relative reduction.
Human review with recommendation paused.
Veriscope human review showing conflicting ownership evidence and a recovery action

Human review with recommendation paused.

07

Human oversight and recovery

Failure and recovery states

Low confidence

When
A decisive claim falls below the review threshold.
Product response
The recommendation pauses and the weakest claim is brought into focus.

Conflicting sources

When
Two authoritative records disagree on a material fact.
Product response
Both values are shown side by side and finalisation is blocked.

Missing or stale data

When
Required evidence is absent or outside its accepted freshness window.
Product response
The interface identifies the missing record and offers a targeted evidence request.

Human override

When
The investigator reaches a different conclusion after reviewing the sources.
Product response
The original model state and the confirmed human rationale remain in the audit history.

Safety and review controls

Visible provenance

Every decisive claim links to named evidence with a timestamp and controller.

Local confidence

Confidence is attached to the claim it describes instead of being reduced to one case score.

Blocked ambiguity

An unresolved conflict prevents clear or escalate actions from being recorded.

Immutable history

Model states, source changes and human actions remain available for audit.

08

AI evaluation

The results cover two aspects of the design: investigation quality at handoff and users' ability to assess AI recommendations. Returns for rework measure the workflow outcome; the controlled study measures human detection of AI errors.

Evaluation methods

Returns for rework

The share of investigations returned for rework fell from 30% to 24%, a 20% relative reduction.

Controlled study of AI error detection

Users' detection of incorrect AI recommendations increased from 65% to 73%. The measure describes how people assess recommendations, rather than the model's accuracy.

Recognition of correct recommendations

Recognition of correct AI recommendations was 90% before and after the change. This measure is reported separately from the detection of incorrect recommendations.

09

Results

−20% Investigations returned for rework: 30% → 24% Relative change in the share of investigations returned for rework
+12.3% AI errors detected by users: 65% → 73% Controlled study of users' detection of AI errors
90% Correct recommendations recognised, before and after Recognition of correct AI recommendations before and after the change

After implementation, the share of investigations returned for rework fell from 30% to 24%. In a controlled study, users' detection of AI errors increased from 65% to 73%, while recognition of correct recommendations was 90% before and after the change. These results describe investigation quality and human review of AI recommendations.

10

Reflection

I focused the design on comparing evidence and resolving gaps before quality review. The results connect those choices to fewer returned investigations and better detection of incorrect AI recommendations. Tracking correct recommendations alongside errors keeps the evaluation focused on informed judgement, with the final decision remaining with the investigator.

Flowcast treasury forecasting workspace

Next case

Flowcast / AI treasury and liquidity copilot

Continue

11

Let’s Connect

Building an AI, FinTech or Enterprise product? I can help shape complex decisions into clear workflows with evidence, controls and measurable outcomes.

Download CV