Fast where it is safe, human where it matters.
A dispute assistant for LATAM Bank, in Spanish and Portuguese. It closes the safe cases in seconds and hands the rest to a person, case ready.
Agents decision
Pick a message, in Spanish or Portuguese, and press Send. Watch the robots hand the case along. The model robots read the message and write the reply. The one in the middle is plain code, and only the backend writes to the bank.
Pick a message
Pick a message and press Send. The robots answer here.
No reconozco el cobro de 389,87 USD en Moda Express.
The messages are examples, and each one is paired with the charge behind it so the Judge has facts to decide on. The Judge is the real order of checks, written as code. The Reader and the Speaker are described, not run.
- 1 understand.agent
Reader
model
Waiting.
- 2 decide.agent
Judge
code
Waiting.
- 3 act.agent
Clerk
backend
Waiting.
- 4 verify.agent
Checker
backend
Waiting.
- 5 escalate.agent
Courier
code
Waiting.
- 6 respond.agent
Speaker
model
Waiting.
Reader (model)
- Reads
- The customer's message.
- Writes
- A guess of the intent, with a confidence.
- Cannot
- It reads. It cannot decide or close anything.
Click a robot to see what it does.
The customer writes
In Spanish or Portuguese, in plain words. The assistant finds the charge and checks it against the policy.
Click the screen to enlarge it.
A person takes it
An approved charge never closes alone. The customer is told a person will review it, and the case moves to the queue with the facts attached.
Click the screen to enlarge it.
The case arrives prepared
Verified facts, the rule that applied, the customer's message and the agent's trace, on one screen.
Click the screen to enlarge it.
The pitch
The three-minute video and the six slides, here on the page. Both can be downloaded.
0:00 / 0:00
1 / 6LATAM Bank dispute assistant
What we measured, and what we did not
372 of 372.
Test cases where it must not resolve alone, and it never did. A small sample never proves zero risk: the upper bound is about 1%.
Cases tested
549, end to end. Offline results on cases the team generated, not a production measurement.
Intent classification
100% with the model against 49% with keywords, on a blind set of 228 messages.
Still missing
The data is synthetic, the login is a sandbox, and there is no load test or alerting. The repository lists what a real bank would need first.
