How the scoring works
What a score is
You give the key points: what every customer must take away. Each AI customer reads your letter the way a person with their traits would, then says in their own words what it told them. Each answer is marked against each key point: understood, partly, missed or got it wrong. Partly counts as half. A score is the share of the key points the AI customers took in.
The four outcomes
- Understanding: all the key points.
- Price and value: the key points about what people pay.
- Products and services: the key points about what the product does and what is changing.
- Support: whether customers know what to do next and where to get help.
The verdict
The report gives one of three verdicts, then one plain sentence on what happened:
- Ready: understanding is 75% or more, for everyone and for AI customers in vulnerable circumstances.
- Needs work: in between.
- Not ready: understanding is below 55%, for everyone or for AI customers in vulnerable circumstances.
A key point missing from the letter, or AI customers likely to do something that harms them, can also lower the verdict.
There is one bar for everyone. No one's answers count more than anyone else's. The report shows the gap between AI customers in vulnerable circumstances and everyone else.
Who the AI customers are
60 read each test. 30 are like your customers (percentages you give, or FCA Financial Lives 2024 figures for the products you offer). 30 are going through something harder, such as a bereavement, money worries or poor health, so you can see who struggles. With extra focus on vulnerable circumstances it is 45 of the 60. Each has an age, a confidence with money and online, and sometimes a circumstance. Meet them on Your customers.
The choices on Run a test pick which of your stored AI customers read it. What each choice changes
What it is good for
Finding the sentences and points people are likely to miss before a letter goes out, and seeing who struggles most.
What it is not good for
It is not a judgement on whether a letter meets any rule. It is synthetic customer data.
How sure you can be
Your headline scores come from the 30 AI customers who follow your customer mix, less any who did not open it. Each score on the report has its own margin beside it, such as "± 5". It comes from that test's own answers: how much the AI customers' scores differed and how many there were, at 95% confidence. So it changes from test to test. A difference smaller than the margin may not be real. Groups under 20 AI customers are not shown.
Stability. Measured on model 1.0: three made-up letters were each tested three times, each time by 100 different AI customers. Understanding moved by at most 7 points. The verdict did not change. Other scores moved more: Support by up to 12 points and the score for customers in vulnerable circumstances by up to 20 points. Your own letter has not been tested more than once unless you retest it.
Versions
Every report records the model and scoring versions and who the AI customers were, so you can see exactly what produced it. Results may vary a little each time a test is run. The model is checked against 20 made-up letters whose results are known. Every rule, in detail
Dutified reports how its AI customers responded. It is feedback, not legal or compliance advice. Decisions stay with your organisation.