How the scoring works, in detail
Dutified reports how its AI customers responded. It is feedback, not legal or compliance advice. Decisions stay with your organisation. Dutified finds the problems: what works, what does not, for whom and why. Your organisation's own writers and compliance team decide what to change.
This page describes consumer model 1.0, the version in use now. Every report records the version it ran on, and past reports never change. Back to Help.
The AI customers
60 AI customers read each test. 30 are like your customers: their ages, how they hear from you, how confident they are with money and going online, and how many are in a vulnerable circumstance. 30 more are going through something harder, such as a bereavement, illness or money worries, so that each circumstance has enough AI customers to say how it went for them.
Each AI customer is given only what a person like them would take in. How the communication reaches them matters: a letter may sit in a pile, most people decide whether to open an email from its subject line, and a long sentence can lose someone part way. Words they would not know are blanked out for them. Some stop before the end.
Known limits of consumer model 1.0, the version in use now. It decides sentence by sentence what each AI customer takes in, so a word they do not know can be blanked in one sentence and shown in the next. And their other traits (age, confidence with money or online, how they hear from you) are drawn separately from their circumstance, so some profiles do not fit their circumstance: for example someone who is not online but is confident online. Both are fixed in candidate model 1.7, which is waiting for the Dutified team's approval. Each report's AI customer pages show exactly what each one saw, so you can check.
The AI customers are simulated customers modelled on published survey data.
The Run a test choices
Dutified keeps a store of 180 AI customers for your organisation. Your 60 are the first of them. Each choice on Run a test changes which of them read the test:
- Product and age lean: the 30 like your customers are picked from the store to fit that product's customers, moved younger or older if you ask. The same choices on a retest pick the same people.
- Extra focus on customers in vulnerable circumstances: 45 of the 60 are going through something harder, not 30. Fewer are left like your customers, so the headline scores rest on fewer AI customers and their margins are wider.
- A saved group or people you pick: they always get a place among the 60. The report opens on them, and everyone else is still shown.
- The Consumer Duty customer check: always runs. It covers the four Consumer Duty outcomes, vulnerable customers and plain English, read by your AI customers.
- The rules check: an optional extra step, off until you ask for it. It checks the words against the FCA rulebook for your products (PRIN 2A and the rules that apply to what you sell) and shows its own findings on the report, separate from what your customers understood. Turning it on later reads the same words again; your AI customers are not read a second time.
The seven questions each AI customer answers
After reading, each AI customer answers the same questions, in their own words, as that person would. They are told not to guess what they did not read and not to answer like an expert.
- What is this asking you to do, if anything?
- What happens if you do nothing?
- What will this cost you, or what are you being charged?
- What is the risk to your money here?
- Is there a deadline, and what is it?
- What would you actually do next, in real life?
- Which words, sentences or numbers confused you?
They also say how sure they feel they understood it, from 1 (no idea) to 5 (completely clear), and how it made them feel, in a few words.
Marking: understood, partial, missed or wrong
A separate AI marker reads every answer against the key points your organisation gave: the things every customer must take away. For each key point it gives one of four marks:
| Mark | What it means | Counts as |
|---|---|---|
| Understood | The answer shows they got the substance, including any date, amount or condition that matters. | One whole point |
| Partial | They have the gist but miss something that matters, such as the date, the amount or who to contact. | Half a point |
| Missed | Nothing in the answer shows they got it, or they say they do not know. | Nothing |
| Wrong | They say something that contradicts it. | Nothing |
The marker also judges what each AI customer says they would do next: a sensible step, a step that is not enough to protect them, or a step that would leave them worse off, such as missing a deadline that matters. It notes whether they know how to act or get help.
An AI customer's understanding score is the share of the key points they got, with a partial mark counting half. Vague or hedged answers get no credit.
How the eight lenses are built
The eight lenses are Dutified's way of reading a communication against the FCA's rules: the four Consumer Duty outcomes, the three cross-cutting rules and the FCA's guidance on vulnerable customers (FG21/1). The FCA does not publish them as a list of eight. Each names where to look in the FCA Handbook or guidance; that is a pointer, not a ruling.
Each key point is sorted under the lenses it is about by the words in it: a point about a charge counts towards price and value, a point about a deadline towards avoiding harm. Every key point counts towards consumer understanding. A lens with no key points of its own uses the overall understanding score and says "not measured separately".
| Lens | Where to look | How it is worked out |
|---|---|---|
| Consumer understanding | PRIN 2A.5 | Average share of the key points each reader understood (partial counts as half). |
| Price and value | PRIN 2A.4 | Understanding of the key points about cost and charges. |
| Consumer support | PRIN 2A.6 | Share of readers who would take a sensible next step and know how to act, together with understanding of key points about what to do. |
| Products and services | PRIN 2A.3 | Understanding of the key points about the product itself and what is changing. |
| Act in good faith | PRIN 2A.2.1 | Understanding of the key points about risks and downsides, less any rule findings on balance. |
| Avoid foreseeable harm | PRIN 2A.2.8 | Share of readers who would not act harmfully, weighted by understanding of deadlines and anything that could cost them money. |
| Enable financial objectives | PRIN 2A.2.14 | Understanding of their options, weighted by whether their next step is sensible. |
| Customers in vulnerable circumstances | FG21/1 | Consumer understanding among readers with a vulnerable circumstance, taken from the second run where each vulnerable group is weighted up. |
Three lenses mix understanding with what the AI customers said they would do. The exact parts:
- Consumer support: 60% understanding of the key points about what to do, 40% the share of AI customers who know how to act.
- Avoid foreseeable harm: 50% understanding of the key points about deadlines and money, 50% the share of AI customers who would not act in a way that leaves them worse off.
- Enable financial objectives: 70% understanding of the key points about their options, 30% how sensible their next step is (a sensible step counts in full, a weak one half, a harmful one not at all).
So a lens that rests on one key point is not just that point's score: for these three, the next step counts too. Each AI customer's page shows their answer, their marks and their next step.
Rule findings from Dutified's rule library take points off the lens they are about: six points for a red finding and two for an amber one, up to a cap. Notes take nothing off. A lens scores Green at 80 or more, Amber from 60 to 79 and Red below 60.
Who struggled, by customer type
The report shows the scores for each type of customer: by age, by circumstance, by how they hear from you, and by confidence with money and online. Each is worked out on just the AI customers in that group, and links to them so you can see exactly who is behind it.
Circumstances come from the AI customers going through something harder, where each has enough of them. Everything else comes from the AI customers like your customers.
The verdict
The report gives one of three verdicts, Ready, Needs work or Not ready, then one plain sentence on what happened: what most AI customers took in, what many missed, and how AI customers in vulnerable circumstances did. The verdict describes how the AI customers got on. It is not a ruling and it does not say whether a communication meets any rule. It is a set of gates, not an average: any one red line makes it Not ready.
| Not ready | Any one of these:
|
|---|---|
| Needs work | Not Not ready, but at least one line for Ready is missed. |
| Ready | All of these:
|
The verdict does not look at each customer group on its own. So a group can do badly while the verdict is Ready: always read Who struggled as well.
One bar for everyone
AI customers in vulnerable circumstances are held to the same bar as everyone else: below 55% is Not ready and 75% or more is needed for Ready. Scores are not weighted: every AI customer counts once. Fairness shows as the gap between AI customers in vulnerable circumstances and everyone else, by circumstance and by age. Tests run before 30 September 2026 keep the settings they ran with, and their reports say so.
The headline numbers
- Understood it well enough to decide: the average understanding score of the AI customers like your customers who opened it, less points for rule findings.
- Customers in vulnerable circumstances: the same, for the AI customers going through something harder.
- Avoiding foreseeable harm: half understanding of the key points about deadlines and money at risk, half the share of AI customers whose next step would not harm them.
- Opened it, stopped reading before the end and would do something that leaves them worse off are shares of AI customers.
- Score of the worst-served tenth: the average of the tenth of AI customers who understood least, among those like your customers whose answer was marked.
Margins of error
Each score comes from a set of AI customers, so another set could give a slightly different score. The report puts a margin beside each score, such as "± 5". It is worked out from that test's own answers: 1.96 times the spread of the AI customers' scores, divided by the square root of how many there were. That is a 95% range. The more the AI customers differ, and the fewer there are, the wider it is, so it changes from test to test.
For the share of AI customers who took in a key point, the report gives the range the share could be in, worked out with the Wilson method, so it never runs below 0% or above 100%.
A difference between two versions that is smaller than the margin may not be real.
Stability
Measured on model 1.0: three made-up letters were each tested three times, each time by 100 different AI customers. Understanding moved by at most 7 points. The verdict did not change. Other scores moved more: Support by up to 12 points and the score for customers in vulnerable circumstances by up to 20 points. Your own letter has not been tested more than once unless you retest it.
Small groups
Groups of fewer than 20 AI customers are not shown in Who struggled. Anything resting on fewer than 100 AI customers' worth of evidence is labelled "small sample". Read those as a guide to where to look, not as a measurement.
Changes to the scoring
Every report says which scoring version it used. A report keeps its version; when the Dutified team removes a finding at review, the report is scored again with the version in use then, and the report and the governance trail both say so. Two versions scored differently can differ because of the method as well as the words, and the comparison says when that is the case.
- 1.4 (29 September 2026): margins worked out on the effective number of AI customers.
- 1.5 (30 September 2026): more everyday words count towards "Act in good faith"; two wording rules (what happens if you do nothing, and how to get help) recognise more ways of saying it; placeholders are never listed as confusing words; the Feedback view leads with everyone who did not fully take a point in.
- 1.6 (30 September 2026): the gap between AI customers in vulnerable circumstances and the others compares like with like. Verdicts and lens scores are worked out as in 1.5.
Feedback and Scores
Every report has two tabs, and both are always there. It opens on Feedback.
- Feedback says in plain words what works, what does not and why, which customers struggled most, the words that confused people, and what to fix, with a fix pointer for each. It has no percentages or margins.
- Scores shows the four outcomes, the margins and the verdict.
Fix pointers and suggestions
Dutified diagnoses, the organisation writes. Each finding comes with a short fix pointer, such as "Give the deadline as a date" or "Show the charge in pounds". Dutified does not write replacement wording for you.
On a finding that quotes your words, you can ask for a suggestion. It is clearly marked "Suggestion". You accept it, edit it or ignore it. Nothing is changed, tested again or approved for you. A new version starts only when you start one, and it includes only the suggestions you accepted or edited.
When you set up a test, we draft key points from the words. You can change them before you run it.
Template fields
Letters made from templates often hold fields such as "[amount]", "[date]" or "[account number]". The AI customers read an example value in their place: a date about four weeks ahead, an amount such as £125.00 or a number such as 12345678. A name field such as "[Customer name]" stays as it is. The report says what each field was read as. Your saved words, the marked letter and the Word download keep the fields as you wrote them.
Retests
A new version of a test is read by the same AI customers as the version before, with the same traits, reading the same way. So any change comes from the words, not from a new set of AI customers. Only those whose words changed are asked again; the others keep their answers.
Accessibility of your communication
Accessibility checks on your own communication are coming soon. Nothing in a report claims that a communication is accessible.
In short
- 60 AI customers read the communication as people like them would.
- Their answers are marked against your key points.
- The marks become four outcomes and eight lenses, for every type of customer.
- Margins of error and small sample labels say how far to trust each number.
- One bar for everyone, and no weighting.
- Each finding has a fix pointer. Your organisation writes the words and makes every decision.
No customer data
Dutified does not ask for customer data. Personal details found in a communication are replaced with placeholders in your browser or on our server before any AI customer reads it.
Feedback, not advice
Dutified reports how its AI customers responded. It is feedback, not legal or compliance advice. Decisions stay with your organisation.
Rule references say where to look in the FCA Handbook. They are not a ruling.