Proven Duty
File reviews & suitability

What Do 500 Scored Suitability Files Say About Consumer Duty Readiness?

Rubric check-area patterns from anonymised suitability file scorecards: where Consumer Duty evidence holds up and where it falls down at file level.

Nick Thorp7 min read

What does scoring a suitability file actually test?

A scored review tests one thing: whether the evidence the FCA expects is present in the file, read cold. COBS 9.2.1R requires a firm to take reasonable steps to ensure a personal recommendation is suitable for the client, and the reasoning sits in the suitability report. If the reasoning is not in the document, the review cannot credit it — intentions, calls the adviser remembers, and knowledge in the adviser's head do not count.

Regarding file-level evidence, our rubric scores each check area against a published citation — COBS, PROD, PRIN 2A or FG21/1 — and returns Pass, Amber or Fail with a line-level reference. Human override sits on top: the tool scores, the reviewer decides. That distinction matters, because supervision tests records, not assurances.

Regarding the brief of this post, every pattern below comes from anonymised scorecards produced by that rubric, framed qualitatively. No client is identifiable, and no figure is claimed beyond what the pattern shows.

Supervision tests records, not intentions. (Source: FCA COBS 9.2.1R)

Which check areas score best across anonymised reviews?

Client fact-find sections score strongest. Attitude to loss, capacity for loss, objectives and time horizon are documented reliably across the files we have seen, because firms have captured them for years under pre-Consumer Duty obligations. The muscle memory exists, and the report templates carry the fields.

Costs and charges disclosure performs reasonably too. The FCA's Consumer Duty publications hub has kept the price and value expectations visible since the Duty came into force, and most templates now state charges clearly.

Regarding the gap between strong and weak areas, the split is instructive: the sections that evidence well are the ones with a long pre-Duty history. The sections that struggle are the ones the Consumer Duty added or sharpened. Readiness, in practice, is a question of how far a firm's documentation has caught up with PRIN 2A.

Supervision tests records, not intentions. (Source: FCA PRIN 2A)

Where do suitability files most often lose marks?

Three check areas fail or amber most often, and the pattern repeats across firms of similar size:

Check areaRubric citationCommon failure pattern
Consumer understandingPRIN 2A.3The report explains the product fluently but records no evidence the client could follow it — no plain-language framing, no check of comprehension.
Alternatives consideredCOBS 9.2.1RThe recommendation is asserted as suitable without naming what else was weighed or why it was set aside.
Ongoing service rationalePRIN 2A.6The service promise is stated, but the file does not connect the service to the client's objectives or record what it must deliver.

Table: The three check areas that most frequently score Amber or Fail in anonymised rubric reviews.

Regarding why these three cluster together, each one asks the file to demonstrate reasoning about the client's position rather than describe a product. Product description is easy to template. Client-specific reasoning is harder, and it is exactly what the Duty's outcomes require. Industry commentary on the four outcomes has observed the same weighting problem, with reports leaning on products and value while underweighting consumer understanding and support.

Why is the alternatives section the sharpest Consumer Duty test?

A personal recommendation only carries meaning against what it was chosen instead of. COBS 9.2.1R asks for suitability; the alternatives check asks the file to show the comparison. Files that name realistic alternatives — a lower-cost route, doing nothing, a different wrapper — and give the reason for the choice demonstrate the reasoning directly. Files that skip the comparison leave the reviewer an assertion, and an assertion cannot score Pass.

Regarding fair value connections, the alternatives check also links to PROD 3.3 and the fair value framework expectations under PRIN 2A.4. A file that records why a chosen route beat the alternatives is, in effect, evidencing value judgement at the individual client level — the layer the fair value outcome is hardest to demonstrate at.

Regarding proportionality, CP26/23 has put the calibration of Consumer Duty expectations under consultation, but the direction of travel does not remove the need for documented reasoning. It changes the scale of what is proportionate, not whether the file must show the comparison.

Supervision tests records, not intentions. (Source: FCA COBS 9.2.1R)

What do vulnerability entries in files reveal?

Vulnerability logging scores as a split pattern. Many files now flag characteristics of vulnerability — health events, life changes, capability concerns — which shows FG21/1 has landed in practice. The failure sits one line below the flag: the file records that the client may be vulnerable but not what the firm did about it. FG21/1 expects the firm to recognise the characteristics and then record the adjusted approach — different communication, extra time, a family member present, a slower decision.

Regarding the FCA's multi-firm work in this area, its March 2025 vulnerable-customers review found positive practice alongside a consistent weakness in evidence — firms acting but not recording. Our file-level pattern matches that finding exactly: the flag exists, the adjustment does not.

A flag without a recorded response is an Amber at best, because the reviewer cannot tell whether support was actually tailored. More on this evidence gap sits in our guide to vulnerable customers under the Consumer Duty.

Supervision tests records, not intentions. (Source: FCA FG21/1)

How should a small firm read its own review findings?

The grade on any single file matters less than the pattern across a sample. Review programmes exist to surface recurring issues, emerging risks and weaknesses in training, process or supervision — that is the stated purpose of a meaningful sample. Where the same check area fails across multiple files from different advisers, the root cause is the template or the process, not the individual.

Regarding what to do with the finding, the practical unit of change is the template or the workflow, not the file. Rewriting the alternatives prompt in the report template, or adding a vulnerability-response field, clears the same failure across every future file without retraining anyone.

Regarding evidence for supervision, the sequence matters: scores over time, the human judgement applied, the change made, the scores after the change. That closed loop is what Consumer Duty outcomes monitoring is for, and it is far more persuasive in a supervision conversation than a folder of perfect individual files. Our file review and suitability hub sets out the full method.

Supervision tests records, not intentions. (Source: FCA FG22/5)

About the Author: Nick Thorp is the founder of Proven Duty, built AI compliance tooling for UK advice firms, and writes about what Consumer Duty means at file level — evidence, records and outcomes rather than policy summaries.

Frequently asked questions

What does scoring a suitability file actually test?

A scored review tests whether the evidence the FCA expects is present in the file itself. PRIN 2A and COBS 9.2.1R require reasoned, suitable advice; if the reasoning is not recorded, the review cannot credit it. Scoring turns that expectation into check areas with citations and a Pass, Amber or Fail.

Which check areas score best across anonymised reviews?

Attitude to loss and capacity for loss, and the factual client data sections, tend to score strongest. Firms have documented these elements for years under COBS obligations, and the supporting evidence is usually already in the report template. The recurring weakness sits in outcome-specific reasoning, not basic fact-gathering.

Where do suitability files most often lose marks?

The same three areas recur: consumer understanding evidence, alternatives considered, and ongoing service rationale. Files describe the recommendation fluently but rarely record why the client should understand it, what else was weighed, or what the service must deliver. These map to PRIN 2A.3, COBS 9.2.1R and PRIN 2A.6.

Why is the alternatives section the sharpest Consumer Duty test?

COBS 9.2.1R requires a personal recommendation; a recommendation only carries meaning against what it was chosen instead of. Files that name realistic alternatives and give the reasoning for the choice demonstrate the comparison. Files that assert superiority without comparison leave the reviewer nothing to score except the assertion.

How should a small firm read its own review findings?

Individual grades matter less than the pattern. Across a sample of files, recurring check-area failures point at process weaknesses rather than one-off slips, which is the insight a review programme is supposed to produce. Fix the template or the process once, and the pattern changes across every future file.

What turns review findings into supervision-ready evidence?

A documented pattern of review, root-cause analysis and change is what supervision recognises as monitoring. A board-ready evidence pack that shows check-area scores over time, the human override applied and the action taken demonstrates the Consumer Duty cycle at file level, rather than a single snapshot of file quality.

Guidance based on published FCA material. This article is not regulatory advice.

See what your files are missing

Run one suitability report through the AI file review and get a Pass / Amber / Fail score against Consumer Duty checks.

Try the free file review