Proven Duty
File reviews & suitability

How Many Files Is a Defensible Review Sample?

No FCA rule fixes a file review sample size. Defensibility rests on a named method, coverage across outcomes, and findings on record — not the count alone.

Nick Thorp6 min read

What does 'defensible' mean for a file review sample?

Regarding file review sample size, the word that carries the weight is 'defensible', not 'correct'. A defensible sample is one a third party — a supervisor, a visiting associate, an ombudsman — can trace from end to end: the population it drew from, the method that picked the files, the findings per file, and the actions that followed.

Four components carry that weight. A stated method names how files were chosen. Coverage shows the sample touches the outcomes under PRIN 2A rather than one product or one adviser. Volume gives the sample enough files for patterns to surface. Records keep findings and follow-up in writing, per file.

Counts sit inside that structure. A sample of 30 with no named method is a guess; a sample of 30 with a method, spread across the firm, and findings on record is evidence. The file review and suitability hub covers what the file itself must show; the sample decides how much of it supervision sees.

Supervision tests records, not intentions. (Source: FCA FG22/5)

Do the FCA rules set a fixed file review sample size?

No FCA publication fixes the number. COBS 9.2.1R states the suitability standard a file must evidence; PRIN 2A sets the outcomes the firm must monitor; FG22/5 guides firms on meeting the Duty. None of the three names a count for internal review samples. The FCA's Consumer Duty publications hub gathers the guidance; the count stays with the firm.

Regarding Consumer Duty monitoring, the FCA's own practice points the same way. The September 2026 good and poor practice review of payments firms examined 'a sample of payments firms' — the regulator samples too, and it publishes the method and the findings rather than a quota.

CP26/23, the June 2026 consultation on proportionality, keeps the calibration question open for small firms. Consultation papers consult; final rules decide. A firm watching that space reads the final text before recalibrating anything.

The Handbook states the standard; the sample is the firm's evidence for it. (Source: FCA COBS 9.2.1R)

Which sample size methods can a firm cite?

Three families of method appear in the published literature, and each comes with a number a firm can put its name to.

Reliability research offers the closest analogue to file review. Inter-rater agreement studies — kappa studies — ask whether two raters reach the same verdict on the same subject, which is the same question as two reviewers scoring one suitability report. The tutorial on sample size calculation for inter-rater reliability records a common rule of thumb: 50–100 subjects as a starting point, varying with the study.

Pilot-study guidance offers the small-numbers anchor. One sample size determination paper for pilot studies works through a 20% non-response allowance and lands on 30 respondents as sufficient to assess the reliability of an instrument.

Survey methodology offers the calculation route. A sample size calculator estimates completed responses for a population proportion at 95% confidence under simple random sampling.

ApproachNumber it yieldsSource
Reliability (kappa) rule of thumb50–100 subjects as a starting pointPMC tutorial on sample size calculation
Pilot-study minimum30 respondents, allowing for 20% non-responsePilot study sample size determination paper
Proportion calculationCalculated at 95% confidence for the firm's own populationSurvey sample size calculator
FCA practice'A sample' — no published countFCA payments firms review, September 2026

Table: Sampling approaches and the counts behind them. The numbers come from methodology sources, not FCA rules.

Regarding sample size methods, the citation matters as much as the count. A firm that writes '30 files, per pilot-study guidance' has a method; a firm that writes '30 files' has a habit.

How does a small firm spread the sample across the Duty?

Coverage dimensions decide whether the sample describes the firm or one corner of it. The outcomes under PRIN 2A form the first split: products and services, consumer understanding, price and value, consumer support. New advice files and ongoing service files form the second — TR24/1 para 1.40 sits over the ongoing service side, where charges continue and reviews go overdue. Vulnerability evidence forms the third: FG21/1 expects records of vulnerable customers' needs, and those records live in the same files.

Advisers form the fourth dimension. Regarding coverage, a sample drawn wholly from one adviser's new business describes one adviser's month; supervision asks about the firm. A sole-trader firm has one adviser, so the split runs across product types and outcomes instead.

Volume interacts with coverage. Thirty files spread across four outcomes and every adviser surface patterns that thirty files from one product line hide.

Vulnerability evidence sits in the same files, under the same sampling discipline. (Source: FCA FG21/1)

What makes the count defensible in supervision?

The written chain makes the count defensible. A sampling policy names the method, the population, and the period. Per-file records show what was checked and what was found, with citations to the rule tested — COBS 9.2.1R for suitability, PRIN 2A.3 for consumer understanding. Action logs show what changed after an Amber or Fail finding, and re-reviews show the fix landed.

Regarding supervision, the examiner reconstructs that chain from population to sample to finding to action. The FCA's March 2026 consumer understanding review, as summarised in Aveni's analysis of the four outcomes, found that sales data and the absence of complaints do not prove customers understood anything. The same logic applies to sample counts: a count proves effort, findings prove monitoring.

The outcomes monitoring hub covers what the Duty expects firms to monitor; the sample is where that monitoring becomes visible. TCC Group's piece on whether suitability reviews provide insight or just oversight draws on the FCA's recent review of outcomes monitoring under the Duty and asks the same question of the file population.

Outcomes live in records; the absence of complaints is not a record. (Source: FCA PRIN 2A.3)

About the Author: Nick Thorp is the founder of Proven Duty, built AI compliance tooling for UK advice firms, and writes about what Consumer Duty means at file level.

Frequently asked questions

What does 'defensible' mean for a file review sample?

A defensible sample is one a third party can trace end to end: a named method, a defined population, coverage across outcomes and advisers, enough files for patterns to surface, and a written record of findings and actions. The count alone carries no weight in supervision.

Do the FCA rules set a fixed file review sample size?

No. COBS 9.2.1R sets the suitability standard, PRIN 2A sets the outcomes, and FG22/5 guides firms on the Duty — none publishes a file count for internal review samples. The FCA's own September 2026 review examined a sample of payments firms without announcing a quota.

Which sample size methods can a firm cite?

Three families appear in the methodology literature: reliability studies using 50–100 subjects as a rule of thumb, pilot-study guidance treating 30 as a workable minimum after a 20% non-response allowance, and statistical calculation of a population proportion at 95% confidence. Each number needs its citation named.

How does a small firm spread the sample across the Duty?

Coverage runs across four dimensions: the outcomes under PRIN 2A, new advice versus ongoing service under TR24/1 para 1.40, vulnerability evidence under FG21/1, and individual advisers. A sample drawn from one adviser's new business describes one adviser; supervision asks about the firm.

What makes the count defensible in supervision?

The written chain: a sampling policy naming method and population, per-file findings with rule citations such as COBS 9.2.1R, action logs, and re-reviews. The FCA's March 2026 consumer understanding review found sales data and the absence of complaints prove nothing — counts prove effort, findings prove monitoring.

Guidance based on published FCA material. This article is not regulatory advice.

See what your files are missing

Run one suitability report through the AI file review and get a Pass / Amber / Fail score against Consumer Duty checks.

Try the free file review