Proven Duty

AI Transparency Documentation

Last updated: August 2026

This document explains how Proven Duty uses artificial intelligence, what data the AI receives, how it reaches its conclusions, and what its limitations are. It is intended to be shared with FCA supervisors when asked "how does your AI tool work?"

Maintained by Nick Thorp, founder — every rubric or process change on this page requires human review before it ships.

1. What the AI Does

Proven Duty uses AI (Google Gemini) in four features:

  • File Review Automation: Analyses uploaded suitability report PDFs against a Consumer Duty compliance rubric, scoring each check area as Pass, Amber, or Fail with explanatory annotations.
  • Vulnerability Signal Detection: Scans uploaded meeting notes and correspondence for language patterns associated with the FCA's FG21/1 vulnerability categories (health change, bereavement, financial stress, comprehension difficulty, life transition, emotional distress, coercion).
  • Regulatory Digest Generation: Summarises recent FCA publications relevant to the firm's advice types, tagging each entry as informational, action required, or urgent.
  • Paraplanner Assist: Drafts suitability report sections from a client fact-find and product recommendation, covering circumstances, rationale, alternatives, vulnerability considerations, costs, and risks.

2. What Data the AI Receives

The following data is sent to Google Gemini for processing:

  • File reviews: Extracted text from uploaded PDF suitability reports. No images, no file metadata.
  • Vulnerability scans: Extracted text from uploaded meeting notes or correspondence PDFs.
  • Regulatory digests: Text content from FCA publications (publicly available).
  • Suitability drafts: Client fact-find data (name, financial circumstances, risk profile) and product recommendation details entered by the adviser.

Client data sent to Gemini is processed under Google's enterprise data processing terms. It is not used for model training. The Gemini API is configured to prevent data retention for training purposes.

3. How the AI Reaches Its Conclusions

The AI uses structured prompting with the following approach:

  • System instructions: Each feature has a defined system prompt that sets the AI's role, constraints, and output format.
  • Structured output: The AI is required to return JSON output conforming to a strict Zod schema. Invalid responses are rejected and the request is retried (up to 3 times).
  • Rubric-based scoring (file reviews): The AI evaluates the document against a versioned checklist of Consumer Duty compliance checks. Each check carries explicit written pass/amber/fail criteria and its FCA Handbook rule references, which are supplied to the model for every review. All rule citations are audited against the FCA Handbook.
  • Category-based detection (vulnerability): The AI classifies detected vulnerability signals into FCA FG21/1categories with severity ratings (low, medium, high).
  • Section-based drafting (paraplanner): The AI generates specific report sections with clear structure, which the adviser can edit before sign-off.

The AI model used is Google Gemini 3.5 Flash-Lite, with automatic fallback to OpenAI GPT-4o-mini if Gemini is unavailable. We do not fine-tune the model on customer data.

4. How the Compliance Rubric is Maintained

The compliance rubric that file reviews are scored against is a controlled document with a full version history:

  • Rule citations audited: Every rule reference in the rubric (COBS, PRIN 2A, FCA guidance papers) has been verified against the FCA Handbook, with the evidence trail retained.
  • Coverage checked against regulator evidence: The rubric's checks are mapped against the FCA's own published file-review findings (including its Retirement Income Advice thematic review, TR24/1) and Financial Ombudsman suitability guidance, so the checks reflect what the regulator actually looks for.
  • Versioned changes only: Rubric changes are made through versioned database migrations — never silent edits — so any score can be traced to the exact rubric version that produced it.
  • Weekly regulatory monitoring: An automated weekly scan checks new FCA publications against the rubric and alerts us to potential impacts. No rubric change is ever applied automatically — every change requires human review and approval.

Limitation we state plainly: the rubric has not yet been validated by practising compliance consultants. We welcome review by firms' own compliance consultants and regulatory specialists, and feedback directly informs rubric revisions.

5. Known Limitations

The AI can and does produce incorrect assessments. Known limitations include:

  • False negatives: The AI may miss compliance gaps that a human reviewer would identify. This is particularly likely for subtle or context-dependent issues.
  • False positives: The AI may flag issues that are adequately addressed in the document, or identify vulnerability indicators that do not reflect actual vulnerability.
  • Inconsistent scoring: The same document may receive different scores on different analyses. We mitigate this through structured prompting and validation schemas, but inconsistency cannot be fully eliminated.
  • Text extraction errors: PDF text extraction may miss content (especially in scanned documents or those with complex layouts), leading to incomplete AI analysis.
  • Training data bias: The underlying model reflects biases in its training data. While we use structured prompts to reduce bias, the AI may respond differently to documents from different demographics, writing styles, or product types.
  • Suitability draft inaccuracies: Paraplanner Assist drafts may contain reasoning that is not appropriate for the client's specific circumstances, incomplete cost and charges disclosures, alternative consideration sections that do not reflect the full range of products actually considered by the firm, and risk warnings that are generic rather than product-specific. The AI does not have access to the firm's own fair value assessments, product governance arrangements, or panels.
  • Context limitations: The AI does not know the firm's internal policies, previous advice given to the client, or regulatory history. A suitability draft must be verified against the firm's own records and procedures before approval.
  • Recency: The AI model has a training data cutoff and may not reflect the most recent FCA guidance or rule changes.

6. Human Oversight

Every AI output is subject to human review before it has any compliance effect:

  • File reviews: Scores are displayed to the compliance team for review. No automated actions are taken based on scores alone.
  • Vulnerability signals: Each signal requires human confirmation (confirm, dismiss, or action) before any client-facing change.
  • Suitability drafts: Each section must be explicitly reviewed and approved by the named adviser. Sections cannot be marked final without adviser approval. All edits are logged in an audit trail.
  • Regulatory digests: Summaries are informational only. No automated compliance actions result from digest content.

The section-by-section approval workflow is designed to ensure each part of the suitability report has been individually read and assessed by the named adviser. Bulk-approving all sections without individual review is a misuse of the tool and a breach of the Terms of Service.

7. Accuracy Monitoring

We monitor AI accuracy through:

  • Schema validation: Every AI response is validated against a strict Zod schema. Invalid responses are rejected and retried.
  • Output logging: All raw AI responses are stored for review and retrospective analysis.
  • Reference verification: For file reviews, every quote the AI cites as evidence is checked against the source document, and unverifiable quotes are flagged.
  • Regression testing: Scoring changes are tested against a fixed corpus of scored example reports before any rubric, prompt, or model change is released. Changes that shift agreement beyond defined thresholds are blocked.
  • Weekly rubric monitoring: New FCA publications are scanned weekly for potential impact on the compliance rubric (see section 4).
  • Feedback mechanisms: Firms can flag incorrect assessments, which we use to improve prompts and rubrics.

8. Bias Monitoring

We are aware that AI systems can exhibit bias. Our approach to monitoring and mitigating bias includes:

  • Structured prompts that focus on objective criteria (rubric checks, FG21/1 categories) rather than subjective assessments.
  • Scoring criteria defined in the rubric that are applied uniformly regardless of client demographics.
  • Monitoring for statistically significant differences in scoring patterns across demographic groups (where data is available).

We acknowledge that bias monitoring is an ongoing process and we welcome feedback from firms on any perceived bias in AI outputs.

9. Data Handling

Client data sent to Gemini is processed under Google's Cloud Data Processing Addendum. The API is configured so that:

  • Data is not stored by Google beyond the duration of the API request.
  • Data is not used for model training or improvement.
  • Where possible, EU-based API endpoints are used for processing.

10. Incident Response

If the AI produces a materially incorrect assessment:

  • The firm should flag the assessment through the platform or by contacting support.
  • We will investigate the root cause (prompt issue, model limitation, data extraction error).
  • If a systemic issue is identified, we will notify all affected firms and update the prompt, rubric, or processing pipeline as appropriate.
  • Corrected assessments can be re-generated at no cost.

11. Contact

For questions about AI usage: support@proven-duty.co.uk