CCoffeez for Closers
← All frameworks
InnovationPavan Agarwal

Explainable High-Stakes AI Validation

Make every automated decision traceable, testable, and accountable

Difficulty
Expert
Time to result
~months to results
Steps
6
Confidence
96%

This framework treats explainability and testing as part of the AI architecture rather than as compliance added later. First, define the consequential decision the system will make and require it to produce an understandable account of why it reached that result. Preserve reasoning logs so operators can inspect and audit behaviour. Build a reusable test set from real historical cases, remove personally identifiable information, and alter selected inputs to cover more scenarios. Run tens of thousands of cases repeatedly to detect incorrect or unstable behaviour. Finally, keep accountability with the operator: the model cannot be used as an excuse, and the organisation should bear the consequences of an erroneous approval. The output is an automated decision that can be traced, tested, and defended.

Origin

Pavan Agarwal says Sun West built Angel AI from scratch for mortgage and financial decisions so it could output its reasoning, support internal testing, and satisfy explainability expectations.

Core principles

  • 01An AI decision must include an understandable reason
  • 02Decision logs make model behaviour auditable
  • 03Real historical cases provide grounded test scenarios
  • 04Continuous testing catches behaviour that drifts off course
  • 05The operator remains accountable for automated decisions

How to run it

  1. 1

    Define the consequential decision

    Specify exactly what the system is allowed to approve, reject, flag, or request. Treat the organisation, not the model, as accountable for the result.

    Pro tip Start with one bounded decision whose inputs and acceptable outcomes are already understood.

    Watch out Do not use the AI itself as an excuse for a decision you cannot defend.

  2. 2

    Generate an explanation

    Make the system return a clear reason with each decision, such as the particular evidence behind an approval or an unsatisfactory credit assessment.

    Pro tip Write explanations for the affected customer or reviewer, not for the model engineer.

    Watch out Weights or internal model values are not a meaningful explanation to a customer or regulator.

  3. 3

    Record the reasoning

    Automatically log the reasoning produced during the decision so the team can reconstruct and audit what happened.

    Pro tip Keep the decision, explanation, relevant inputs, and model version together.

    Watch out An explanation that is not retained cannot support a later audit.

  4. 4

    Build grounded test cases

    Turn real historical cases into test data, remove personally identifiable information, and vary selected facts to create additional scenarios.

    Pro tip Include both ordinary cases and exceptions that previously required human correction.

    Watch out Do not expose customer personal information in the test set.

  5. 5

    Test continuously

    Run a large automated scenario suite repeatedly and inspect failures to ensure the system has not gone off the rails.

    Pro tip Convert every corrected error into a lasting regression case.

    Watch out A one-time launch test cannot establish that later behaviour remains correct.

  6. 6

    Back the decision

    Set an explicit policy for what the organisation will do when the system makes a wrong decision. This turns accountability from a statement into an operating commitment.

    Pro tip Choose a remedy proportionate to the consequence of the automated decision.

    Watch out Do not offer a guarantee that the organisation cannot financially or operationally honour.

In the wild

Angel AI mortgage approval

Angel AI reviews a mortgage file, makes an approval decision, and provides the reasoning behind it. Sun West tests the system with anonymised and altered versions of real loan cases, while its warranty means the company funds and balance-sheets a loan if Angel approved it incorrectly.

Pavan reports more than 200,000 transactions since 2018 and about four loans that Sun West had to balance-sheet under the warranty since 2019.

Illustrative insurance eligibility check

An insurer limits automation to one eligibility decision. Each result names the policy facts used, stores an audit log, and is tested against anonymised historical claims plus altered edge cases. Every corrected mistake becomes a regression test, and a human-owned remedy applies when the system is wrong.

Reviewers can reconstruct the decision and detect regressions before they affect more applicants.

Common mistakes

Treating the model as the explanation

Pointing to a neural network's weights does not explain a consequential result in terms a customer or regulator can understand.

Testing only before launch

A model can make mistakes or change over time, so test scenarios need to run continuously rather than once.

Outsourcing accountability to AI

The organisation cannot defend a poor outcome by saying that the AI made the decision.

Is it for you?

Best for

It is best for teams automating regulated financial, credit, insurance, or other high-stakes decisions.

Not ideal for

It is not ideal for low-consequence creative tasks where formal decision explanations and extensive validation would add little value.

From the transcript

you can use the AI to make a decision but then you also have to be able to explain how the AI made that decision

Pavan Agarwal · 33:00

we had to build it from scratch to be able to spit out the the reasoning for his decisions along the way

Pavan Agarwal · 34:30

we're constantly testing the AI to make sure that it doesn't go off the rails

Pavan Agarwal · 36:30

From the episode

A.I. in Finances and Mortgages ft. Pavan Agarwal

Pavan Agarwal