Decisions in minutes · auditable · explainable Straight-through processing as the default AI platform for insurance LGPD-compliant Decisions in minutes · auditable · explainable Straight-through processing as the default
Back to Insights & News
· Artigo

Shadow Mode Underwriting: Test AI Without Risking a Book

Shadow mode underwriting is running an AI model on live submissions without letting it decide anything. The model sees the same risks the underwriters see, its answer is logged and compared rather than acted on, and the record becomes the evidence a regulator or a reinsurer will actually accept.

Shadow Mode Underwriting: Test AI Without Risking a Book

Shadow mode underwriting is running an AI model on live submissions without letting it decide anything. The model sees the same risks the underwriters see, produces its own answer, and that answer is logged and compared rather than acted on. It is the cheapest way to find out whether a model works on your book, and it is the evidence a regulator, a reinsurer, or a nervous chief underwriting officer will actually accept.

What is shadow mode underwriting?

Shadow mode underwriting is a controlled evaluation period in which an AI or rules-based underwriting system scores real, live submissions in parallel with the existing human process, while having no authority to quote, decline, price, or bind anything.

The name comes from software engineering, where a new service runs alongside the old one on production traffic, its outputs discarded, purely to see whether it would have behaved correctly. In underwriting the mechanic is identical. Every submission that reaches the team also reaches the model. The underwriter works normally and never sees the model output while deciding. Afterwards, the two answers are stored side by side.

The critical constraint is the one people get wrong: the model must not be visible to the decision-maker during the decision. The moment an underwriter can see the score before committing, the comparison is contaminated, because you are no longer measuring the model, you are measuring a human being anchored to it.

Why shadow mode exists

Three problems make a straight go-live irresponsible on an underwriting book.

The first is that model performance on a vendor's benchmark says nothing about performance on your portfolio. Appetite, broker mix, data quality, and the local market all shift the answer. A model that scores well on a North American commercial property book can be useless on Brazilian cargo.

The second is regulatory. Supervisors increasingly expect insurers to demonstrate that an automated decision was validated before it affected a customer, and to be able to explain any individual outcome. A shadow period generates exactly that evidence: a dated, versioned record of the model's behaviour on real risks before it had any authority. The global picture is covered in AI underwriting regulation in 2026, and the Brazilian view in SUSEP and AI regulation.

The third is organisational. Underwriters are asked to hand judgment to a system. The fastest way to lose that room is to launch a model and let it be wrong in public. The fastest way to win it is to let the underwriters see, for three months, where the model agreed with them and where it did not, before anything is at stake.

How a shadow run works in practice

  1. Freeze the scope. One line of business, one segment, one geography. A shadow run across a whole book measures nothing because the failure modes cancel each other out.
  2. Define the decision being shadowed. Appetite in or out, refer or auto-quote, technical price, risk score band. A model shadowing four different decisions at once produces four weak signals instead of one strong one.
  3. Wire the model to live traffic. Real submissions, real data quality, real gaps. Never a clean historical extract. The mess is the point.
  4. Blind the humans. Model output is written to a store the underwriting team cannot see during the working day.
  5. Log the full record. For every submission: inputs, ruleset and model version, model output, reason codes, human decision, timestamp, and eventual outcome where it becomes known.
  6. Review on a cadence. Weekly for the first month, then fortnightly. Every review looks at disagreements, not at the headline agreement rate.
  7. Track the ones that bound. Agreement is a proxy. Loss experience and conversion on the risks that were actually written are the real signal, and they arrive later.

What to measure

Agreement rate is the number everyone asks for and the least informative one on its own. Four measurements matter more.

  • Disagreement analysis by direction. The model wanted to decline and the human wrote it, or the model wanted to quote and the human declined. These are two completely different business problems and must never be pooled into one percentage.
  • Disagreement analysis by cause. Missing data, a rule the model does not have, a broker relationship the model cannot see, or a genuine model error. Only the last one is a modelling problem. The others are process problems wearing a model costume.
  • Coverage. How often the model returned a usable answer at all. A model that agrees 95% of the time but abstains on a third of submissions has a data pipeline problem, not a good result.
  • Stability. The same risk, resubmitted with trivial differences, should score the same. Instability in shadow is the clearest predictor of trouble in production.

There is a hard rule underneath all of this: a high agreement rate is not proof the model is good, it is only proof the model is not obviously bad. A model that agrees with your underwriters 92% of the time has replicated your existing book, including its existing mistakes. The value is concentrated in the 8%, and the whole point of the review cadence is to understand it.

How long a shadow period should last

Long enough to see enough risks, and long enough for at least some of them to season. In practice:

  • Volume first, not calendar time. A book doing 200 submissions a week reaches statistical usefulness far sooner than one doing 20. Fix a target submission count, not a number of weeks.
  • Cover the seasonality that matters. Renewal peaks, monsoon or hurricane windows, and fiscal year-end all change submission mix. A shadow run that misses the peak has not tested the peak.
  • Six to twelve weeks is the common band for a defined segment with reasonable volume. Under four weeks is a demo. Over six months usually means nobody wants to decide.

Graduating out of shadow

Shadow mode is a phase, not a permanent state, and the exit should be defined before it starts. The usual path has three steps and it is deliberately gradual.

First, advisory mode: the model output becomes visible to the underwriter as a recommendation, with the human still deciding everything. Agreement typically jumps here, which is anchoring, not improvement, so this phase measures adoption rather than accuracy.

Second, limited authority: the model gets to auto-decide a narrow, well-understood band. Usually the cleanest risks that would have been auto-quoted anyway, or clear out-of-appetite declines. Everything else still refers to a human.

Third, expanded authority, band by band, each expansion justified by the data from the previous one.

The exit criteria should be written down at the start and should include a floor on coverage, a ceiling on unexplained disagreement, evidence of stability, and a documented rollback procedure. The wider pilot design around this is set out in how to run an AI underwriting pilot.

What shadow mode does not fix

It does not fix bad data. If submissions arrive as unstructured PDFs and email threads, the model will underperform in shadow for reasons that have nothing to do with the model. Underwriters already lose around 40% of their time to administrative work rather than risk selection, according to Deloitte, and a shadow run on top of that mess will mostly measure the mess. Fix ingestion and extraction first.

It does not fix an undefined appetite. If two senior underwriters disagree about whether a class is in appetite, the model cannot be right, because there is no right answer to compare against.

And it does not, by itself, satisfy an auditor. What satisfies an auditor is the record: versioned rulesets, reason codes on every output, and the ability to replay a decision months later. That requirement is the same one that governs live underwriting and is covered in how to audit AI underwriting decisions for compliance and making underwriting decisions auditable.

The architectural implication is that shadow mode is only cheap if the model can be attached to live submission flow without touching the policy administration system. If evaluating a model requires a core integration project, nobody will evaluate more than one model, which is precisely how insurers end up committing to the first vendor they saw. An external layer that reads live traffic and writes to its own log, leaving the system of record untouched, is what makes shadow mode a routine exercise rather than a programme.

WIR Innovation is an external AI layer for insurers and MGAs that automates submission intake, quotation, and underwriting decisioning without replacing the core system, with ML calibrated to the insurer's own risk appetite and underwriting manual, and every decision explainable with a full audit trail. It has applied this pattern in a proof of concept with a global insurer in the Transport line.

Frequently asked questions

What is shadow mode underwriting?

Shadow mode underwriting is a controlled evaluation period in which an AI or rules-based underwriting system scores real, live submissions in parallel with the existing human process while having no authority to quote, decline, price, or bind anything. The model's output is logged and compared against the underwriter's decision afterwards. The underwriter must not see the model output while deciding, otherwise the comparison measures anchoring rather than model quality.

How long should a shadow mode period last?

Set a target submission volume rather than a number of weeks, because a book doing 200 submissions a week reaches statistical usefulness far sooner than one doing 20. The run should also cover the seasonality that matters, such as renewal peaks or catastrophe windows. For a defined segment with reasonable volume, six to twelve weeks is the common band. Under four weeks is a demonstration rather than a test.

What should you measure during a shadow run?

Disagreement analysed by direction, since the model wanting to decline a risk the human wrote is a different problem from the model wanting to quote a risk the human declined. Disagreement analysed by cause, separating missing data, missing rules, and genuine model error. Coverage, meaning how often the model returned a usable answer at all. And stability, meaning the same risk scores the same way on resubmission. Headline agreement rate on its own is the least informative number.

Is a high agreement rate proof the model is good?

No. A high agreement rate only proves the model is not obviously bad. A model that agrees with your underwriters 92% of the time has largely replicated your existing book, including its existing mistakes. The value sits in the disagreements, which is why the review cadence should focus on understanding the minority of cases where the two answers differ rather than on celebrating the majority where they match.

How do you move from shadow mode to live underwriting?

In three graduated steps with exit criteria defined before the run starts. First advisory mode, where the model output becomes visible as a recommendation while the human still decides. Then limited authority, where the model auto-decides a narrow well-understood band such as the cleanest risks or clear out-of-appetite declines. Then expanded authority band by band, each expansion justified by data from the previous one, with a documented rollback procedure throughout.

What problems does shadow mode not solve?

It does not fix bad data. If submissions arrive as unstructured PDFs and email threads, the shadow run mostly measures the quality of your ingestion, so extraction should be fixed first. It does not fix an undefined appetite, because if senior underwriters disagree on whether a class is in appetite there is no correct answer to compare the model against. And it does not by itself satisfy an auditor, who needs versioned rulesets, reason codes, and replayable decisions.