Fraud scoring, in the words the model uses

A fraud score is not an accusation and not evidence. Seven terms decide what it is, what it can support, and why the claimant it is wrong about pays a cost nobody in the insurer measures.

Updated September 13, 2026 Advanced

Almost every argument about fraud models is really an argument about a word. The terms below are the ones a claims professional meets in a scoring conversation, defined as they are generally used in this field rather than as any vendor defines them. The subject is easier to hold once they are fixed.

Score. A number attached to a claim, meant to order claims by how much they resemble claims previously found or believed to be fraudulent. It is an ordering device. It is not a probability that this claimant is dishonest, even when it is expressed as one, and it is not evidence of anything.

Feature. An input the model uses. Some are about the claim (time between policy inception and loss, time of day, delay in reporting, injury claimed without vehicle damage, prior claims on the policy). Some are about the parties (an address, a vehicle, a repairer, a medical provider, a representative). Some are derived from text or images. A feature is not a reason; it is a correlate.

Link or network analysis. The part of this field that does something a handler cannot do. Rather than scoring one claim in isolation, it builds a graph across many claims — vehicles, addresses, phone numbers, clinics, repairers, representatives, bank details — and looks for structures: the same vehicle appearing in unrelated collisions, a cluster of claimants who share a treating provider, a repairer that recurs across claims with no other connection. Organised fraud has a shape in that graph, and the shape is visible only from above.

Base rate. The proportion of claims in the population that are actually fraudulent. It is low, it is not knowable precisely, and it governs everything that follows.

False positive. A claim the model flags that turns out to be honest. Its counterpart, the false negative, is a fraudulent claim the model passes. Every threshold choice trades one against the other, and the two errors land on different people.

Referral threshold. The score above which a claim goes to a special investigation function rather than being handled normally. Like every threshold in automated claims handling, it is set by people, adjustable in an afternoon, and invisible in the file.

Explainability. The property of being able to say why this claim received this score, in terms a person can check. It is not the same as interpretability of the model in general, and neither is the same as the duty to give a claimant reasons for a decision, which is a legal obligation about outcomes rather than a technical property of a system.

The subject, in those terms

Fraud models do one thing well, and it deserves to be stated before the objections. A single handler sees one claim. Repeated and organised fraud is a pattern across files that were handled by different people, in different months, in different places, and the graph is where that pattern becomes visible at all. An insurer that can identify a cast of parties recurring across unrelated losses is detecting something real, and it is detecting the kind of fraud that costs the most and that honest policyholders pay for. On that ground the case is strong.

The base rate is what complicates everything else. Because fraudulent claims are a small share of all claims, the flagged population is dominated by the majority class even under a model that performs well. This is arithmetic rather than a criticism of any particular system: apply a good test to a rare condition and most positives are still false. A model can be better or worse — it cannot escape this — and the only lever available is the threshold, which trades the two error types against each other rather than reducing both.

So the question is never “is the model accurate”. It is what happens to the claims in the flagged population that do not belong there, and the answer to that question is set by process design rather than by data science.

The labels are the deeper problem

A fraud model is fitted on historical claims labelled fraudulent or not. Where do the labels come from? From investigations that were opened, pursued and concluded — which is to say, from the judgements of the people who handled claims before the model existed.

That has a consequence worth stating plainly: the model learns the pattern of what past investigators pursued. If investigations concentrated in particular postcodes, on particular repairers, on claimants with particular names or particular ways of writing a claim, the model reproduces that concentration and reports it as a finding. The claims that were defrauded and never investigated are labelled clean, because nobody found them, so the model learns that they look normal.

The result is a system that can be measured as accurate against its own history and still be wrong about the world in a stable, directional way. Where a feature correlates with a protected characteristic without naming it — a postcode, a vehicle age, a preferred language, a hire arrangement — the model can produce a differential outcome that nobody chose and nobody can see from inside the score. That is the unfair-discrimination exposure regulators have started to describe, and it does not require anyone to have had a discriminatory intent.

What a false positive costs, and who counts it

An honest claimant whose claim is flagged does not experience a score. They experience a claim that stops. The repair authority does not arrive. Someone asks for documents that were not asked for before: bank statements, phone records, a recorded interview, proof of ownership, an explanation of why the claim was reported three days later rather than on the day. The tone of the correspondence changes. Where a hire vehicle was in place, it may end. Where there was an injury, treatment decisions get made around a claim that is not paying.

None of that appears in an insurer’s reporting. Fraud saved is counted, celebrated and used to justify the programme. The delay imposed on the claims that were released without a finding is counted nowhere, by no one, and it is borne entirely by people outside the company — which is exactly the structure that makes it easy to expand.

There is a second cost that is harder to see and more serious. Some people, faced with an investigation, withdraw a claim they were entitled to. A withdrawn claim is recorded as a saving. It is indistinguishable in the numbers from a fraud deterred, and there is no mechanism inside the insurer that tells the two apart.

Reasons, and what a score cannot supply

When a claim is declined, or reduced, or handled on the basis of suspicion, the claimant is owed an account of why in most systems we are aware of, and the specific obligation is a jurisdiction question. What is not a jurisdiction question is whether a score can serve as that account. It cannot. A score says this claim resembles other claims; a reason says what about this claim is wrong. The gap between those is where the complaint, the regulatory enquiry and the bad-faith allegation live.

The discipline that follows is unglamorous and it is worth adopting before anyone requires it. Every referral should carry, in the file, the claim-specific facts that a person verified — not the score, and not a paraphrase of the score. Every release without a finding should be recorded as such, with the elapsed time. And the model’s version, the inputs it saw and the threshold in force should be retrievable for a claim months later, because a claim that becomes contentious becomes contentious long after the score was produced.

An insurer that can do those three things can defend its programme to a regulator, to a court, and to a claimant. One that cannot has a detection capability and no account of how it was used.

Fraud detected is a number every claims committee sees. Honest claimants delayed is a number nobody produces, so the trade-off between them keeps being decided with one side of the ledger blank, and the side that is blank is the one that belongs to somebody else.

Ariski's take

Fraud detection is the application in claims where we are most willing to say the models earn their place: organised, repeated, networked fraud is a pattern problem across many files, and a pattern across many files is precisely what a person handling one file cannot see. The objection we hold is not to the scoring. It is to an asymmetry in how the results are read. An insurer counts what the model caught and carries no number at all for the honest claimant whose payment was delayed while a unit looked at them. Augmentation here means the model finds candidates and a person decides what the candidate deserves — and it stops being augmentation the moment a score can start a process that a claimant cannot see, answer, or appeal.

Rules in your jurisdiction

Deadlines, fault rules and minimum coverage differ by state and country. Pick yours to see the rules that apply to this topic.

Select a jurisdiction to see its rules.

Frequently asked questions

Can we tell a claimant their claim was flagged by a model?

Say what is happening and why, without handing over the detection method — those are different disclosures and conflating them is how insurers end up saying nothing. A claimant can be told that their claim requires further investigation, what specifically is being verified, what is needed from them, and when they will hear back. What cannot usefully be given is a feature list, because publishing the features of a fraud model degrades it for the cases it is meant to catch. Where a regulator requires reasons for an adverse decision, the reason has to be about the claim rather than about the score: a score is not a reason, it is a routing outcome, and a decision that cannot be explained without referring to the score is a decision the insurer is not yet in a position to make. What your own rules require by way of reasons, and whether anything in them speaks to automated decision-making or profiling, is not something this answer states for any jurisdiction; the regulator for yours is named in the rules below.

Our referral rate went up and our confirmed-fraud rate did not. What does that mean?

Most likely that the threshold moved into a region where the added candidates are mostly honest, which is what happens at the tail of any scoring distribution. It can also mean the referrals are arriving in a form the unit cannot work with, or that confirmation is being defined more strictly than before, so check those before concluding anything about the model. The number worth adding to the report is the one nobody keeps: how long a claim spends in investigation before it is released without a finding, and how many contacts the claimant went through in that time. Without it you are optimising one side of a trade-off and treating the other as free.

Is a score admissible or disclosable if the claim ends up in litigation?

Treat it as a document that will be read by someone hostile, because that is the safe assumption. Whether it is in fact obtainable by a claimant, and what privilege attaches to it, is a question of procedural law that this answer does not settle for any jurisdiction. Two practical consequences follow regardless of the answer. First, whatever the handler wrote next to the score becomes part of the story of how the claim was handled, so a note recording that a score prompted a check reads very differently from a note recording that a score prompted a conclusion. Second, an insurer that cannot reconstruct why a particular claim scored as it did is in a poor position months later, which makes retention of the inputs and the model version a litigation question rather than only a governance one.