Accuracy Is the Wrong Question: Deepfake Detection KPIs for High-Volume Claims
Why accuracy misleads in deepfake detection, how prevalence and error costs set the operating point, where AUROC fits, and what to ask a vendor instead.

Every deepfake detection vendor gets the same first question: “What’s your accuracy?”
It sounds like the right question. It is the one number that cannot tell you whether the detector will work in your claims pipeline. This article explains why, and what to ask instead. One example runs through it: an insurer receiving 100,000 claims a month with photos, documents or audio attached, and a fraud team that can properly review about 3,000 of them.
The Short Version
Accuracy tells you almost nothing. When most claims are genuine, a detector that flags nothing scores in the high nineties. A vendor’s “99%” can mean catching three quarters of the fakes, or all of them.
Two rates run your operation. False alarms set how much work your team gets and how much of it is noise. Misses set how much fraud you pay. One threshold trades them against each other, and the right setting is yours to choose, not the vendor’s.
Your team’s capacity is the real limit. The cheapest way to run a detector sends far more claims to review than most teams can handle. So the detector’s real job is to rank the queue, and the number to watch is confirmed fakes per analyst hour.
The detector you choose decides whether the best setting is affordable. On the same stream, with the same team, a best-in-class detector costs a fraction of what a mid-range one costs to run, and only the best one can be run at its cheapest setting at all.
Once the detector is good enough, your analysts are the limit. Reviewers do not confirm every fake put in front of them. What arrives with the flag matters as much as the flag.
Ask a vendor for four things instead of accuracy: detection rate at the false alarm rate you can afford, alerts per 10,000 items on media like yours, a ranked score with evidence behind it, and how the threshold is set to your volumes and your team.
Why “99% Accurate” Means Almost Nothing
Start with what accuracy actually measures: the share of all decisions that were right.
That sounds sensible until you remember what a claims stream looks like. If 2% of claims carry synthetic media, then 98% are genuine, and a detector that never flags anything gets 98% of its decisions right. It scores 98% accuracy while catching no fraud at all. A detector that catches most of the fakes and raises a few false alarms along the way can score lower.
Accuracy is a statement about the mix of your data, not about the detector. It rewards silence.
The “99%” on a vendor slide is worse than uninformative, because it hides the one thing you need to know: what happens to the fakes. On our stream, a detector at 99% accuracy with 0.5% false alarms could be missing a quarter of the fakes, or almost none of them. The same headline covers both. What separates them is a second and a third number, which is where the useful conversation starts.
The Two Numbers That Actually Run Your Operation
Every detector makes one of four calls on each claim: it passes a genuine claim (right), passes a fake (a miss), flags a genuine claim (a false alarm), or flags a fake (a catch).
Two rates fall out of that, and they are paid for in different currencies.
The false alarm rate is the share of genuine claims that get flagged. It sets your team’s workload and, more importantly, how much of that workload is noise. It is paid in analyst hours and in honest customers waiting for their money.
The miss rate is the share of fakes that get through. It sets the fraud you pay. It is paid in claims money.
Here is the part most buyers miss. Those two rates are not fixed properties of a detector. Every detector produces a score, and the flag is simply that score compared with a threshold. Move the threshold up and you get fewer false alarms and more misses. Move it down and the reverse. A detector is a whole range of possible settings, and the choice of setting belongs to you, because you are the one paying for both kinds of mistake.
Two other figures turn up on vendor decks, and neither replaces these two. AUROC compresses a detector’s entire range of settings into a single score, weighting the settings you would never use as heavily as the ones you would; a detector with an impressive-sounding AUROC can still miss one fake in six at the false alarm rate you can afford. EER is the setting where the false alarm rate and the miss rate happen to be equal, which is a fair way to line detectors up against each other and a setting no claims team would ever run, because it means flagging roughly one genuine claim in twenty. Even ASVspoof, the main academic benchmark for audio deepfakes, now scores detectors on a cost-weighted measure rather than on accuracy or EER. Treat both as tie-breakers between detectors, never as a description of what you will get.
What Your Analysts Actually See
Analysts never see accuracy. They see a queue, and the only thing that matters to them is how much of it is real.
That share, precision, is driven by a simple relationship: at any given detection rate, it depends on the false alarm rate set against how common fakes are. If one claim in fifty is fake and the detector flags one genuine claim in twenty, the queue is mostly noise, however good the detector is at catching fakes. For most alerts to be real, false alarms have to be rarer than fakes themselves.
On our stream, a detector catching 90% of fakes fills the queue very differently depending on its false alarm rate:
| False alarm rate | Claims in the queue | Share that is real |
|---|---|---|
| 5% | about 6,700 | one in four |
| 1% | about 2,800 | two in three |
| 0.5% | about 2,300 | four in five |
Same fakes caught in every row. The difference between a queue analysts trust and one they learn to ignore is almost entirely the false alarm rate.
And the share of fakes is not one number. It changes by claim type, by channel, by season, and by whatever fraud tactic is circulating this quarter. A setting that gives a clean queue on motor claims can give a queue that is half noise on property claims. The setting has to be chosen per stream and revisited, not shipped as a default.
Putting a Price on Each Mistake
To choose a setting you need two prices: what a miss costs you, and what a review costs you.
A miss is a fraudulent claim paid. A defensible figure is the average value of the fraudulent claims you already catch; call it $4,000. A review is analyst time, whether the claim turns out to be genuine or fake; call it $40. That makes a miss worth a hundred reviews, and it leads to a rule of thumb worth remembering: it is worth reviewing any claim with more than a 1% chance of being fake.
Follow that rule and something surprising happens. The cheapest way to run a detector is far more aggressive than anyone expects. For a good detector on our stream, the cheapest setting flags about one genuine claim in thirteen and sends close to 9,400 claims a month to review. The setting with the best accuracy, the one the vendor ships by default, turns out to be the most expensive of all, because it pays out around $1.9m a month in missed fraud to keep the queue short.
| Setting | Claims reviewed | Fakes missed | Total cost per month |
|---|---|---|---|
| Vendor default: 0.5% false alarms, 99.0% accuracy | about 2,000 | about 475 | about $2.0m |
| 3% false alarms | about 4,800 | about 160 | about $830,000 |
| Cheapest: about 7.6% false alarms | about 9,400 | about 60 | about $630,000 |
Few fraud teams can review 9,400 claims a month. That is the finding that changes the whole problem.
Your Team’s Capacity Is the Real Limit
If the team can clear 3,000 claims a month, then 3,000 is the setting, whatever the cost table says. For our good detector that is about 1.3% false alarms: 86% of fakes caught, 57% of the queue real, and around $720,000 a month cheaper than the vendor’s default.
Once capacity is the limit, the question the detector answers changes. It is no longer “flag or pass”. It is “which 3,000 claims go to the top”.
A detector that only says yes or no cannot answer that. At a loose setting it hands the team thousands of flags with no way to order them, and the fakes are buried among false alarms; the real detection rate becomes whatever fraction of the pile gets looked at before the month ends. A detector that returns a score lets the team work the highest-risk claims first, and the measure that matters becomes confirmed fakes per analyst hour. On our stream that is about 1.7 an hour for the good detector at its capacity setting, against one every sixteen hours for reviewing claims at random.
Capacity also has a price you can put a number on. Every extra block of reviews catches more fakes, and the fraud avoided is usually several times the cost of the reviews. That is the business case for more reviewers, for automating the low-risk tail so analysts see only the top of the queue, or both.
The Detector You Choose Matters More Than the Threshold

Three illustrative detector curves. The shaded strip is the only part of the chart a claims team ever operates in. The markers are each curve’s cheapest setting. Only the best-in-class marker sits inside the strip.
It would be easy to conclude from all this that the detector barely matters and the setting does all the work. The opposite is true. The setting picks a point on the curve. The detector picks the curve. And the gap between curves is bigger than most buyers expect.
Public benchmarks of audio deepfake detectors show the spread. Typical detectors, of the kind built into general-purpose platforms, flag around 2.5% of genuine items while missing around 30% of fakes. Mid-table detectors sit near 5% on both. The best sit near half a percent on both. Take those three levels as illustrative curves and put them on our stream, with the same 3,000-review team:
| Detector | Fakes caught | Cost per month | Confirmed fakes per analyst hour |
|---|---|---|---|
| None: review 3,000 claims at random | 3% | about $7.9m | 0.06 |
| Typical | 65% | about $3.0m | 1.3 |
| Good | 86% | about $1.3m | 1.7 |
| Best-in-class | over 99% | about $135,000 | 2.2 |
Doing nothing at all costs $8m a month. Any detector changes that; the rest of the table is the case for choosing the right one. The good detector costs roughly nine times as much to run as the best-in-class one, the typical detector around twenty times, and only the best-in-class curve can be run at its cheapest setting inside the team’s capacity. The other two have to be held back to a setting that misses fakes they could have caught.
Two honest caveats. Every vendor’s real-world curve sits below their benchmark curve (see the lab-to-production gap), so measure on your own media. And when a team cannot review even the fakes the best detector finds, the gap between detectors narrows, and more capacity is worth as much as a better detector.
Try It With Your Own Numbers
The calculator below runs the same arithmetic as the tables above. Put in your volume, your team’s capacity and your costs, pick a detector level or type in a benchmark result, and it returns the setting where your capacity binds, what it costs, and what more review capacity would be worth.
- Flags
- -
- Fakes flagged by detector
- -
- Fakes confirmed by analysts
- -
- Fakes paid
- -
- Real alerts in queue
- -
- Confirmed fakes per analyst hour
- -
Analysts Miss Things Too
Everything so far assumes that a fake put in front of an analyst is always confirmed. It is not. Reviewers confirm some share of the real fakes they see and wave the rest through, and that share is rarely measured. Eighty percent is a reasonable planning figure. Time pressure, a noisy queue and media with no visible clue all push it lower.
This changes where the value in the system sits. With a best-in-class detector, almost no fakes get past the detector, yet at an 80% confirmation rate hundreds a month still get paid, nearly all of them flagged and then cleared by a person. Once the detector is good enough, the analyst is the bottleneck, and the job shifts from finding fakes to helping analysts confirm them: evidence they can see, the part of the image or audio that was manipulated, and a consistent way to record the call. Automate the prioritisation and the evidence collection; keep adverse decisions on a claim with a governed human process.
A noisy queue makes this worse. Reviewers who see mostly false alarms learn to clear things. A cleaner queue is not just cheaper to review; it is reviewed better.
What To Do on Monday
Estimate how common fakes are in each stream, separately for each claim type and channel you plan to screen.
Price a miss and a review. The average fraudulent claim you already catch, and loaded analyst time per review.
Build the cost table on your own media, not on a clean benchmark.
Apply your capacity. Find where the queue meets what the team can clear, work the queue by score, and track confirmed fakes per analyst hour.
Re-check quarterly. The share of fakes and the tactics behind them drift. The setting is a dial, not a property of the detector.
What to ask a vendor instead of accuracy
- Alerts per 10,000 items at the proposed setting, on media that looks like yours.
- Detection rate at the false alarm rate you can afford.
- Whether alerts come with a ranked score and the evidence behind it.
- How the threshold is set against your volume and review capacity, and how often it is revisited.
A vendor who can answer those four has evaluated the detector the way you will use it. One who can only quote accuracy has not.
How deetech™ Approaches This
A score, not just a flag. Every analysed item returns a risk score alongside the yes/no result, so a claims team can rank its queue and work from the top.
Settings chosen for your stream. During a proof-of-concept the threshold is set against your share of fakes, your costs and your review capacity, and revisited as the stream changes.
Measured where it will run. Performance is evaluated on compressed, re-encoded, variable-quality media across images, documents, audio and video.
Evidence with the verdict. A flag arrives with the findings behind it, so the analyst is confirming a documented case rather than second-guessing a score.
Free operating-point assessment
Send us four numbers: your monthly volume of claims with media, how many your team can properly review, your average fraudulent claim value, and what a review costs you. We return the cost table for your stream, the setting where your capacity binds, and what each extra block of review capacity is worth in fraud not paid. No cost, no commitment.
The Bottom Line
Accuracy measures how a detector treats the majority, and in claims screening the majority is genuine. It rewards silence, and the setting with the best accuracy is often the most expensive way to run a detector.
What matters is the false alarm rate, because it sets the size and quality of the queue; the miss rate, because it sets the fraud paid; and the review capacity that decides where between them you can afford to sit. The setting picks the point. The detector picks the curve. And once the curve is good enough, the analyst confirming the flag becomes the limit.
Ask for the detection rate at the false alarm rate you can afford. Ask for alerts per 10,000. Ask for the score and the evidence. Do not buy on accuracy.
Related Reading
- The Lab-to-Production Accuracy Gap
- Insurance Fraud ROI Calculator
- The True Cost of Deepfake Fraud in Insurance
- Why Forensic-Grade Detection Is Not Enough
The worked example uses 100,000 claims a month, a 2% share of synthetic media, 3,000 reviews of capacity, $4,000 per missed fraudulent claim and $40 per review. Costs are computed from the underlying model and rounded only for display. The three detector levels are illustrative curves, each an equal-variance binormal ROC drawn through a published benchmark operating point rather than a measurement of any vendor. Substitute your own figures before drawing conclusions for your portfolio.