September 21, 2026

AML False Positives: When Compliance Tooling Blocks Legitimate Flows

Sagar Prasad
Portfolio Manager
AML False Positives — compliance screen stopping illicit and legitimate flows
In This Article
Share
Questions? Speak to our Team

Studies across financial services put as much as 95 percent of AML alerts in the false positive category, and academic work on rule-based crypto monitoring reports rates exceeding that threshold. Elliptic's analysis from April 2026 makes the structural point clearly: crypto compliance programs see higher false positive rates, and not because crypto transactions are inherently more suspicious. Wallet addresses carry no identity information, so without rich customer data a monitoring system has less context and defaults to broader, less precise alerting, while velocity and volume triggers designed for slower banking environments fire constantly against blockchain throughput. The cost is routinely described as wasted investigation hours, which understates it: a compliance function generating ten alerts to find one real risk is not merely inefficient, it is a detection system whose true signals are buried in its own noise.

The Trigger and the Mechanics

The trigger is a threshold set under asymmetric costs. Academic work on algorithmic compliance frames it precisely: false negatives allow illicit activity through, false positives generate investigative burden, and optimal classification depends not only on predictive accuracy but on how misclassification costs enter downstream enforcement decisions. Those costs are not symmetric in practice. A compliance officer who under-blocks faces an enforcement action; one who over-blocks faces an unhappy customer. The system optimizes for the examiner's inspection rather than the customer's access, and thresholds drift conservative because nothing pushes them back.

Two crypto-specific mechanics amplify it. Screening operates on addresses rather than accounts, and addresses are persistent public identifiers — so a flag does not expire with a transaction the way a banking alert does. It attaches to the identifier and is independently consumed by every other institution screening the same chain. And indirect exposure screening looks several hops from a sanctioned address, so legitimate counterparties inherit risk scores from transactions they were never party to and could not have declined.

Where the Losses Land

Four exposures follow. The legitimate customer bears the immediate cost: a treasury movement, payroll run, or settlement leg frozen mid-flow, with no defined timeline for release. The counterparty several hops removed bears an inherited score it never chose. The compliance team bears alert fatigue, and this is where the failure becomes serious — at a ten-to-one ratio, analytical attention is the scarce resource, and a genuine alert competes for it against nine that are not, so the false positive problem produces false negatives. And at the institutional level, the rational response to persistent ambiguity at scale is not better investigation but exit: de-risking an entire customer category is cheaper than adjudicating it, which converts a tooling parameter into financial exclusion.

What to Watch, Real Defenses, and What Does Not Work

The indicators are measurable and most firms already have them. Alert-to-escalation ratio by rule, which identifies exactly which rule is generating noise. Median time-to-clear, since a rising figure means the backlog is growing faster than the team. Repeat flags on the same counterparty, indicating a rule mismatched to a known customer's normal behavior rather than a recurring risk. And customer attrition following screening delays, which is the de-risking signal before anyone calls it that.

Real defenses are configuration rather than procurement. Risk-appropriate rules calibrated per customer segment, so a high-throughput institutional counterparty is not measured against retail velocity thresholds. Behavioral and contextual signals — transaction patterns, wallet history, counterparty type — rather than static amount triggers. And documented threshold governance: who set the parameter, on what basis, and when it was last reviewed.

One fake defense inverts the logic and deserves naming. A low false-positive rate is not the goal and is itself a warning sign — it indicates rules narrow enough that fewer alerts fire, which means legitimate threats are passing undetected. Any firm reporting near-zero false positives has tuned for a comfortable queue rather than for detection. The related failure is treating a vendor's default configuration as a risk decision, when defaults are built for the median client and adopting them unexamined outsources a judgment the institution remains accountable for.

The Playbook, Residual Risk, and the Scale Question

The playbook is threshold governance as a standing discipline: measure alert-to-escalation by rule, tune the worst offenders quarterly with documented rationale, segment rules by customer type, feed identity and attribution data to the screening layer, and keep a written record of every parameter decision — because the defense of a tuned threshold is the documentation behind it, not the number itself. The residual risk is that the asymmetry is permanent: as long as under-blocking carries enforcement risk and over-blocking does not, thresholds drift conservative, and no model improvement removes that incentive.

For compliance tooling to support 10x institutional adoption, three things must become true. Threshold governance must be examinable as a discipline, so a firm can defend a calibrated threshold as readily as a conservative one. Context must reach the screening decision rather than the post-alert review, which separates a system that scores addresses from one that understands counterparties. And remediation must exist: today a flag propagates across institutions while a correction does not, so there is no mechanism for a wrongly-scored address to be cleared and for that clearing to travel. The constructive signal is that the tooling is improving measurably — machine learning deployments report 60 to 70 percent reductions in non-actionable alerts, and ensemble graph models on published datasets achieve high recall of illicit activity at false-positive rates below 1 percent. The gap is no longer primarily technical. It is that the incentive to tune sits with the party that bears none of the cost of over-blocking.

For informational purposes only. Not an offer to buy or sell any security. Available only to accredited investors who meet regulatory requirements.

Recommended blog posts