A question sits underneath almost every compliance program that screens counterparties — vendors, customers, distributors, end-users, suppliers: how do we know the screening works? The instinct is to point at the program’s output. We screen every transaction. Our lists are current. Our alerts are cleared within a day. Those are real numbers and they are worth having. But none of them answers the question that was asked, and the reason is structural.
What a Screening Program Actually Learns From
A screening program is trained, formally through a model or informally through analyst experience and rule changes, on its own history. The counterparties that generated alerts were investigated. Those investigations produced outcomes, and the outcomes refined the rules. The counterparties that cleared were cleared without inquiry and produced no outcome at all.
That asymmetry is the whole problem. The program has evidence only about the population it already suspected. About the population it cleared — which is nearly all of it — it has nothing, because it never asked. Every performance number the program reports is therefore conditional on its own past decisions: the hit rate, the false-positive rate, the clearance accuracy all describe the slice of the world the program chose to look at.
The practical consequence is that a clean record on a class of counterparty is not evidence the class is clean. It is evidence the program never examined it. A profile that has never triggered an alert looks safe for exactly the same reason an unopened box looks empty. This is why compliance programs are repeatedly surprised by enforcement actions involving counterparties they had screened and cleared many times over.
The clearance was never a finding. It was the absence of a question.
A program cannot state its own failure rate if it has only ever examined the counterparties it flagged. That number — the rate of disqualifying conduct among the counterparties you cleared — is the one an examiner, an opposing expert, or a whistleblower will eventually put to you, and process metrics cannot produce it.
The Capability Nobody Measures
There is a second structural limit, and it sits upstream of everything the screening rules do. A screen is only as good as its ability to recognize that two nominally distinct counterparties are the same operation.
Evasion networks, sanctions front companies, and diversion schemes change corporate identity far faster than they change the substrate underneath it — the people, the physical premises, the freight forwarders, the banking relationships, the registered agents. A program that treats each new corporate name as a new and unknown entity discards the single most stable signal available to it, and re-clears an operation it has already seen under a different name.
Every screening program performs this entity resolution, whether it knows it or not. Almost none measure how well. The resolution quality is invisible to program management and to auditors alike, which means the binding constraint on the program’s accuracy is a number nobody reports. A program that cannot say what fraction of its screened counterparties it has resolved to a persistent identity does not know how it performs against the counterparties most likely to be hiding.
Ownership-based rules — the OFAC 50 percent rule, foreign-ownership-and-control determinations, research-security affiliation checks — are graph problems. A program running them without resolving entities is performing the analysis it is able to perform, not the analysis the rule requires.
The Attributes You Trust Most Are the Ones an Adversary Can Change for Free
Screening weight tends to concentrate on the attributes that are cheapest for a sophisticated counterparty to alter.
A corporate name changes in days for a nominal fee. A registered address changes as fast. A declared end-use or end-user statement costs nothing to write. A self-classification, a declared value, a self-attested compliance score — all are supplied by the counterparty and cost nothing to misstate. Against those sit the attributes that are genuinely expensive to fake: an operating history, a verified physical footprint, relationship tenure with banks and brokers and forwarders, a transaction pattern consistent with the stated business. Those mostly go unscreened.
An attribute’s value to a screen is not its predictive power. It is its predictive power discounted by the cost to the counterparty of falsifying it.
A feature that predicts well and is free to change is a liability, not an asset: it produces accuracy against unsophisticated actors, who never bother to change it, and close to none against the counterparties who will — while generating documentation that the program appeared to be working. Denied-party list screening is the pure case. It is highly accurate against entities that have not changed their names, and structurally uninformative against those that have.
A program optimized on historical accuracy alone will load its weight onto the cheap attributes, because the enforcement record it learned from is dominated by the unsophisticated actors who left them unchanged. It ends up optimized against the population least likely to cause a serious loss.
The Fix Is One Inexpensive Design Change
The correction to all three defects is a single addition, and it is cheap.
Sample your own clearances. Take a small fraction of the counterparties your screening cleared — one to three percent is enough at commercial volume — selected at random, independent of any score or alert or analyst judgment, and examine them properly. Stratify the sample by jurisdiction, by commodity or service class, by counterparty tenure, so each part of your book is represented.
This is not a novel technique. It is exactly what the Internal Revenue Service does to estimate the tax gap: a randomly selected audit stratum, chosen independently of the enforcement-selection rules, precisely so it reveals what routine enforcement misses. The design has been in use for decades in a domain with far higher stakes than most compliance programs face.
That single stratum produces the one figure no process metric can: an unbiased estimate of your clearance failure rate, with a confidence interval. It lets you state, in an audit or a courtroom, that among the counterparties your program cleared, the estimated rate of undetected disqualifying conduct is some measured value. Without it you can say nothing at all about the counterparties you cleared — which is nearly all of them. The objection to the sample is never its cost, which is trivial. The objection is that it generates a documented estimate of the program’s own failure rate. That is the argument for it, not against it, and the reason is legal, addressed below.
Around that core, four supporting moves follow directly from the three defects:
- Weight attributes by how hard they are to fake, and cap the weight assigned to cheap-to-change attributes even where they predict well. Accept slightly lower measured accuracy in exchange for accuracy against the counterparties who matter.
- Resolve counterparties into a persistent entity graph, so a re-registered shell is recognized rather than cleared as new, and report the resolution rate as a program metric.
- Let trusted status decay. Any status that lowers scrutiny — approved end-user, certified supplier, an accepted remediation plan, a prior favorable review — should carry an expiry and be suspended by behavioral triggers, not left standing until a formal revocation. The status a program trusts most is the most valuable thing for a bad counterparty to acquire.
- Report effectiveness, not activity: clearance failure rate from the sample; lift over random, meaning how much better the alerts do than chance; and the fraction of your transaction space that has received zero non-random review, which is the leading indicator that illicit activity has moved into a segment your rules do not reach.
Every item on this list is measurable with data the program already holds. What is missing is not capability. It is the decision to measure the outcome instead of the activity.
Why This Is Now a False Claims Act Question
Where a contractor certifies compliance with an export-control, supply-chain, or cybersecurity requirement as a condition of payment, that certification is discoverable, and the screening behind it becomes evidence. This is not a hypothetical exposure. It is where the enforcement has gone.
The Department of Justice’s Civil Cyber-Fraud Initiative, running since 2021, uses the False Claims Act against contractors that misrepresent their cybersecurity posture. The object being certified is a program, not an outcome, which is the point: no breach, no exfiltration, and no attack is required. The false statement is the attestation itself. In fiscal year 2025 the Department reported more than fifty-two million dollars recovered across nine cybersecurity settlements, part of a record year for whistleblower filings.1
The worked cases show the pattern precisely.
A logistics contractor self-reported a perfect cybersecurity score of 110. A later government assessment, run by the Defense Contract Management Agency’s assessment center, scored the same company at negative 170 — a 280-point gap on a single self-attested number, against a possible range that bottoms out at negative 203. The company resolved the allegations for $507,144. No breach was alleged; the covered conduct was billing while knowing the required control set was not implemented.2
Another contractor self-reported 104, near the top of the range. An outside assessment told it the real figure was negative 142, reflecting roughly a fifth of the required controls in place. The company did not correct the reported score until months later, after a federal subpoena. It resolved the matter for $4.6 million.3
A research university’s reported score was premised on an assessment environment that did not correspond to any actual system handling the government’s information. It settled for $875,000. A second university settled for $1.25 million over controls it had represented it would remediate on a schedule it did not keep.4
These are settlements, and the allegations were resolved without any admission or adjudication of liability. But the pattern is unambiguous. A self-attested score, unaccompanied by any measurement of whether the attested controls actually work, is the cheapest-to-fake attribute in the whole system — and it is now the precise object around which enforcement is built.
When the government pays for a compliance program rather than a compliance outcome, the question of whether a misrepresentation was material turns on whether the program the contractor ran was the program the government paid for. A score with nothing behind it documents a program without evidencing one. The distinction is exactly what the random-sample audit is built to establish.
Measurement Is the Defensible Posture. Non-Measurement Is Not.
There is a tempting misreading of everything above, and it needs to be met directly, because it is the reading a hostile party would prefer you adopt. The misreading runs: if measuring my own failure rate creates a record of what I knew, it is safer not to measure. Under the False Claims Act’s scienter standard — which reaches actual knowledge, deliberate ignorance, and reckless disregard5 — that reasoning is not just wrong, it is dangerous.
Three reasons it is wrong.
First, the exposure that has actually cost companies money is the false certification, not the discovery of a residual risk. A program that measures its failure rate and acts on it — tightening screening where the sample shows leakage — converts the measurement into evidence of diligence. The adverse posture attaches to measuring and then doing nothing, which is a governance failure independent of any statute.
Second, materiality is a partly separate question, and disclosure builds a defense on it. Where the government continued to pay with knowledge of a disclosed weakness, that continued payment is evidence the requirement was not material — a defense at least one contractor has already argued in live litigation, on the theory that the agency never treated the certification as the essence of the bargain and kept paying after learning of the alleged noncompliance. That defense is available only to a contractor who measured and disclosed. Silence forfeits it.
Third, the alternative to measurement is not safety. It is a certification with no evidentiary basis, which collapses the moment the underlying conduct surfaces, with nothing to fall back on. Choosing not to look does not make the exposure smaller. It makes it undocumented, and a documented decision not to examine a known blind spot is closer to the deliberate-ignorance the statute reaches than to a safe harbor.
The instinct to avoid measurement in order to avoid knowledge produces the worst available position: the exposure remains, the diligence defense is gone, and the materiality defense is gone with it. Measurement is the posture that can be defended. Non-measurement is the one that cannot.
Three Questions That Settle It for Your Program
None of these requires counsel to answer, though the answers may send you to counsel.
One. Can we state our clearance failure rate? Not our alert volume or our clearance speed — the estimated rate of disqualifying conduct among the counterparties we cleared. If the honest answer is that there is no way to know, the program measures activity, not effectiveness, and it cannot substantiate a reasonable-care representation on the merits.
Two. Would our screening survive a counterparty that changed its name last quarter? If a re-registered entity clears as new, the program is not resolving entities to a persistent identity, and its record against the counterparties most likely to be concealing something is unknown to us.
Three. What sits behind our own — or our supplier’s — attested compliance score? If the answer is the score itself and nothing that measures whether the controls actually work, then the most enforcement-exposed attribute in the whole program is also its least substantiated.
If those three answers are clear and consistent, the program can defend itself, and that is a real outcome — more common than the enforcement record suggests, because settlements are the only cases anyone publishes. If any one of them conflicts — a certification with nothing behind it, a screen that clears renamed shells, a program that cannot state what it missed — that conflict is the finding. It is far better identified before an examiner, an opposing expert, or a whistleblower identifies it.
When This Stops Being a Design Question
Whether a particular certification was accurate when it was made, whether a posted score can be substantiated, and what to do about a representation already submitted that may have been wrong are legal determinations, and the last of them belongs with counsel immediately. What is a design question, and what this assessment addresses, is whether a screening program is built to detect the conduct it exists to catch, and how to know. That question has an answer, and the answer is to measure the outcome rather than the activity.
This assessment is not legal advice and is not a substitute for counsel.