Referral-Grade Evidence: Audit Trails That Survive CMS and an MFCU
The Pondera failure wrongly suspended 600K+ claimants. Here is what a defensible, human-in-the-loop referral packet contains and how confidence scoring prevents wrongful terminations.
In 2020, as pandemic unemployment claims flooded state systems, an automated fraud-scoring platform flagged more than 600,000 legitimate claimants in California alone, freezing benefits for people who had done nothing wrong [1][2]. The scoring engine did its job as designed. It produced a number. Then the number became a verdict, and hundreds of thousands of families lost income while they fought to prove they were real.
That failure is the reason this article exists. Referral-grade evidence means three things stacked together: confidence-scored detection signals, a documented human review step, and a preserved audit trail that reproduces exactly how the decision was made. A model score alone is a lead, not a verdict. If you cannot show a CMS auditor or a Medicaid Fraud Control Unit prosecutor who reviewed the case, what documents they examined, and which model version scored it, you do not have evidence. You have a liability.
I have built and audited these packets. The difference between a program that survives scrutiny and one that becomes the next headline is not detection accuracy. It is what happens between the score and the adverse action.
600,000 Wrongly Flagged: The Failure That Should Haunt Every Program
The California EDD episode is the case study every program-integrity lead should keep on their desk. During the 2020 surge, the state's fraud-detection tooling froze roughly 1.4 million accounts, and later review found that a large share of those freezes hit legitimate claimants [1][2]. The scoring worked. The enforcement design failed.
Here is the mechanical problem. An automated score answers one question: how anomalous does this application look relative to a model. It does not answer whether the anomaly is fraud, whether a human confirmed it, or whether the underlying documents support the flag. When a score triggers automatic suspension, you have skipped every step that turns suspicion into defensible evidence.
Contrast that with defensible enforcement. A score of 0.87 arrives in a case queue. An analyst opens a guided investigation, sees the specific signals that drove the score, reviews the source paystub and ID, records a determination, and either confirms or overturns the flag. Only after that human sign-off does an adverse action proceed. The score started the process. A person, with accountability and a paper trail, finished it.
That distinction is the whole discipline. Everything below is how you operationalize it.
What CMS, Congress, and an MFCU Actually Ask For
Three audiences will examine your program-integrity work, and they ask different questions. If you build for only one, you fail the other two. CMS reviews for MARS-E compliance and program-integrity documentation. MFCUs want prosecutable evidence they can take to a grand jury. Congress and the GAO want assurance that you did not wrongly terminate eligible people [3].
An unexplained model score fails all three at once. It has no reasoning a CMS reviewer can follow, no source documents an MFCU can enter into evidence, and no reviewer accountability that protects you from a wrongful-termination finding. The score is a black box, and every one of these audiences distrusts a black box that touches benefits.
| Audience | What they examine | What fails their review | What passes |
|---|---|---|---|
| CMS program integrity | MARS-E controls, decision documentation, reviewer sign-off | Automated action with no human step or audit log | Confidence-scored signal plus documented reviewer determination |
| MFCU prosecutor | Source documents, chain of custody, reproducibility | Score with no underlying evidence or hashed originals | Packet with subject facts, forensic findings, custody log |
| GAO / Congress | False-positive rate, appeals outcomes, wrongful terminations | Bulk suspensions, no override tracking | Weekly FP metrics, human-override rate, appeal reversals |
The pattern across the table is consistent. Every audience wants to see the reasoning and the person behind the decision. Build the packet once, correctly, and it satisfies all three.
Anatomy of a Referral Packet That Holds Up
A referral packet that survives scrutiny has five fixed sections. Not four. Not "whatever the analyst had time for." Five, every time, so that a reviewer at CMS or an assistant attorney general at an MFCU can act without re-investigating the case.
The five sections are: subject identity facts, detection signals with confidence scores, source documents, reviewer determination and notes, and chain of custody. Fix the structure and quality becomes repeatable instead of dependent on which analyst caught the case.
Consider a concrete broker-ring referral. Detectory's broker-ring graph analytics surfaced 47 applications linked through a shared agent National Producer Number, three reused device fingerprints, and a common bank-routing pattern. The document forensics layer flagged 31 of those applications with paystubs sharing identical metadata and a cloned employer logo. Synthetic-identity intelligence found four SSNs belonging to deceased individuals and six reused across multiple applications. Each finding carries a confidence score. Each score links to the source artifact.
That packet does not say "these look fraudulent." It says: here are 47 applications, here is the graph edge that links them, here are the hashed source paystubs, here is the SSN death-master-file match, and here is the analyst who confirmed each. An MFCU can prosecute from that. A score cannot.
Use this checklist before any packet leaves the queue:
| Packet element | Required content | Blocked if missing |
|---|---|---|
| Subject identity facts | Name, SSN status, application IDs, linked entities | Yes |
| Detection signals | Each signal with calibrated confidence score and source layer | Yes |
| Source documents | Hashed originals of every document referenced | Yes |
| Reviewer determination | Named reviewer, decision, mandatory notes, timestamp | Yes |
| Chain of custody | Access log, model version, document hashes, retention tag | Yes |
Confidence Scoring: Why a Number Is Not a Decision
A confidence score is only useful if it means something reproducible. A calibrated 0.82 should mean that, across your validation set, cases scored 0.82 turned out to be fraud roughly 82 percent of the time. Raw model output that has never been calibrated against outcomes is a vibe, not a probability, and you should never route enforcement on a vibe.
Calibration is what lets you set risk tiers that route cases to different review depths instead of terminating anyone automatically. High-confidence cases get expedited review. Mid-tier cases get a full guided investigation. Low-confidence cases get a lighter touch or a documentation request rather than an adverse action. No tier ends in automatic termination.
The lesson embedded in those numbers is that pressure to catch fraud fast is exactly what produces wrongful terminations. When the political cost of missed fraud is high, teams lower thresholds and skip review. Calibrated scoring plus mandatory human review is the counterweight. It lets you act on real fraud without punishing the innocent people who look statistically unusual because they are self-employed, recently moved, or newly enrolled.
This is the same discipline behavioral-baseline work applies elsewhere in identity security. A single anomaly is a signal, not proof. The mature move is progressive trust: escalate scrutiny as evidence accumulates rather than acting on the first flag.
The Human-in-the-Loop Step You Cannot Skip
The reviewer step is the difference between the California outcome and a defensible program. It is not a rubber stamp. A proper guided investigation presents the analyst with the specific signals, the linked entities, and the source documents, then requires structured output before the case can close.
The reviewer workflow needs four non-optional fields: a determination (confirm, overturn, or request documentation), mandatory free-text notes explaining the reasoning, the reviewer's identity, and a timestamp. The reviewer must also have real authority to overturn a signal. If a paystub the model flagged turns out to be legitimate on inspection, the analyst overturns it, and that overturn is recorded as data.
Override data is not overhead. It is the feedback loop that keeps your program fair. When analysts consistently overturn signals in a certain confidence band or from a certain detection layer, that is your calibration telling you the threshold drifted. Feed override rates back into threshold recalibration, and you close the loop between detection and fairness. A program that never overturns anything is not accurate. It is not looking.
Chain of Custody and Reproducibility for Digital Evidence
Digital evidence has to be reproducible or it is worthless in front of an MFCU. Reproducibility means immutable logging of who saw what, when, which model version scored the case, and what source documents were examined, each hashed so tampering is detectable.
Analysts who work in AWS already trust this pattern. A CloudTrail record pivots on principal, timestamp, source IP, and action, and that four-field pivot is the backbone of every credential-abuse investigation. Apply the same logic to benefits fraud. Every case event should record the acting principal, the timestamp, the action taken, and the artifact touched. When session-and-token-theft investigations moved the industry toward per-principal, per-event logging [5], program integrity should follow the same evidentiary standard.
The model version matters more than teams expect. If you scored a case with model v3.1 and then retrained to v4.0, an appeal or audit six months later must reproduce the exact score the subject received. Log the model version with every determination or you cannot defend the decision later.
| Evidence element | Storage requirement | Retention period | Defensibility risk if missing |
|---|---|---|---|
| Source documents | Hashed, immutable object store | 6+ years (align to CMS) | No prosecutable original; MFCU declines |
| Model version | Recorded per determination | Life of case + appeals | Cannot reproduce the score; appeal reversal |
| Reviewer actions | Append-only access log | 6+ years | No accountability; wrongful-termination finding |
| Confidence scores | Versioned with calibration set | Life of case + appeals | Score cannot be explained to CMS |
| Chain-of-custody log | Tamper-evident, principal + timestamp | 6+ years | Evidence inadmissible |
Set retention to match CMS records requirements and your state's appeals window, whichever is longer. Under-retaining is how a solid case falls apart on appeal two years later.
Metrics That Prove Your Program Is Not the Next Pondera
You cannot manage defensibility without measuring it. Four metrics belong on a weekly dashboard: false-positive rate, human-override rate, referral acceptance rate by MFCU, and time-to-determination. Watch them weekly, not quarterly, because drift causes harm fast.
The two that predict a Pondera-style incident are false-positive rate and human-override rate. If your false-positive rate climbs, your thresholds are too aggressive and you are freezing legitimate people. If your override rate spikes, the model and reality have diverged and your scoring needs recalibration. Set trigger thresholds in advance: if either metric moves more than a set percentage week over week, recalibration is mandatory before the next batch of cases routes to enforcement.
Frequently asked questions
Can we terminate benefits based on the automated score? No. The score initiates review. A documented human determination must precede any adverse action, and the reviewer's identity and notes must be in the record.
What makes a referral packet acceptable to an MFCU? Five sections: subject identity facts, confidence-scored detection signals, hashed source documents, a named reviewer determination with notes, and a chain-of-custody log. The prosecutor should be able to act without re-investigating.
How do we stay audit-ready for CMS? Preserve model versions, reviewer actions, and source-document hashes for at least six years, and produce weekly false-positive and override metrics. MARS-E controls and documented review satisfy most CMS program-integrity requests [3].
What protects us in an appeal? Reproducibility. If you can reproduce the exact score, the documents reviewed, and the human who signed off, you can defend the decision. If any of those is missing, expect a reversal.
Start Here: A 30-Minute Audit of Your Current Evidence Trail
Do this in the next 30 minutes. Pull three recent adverse-action cases. For each one, check two things: is there a named human reviewer with a documented determination, and are the source documents preserved and hashed. If any of the three cases fails either check, you have a Pondera exposure right now, and you should stop automated adverse actions until the human step is enforced.
Start tracking one metric this week: the human-override rate on high-confidence cases. If it is near zero, your reviewers are rubber-stamping and you have no real safety net. If it is very high, your thresholds are miscalibrated and you are flagging legitimate people. A healthy program sits in a measured middle and watches the number move.
Six hundred thousand people lost their income because a score became a verdict with nobody accountable in between. The fix is not better scoring. It is confidence-scored signals, a human who signs their name to the determination, and an audit trail that reproduces the whole decision. Build the packet with five fixed sections, log every action like a CloudTrail event, and measure override and false-positive rates every week. That is what referral-grade evidence means, and it is what keeps your program off the front page.
References
[1] California State Auditor, "Employment Development Department: Significant Weaknesses in EDD's Approach to Fraud Prevention," 2021. https://www.auditor.ca.gov/reports/2020-628.2/index.html
[2] Los Angeles Times, "California froze 1.4 million unemployment accounts. Many were legitimate," 2021. https://www.latimes.com/california/story/2021-01-11/edd-fraud-unemployment-benefits
[3] Centers for Medicare & Medicaid Services, "Program Integrity and Medicaid," CMS.gov, 2025. https://www.cms.gov/medicare-medicaid-coordination/fraud-prevention/medicaid-integrity-program
[4] U.S. Government Accountability Office, "Unemployment Insurance: Estimated Amount of Fraud During Pandemic Likely Between $100 Billion and $135 Billion," GAO-23-106696, 2023 (most comprehensive federal estimate available). https://www.gao.gov/products/gao-23-106696
[5] Verizon, "2025 Data Breach Investigations Report," 2025. https://www.verizon.com/business/resources/reports/dbir/
[6] National Institute of Standards and Technology, "NIST Special Publication 800-63 Digital Identity Guidelines," 2024. https://pages.nist.gov/800-63-3/
[7] Centers for Medicare & Medicaid Services, "Minimum Acceptable Risk Standards for Exchanges (MARS-E) Version 2.2," 2024. https://www.cms.gov/marketplace/technical-resources/security-privacy