AI Policy & Governance, Government Surveillance
Changing Course to Get It Right: The Advisory Committee Reviews Its AI Evidence Rule
When a court is deciding whether someone goes to prison, whether a company pays for a defective product, or whether the government violated a person’s constitutional rights, the evidence supporting that decision needs to be reliable. That principle sounds obvious. But as artificial intelligence (AI) increasingly produces the “evidence” that judges and juries are asked to trust — a facial recognition match, a gunshot-detection alert, a pattern-of-life analysis pulled from a phone — the rules that determine what information gets into court need to keep pace.
That is why we have had a close eye on Proposed Federal Rule of Evidence 707 (“Proposed Rule 707”), which would govern when AI-generated information can be admitted as evidence in federal court — and many state courts where the federal rules are routinely adopted — without a human expert to explain it.
In May, after several public hearings and more than 70 written comments from individuals and organizations, including our own, the Advisory Committee on Evidence Rules[1] chose not to push the draft rule towards the finish line in its current form. Instead, it suggested revising the proposal, narrowing its scope, and inviting further input from technical experts and practitioners at a mini-conference to be held during its fall meeting scheduled for October 15, 2026.
Changing course based on informed advice was exactly the right call.
Why Proposed Rule 707 Matters for Everyone
It is tempting to file Proposed Rule 707 under “technical lawyer stuff.” But we must resist that temptation, because the stakes are very meaningful for anyone who may find themselves embroiled in litigation, whether it be for divorce proceedings, child custody disputes, slip-and-fall or defective product cases, or alleged crimes. AI systems are now generating the kinds of conclusions that, not long ago, only a trained human expert could offer, and courts are being asked to treat those outputs as trustworthy.
Consider the gap Proposed Rule 707 needs to close. If a human forensic analyst takes the stand to testify that a bullet casing recovered from the scene of a crime was fired from a gun registered to the defendant, that analyst can be questioned about their training, their methodology, and their error rate. Rule 702, the long-standing test for the reliability of expert testimony, determines whether a jury ever hears that testimony. But when an AI system spits out the same conclusion, and the police officer who “pressed the button” reads the conclusion into the record, Rule 702’s requirements do not apply. As a result, the AI system’s conclusion can carry all the weight of expert testimony, and perhaps even more, while dodging the scrutiny we apply to experts.
That asymmetry is dangerous for every litigant, and it is especially dangerous for criminal defendants, who often lack the resources to hire their own experts to probe the AI system the government used in its investigation or prosecution. The stakes here are liberty itself. Flawed algorithms that nonetheless look authoritative can put a thumb on the scale in exactly the cases where the scale matters most.
So the Committee was right to recognize that AI-generated evidence raises real reliability concerns and that the existing federal rules of evidence do not squarely address them. We said as much in our comments, and further explain here why the growing prevalence of AI-generated evidence does warrant a new rule.
How Proposed Rule 707 Missed the Mark
CDT and others flagged several problems in Proposed Rule 707. First, the original draft was overbroad, sweeping in all “machine-generated evidence” and arguably capturing some mundane digital outputs, like a server access log that simply records timestamps and IP addresses. Those records raise none of the reliability concerns the Committee was actually worried about, which center on AI systems that make predictions or draw inferences from data. Treating an access log the same as a facial-recognition match would burden courts with needless debates, and disadvantage parties who cannot afford to hire an expert.
Second, the attempted fix to overbreadth — a carveout for the outputs from “simple scientific instruments” — was too vague to do the work. The temperature reading from a digital kitchen thermometer might be excluded, while server access logs might not. Plus, judges in different jurisdictions would inevitably draw lines in different places, producing a patchwork of inconsistent precedents.
Third, and most importantly, Proposed Rule 707 awkwardly borrowed Rule 702’s criteria for evaluating the reliability of human expert witnesses, applying them to evaluations of AI systems’ reliability. But the right questions to ask about human experts are not necessarily the right questions to ask about AI systems. As reflected in Rule 702, human expert reliability flows from education, experience, and professional norms. But an AI system’s reliability flows from entirely different sources, like how it was designed, what data it was trained on, and how it was tested and evaluated. It also matters what those tests revealed about the system’s capabilities and limitations, how it was deployed, and whether the opposing party has been able to meaningfully examine it.
What Changed
Here is the encouraging part: the Committee’s revised draft of Proposed Rule 707, and the accompanying Committee Note that would guide litigants and judges in implementing the rule, reflect a number of the changes we and other commenters urged.
For instance, the revised rule applies only to information generated by “artificial intelligence,” not all “machine-generated” information. As a result, the rule no longer includes the vague carveout for “simple scientific instruments” precisely as we suggested. By focusing on AI systems, rather than every digital output, the revised draft takes aim at systems whose reliability genuinely turns on predictions and inferences, and steps back from covering routine records that look like traditional documentary evidence.
Perhaps even more significantly, the revised Committee Note now also points litigants and judges more clearly toward additional questions that matter. The original Note anticipated that applying the original Proposed Rule 707 to determine whether evidence is admissible would require the court to answer whether the AI system was tested in circumstances like those of the case at hand and whether it can therefore produce an accurate output in the case. The revised Note adds that such analysis will also usually involve asking whether an AI system’s conclusions can be reproduced or audited; whether the AI system has been evaluated by people independent of the developer or deployer; and whether the opposing party has received enough information to test the AI system for themselves. Readers of our prior comments will recognize these additions because we urged the Committee to direct judges to these questions, which are critical for determining whether an AI system can be trusted. However, we still believe these questions should be written into the rule itself, rather than appear in the Committee Note, which is only advisory and may be skipped over by judges and litigants.
That said, another central concern persists. The revised rule still treats expert testimony as the ordinary, but not mandatory, path to admissibility. This would allow a court in “exceptional circumstances” to rely on other proof that an AI system is reliable. We remain skeptical that the questions the Committee rightly wants judges to ask can be answered without an expert. Evaluating whether a system’s training data was representative, whether it was validated under realistic conditions, or whether the opposing party can meaningfully test it generally calls for scientific, technical, or specialized knowledge, which is the precise nature of expert testimony. We hope the Committee will use this opportunity for further review to consider whether any AI-generated evidence that would be subject to the rules for expert testimony if offered by a human should ever come in without an expert.
The Value of Vetting
Ultimately, the Committee decided not to push Proposed Rule 707 forward at its meeting in May. Instead, it decided the more prudent course was to revise the proposed rule and vet the revision in the coming months, with AI experts and litigators providing additional, essential input. A rule of evidence that governs when AI’s word can help send someone to prison is not the place to settle for “good enough.” Build it too broadly, and courts might drown in needless litigation over ordinary records. Build it on the wrong criteria, and unreliable AI outputs might sail through without scrutiny. And, of course, the cost of getting this balance wrong would likely fall hardest on the people least equipped to fight back. A measured, expert-informed process is how you avoid those outcomes.
[1] The Judicial Conference of the United States establishes the rules of evidence for federal courts, and it routinely relies on the work of its Advisory Committee on Evidence Rules.
