top of page
Search

Before the Bench: Court-Side Solutions Against AI Hallucinations

By Nima Movahedi, Policy Fellow - Encode London



Executive Summary


Generative artificial intelligence (GenAI) is transforming the practice of law. Seventy-seven percent of legal professionals now use large language models (LLMs) for tasks such as research, drafting, and citation generation; for millions of self-representing litigants, these tools have made the otherwise impenetrable legal system navigable without costly professional services. 

Yet the technology carries a fundamental risk: generative AI systems routinely fabricate case law, invent citations, and advance erroneous legal reasoning. These false outputs, known as hallucinations, are actively corrupting the integrity of judicial proceedings worldwide.

58–88%

Hallucination rate

from public LLMs

17–34%

Hallucination rate in legal-specific AI tools

~60%

U.S. incidents from

pro se litigants

Within the United States, the institutional response has not merely been insufficient but structurally foreclosed. A deregulatory executive branch, a gridlocked Congress, and federal preemption of state AI laws have produced a regulatory vacuum. Courts cannot afford to wait. This memo consequently advances a graduated series of court-side interventions to protect judicial proceedings from the rise of AI hallucinations – each immediately implementable through existing judicial governance structures and independent of external political actors or resources:


  • Mandatory hyperlinking of cited authorities

Require functional hyperlinks in all e-filed documents, creating a self-executing verification mechanism at the point of filing 

  • Randomized citation auditing 

Utilize existing citation verification infrastructures to selectively verify citations as a probabilistic deterrence mechanism 

  • Targeted screening and education for pro se litigants

Combine an enhanced review of pro se filings with educational resources to reduce AI-hallucinated legal content from the highest-risk population


Background


Since its widespread emergence in 2023, generative artificial intelligence (AI), spearheaded by large language models (LLMs), has rapidly reshaped professional practice across industries. The legal profession has been particularly receptive: LLMs promise increased efficiency, reduced costs, and streamlined access to legal information. A 2025 Thomson Reuters report found that 77% of legal professionals now use generative AI for services ranging from document summarization and legal research to brief drafting and even contract generation. Perhaps more consequential, the open accessibility of free LLMs has helped equalize an otherwise complex and costly legal system for litigants without access to professional services. As Chief Justice John Roberts of the U.S. Supreme Court highlighted in his 2023 Year-End Report on the Federal Judiciary, generative AI holds “great potential to dramatically increase access to key information for non-lawyers…[and] litigants with limited resources”. Despite this promise, the rapid adoption of AI in law has been accompanied by serious ethical concerns regarding the technology’s accuracy, susceptibility to bias, and the lack of accountability frameworks governing its outputs.


The Hallucination Problem 


Central to these concerns is the phenomenon of “hallucinations”. Cambridge University defines these as instances where artificial intelligence is designed with “qualities that the human brain has” and “produces language in a way that seems human,” that is, in fact, false. Within a legal system, “where adherence to authorities is paramount”, the possibility that AI may invent convincingly false legal citations or advance erroneous interpretations “can lead to nonsensical- or worse, harmful and inaccurate -legal advice or decisions”. 


The empirical evidence substantiating this possibility is considerable. In the first systematic study of legal hallucinations across public-facing LLMs, Dahl et al. found that popular models hallucinate at alarmingly high rates: “between 58% of the time with ChatGPT 4 and 88% with Llama 2” when presented with verifiable questions about U.S. court cases. Critically, this risk is not distributed equally; the authors note that “the risks are highest from those who would benefit from LLMs most, under-resourced or pro se litigants…[potentially advancing] existing inequalities in access to legal services”.


The Limits of Legal-Specific AI


Recognizing these limitations in general-purpose models, major legal technology providers have developed proprietary AI tools built from the ground up with curated legal databases and jurisprudential reasoning. Two competitors, Thomson Reuters and LexisNexis, reflect a new market race: Thomson Reuters has released its Westlaw AI-Assisted Research, claiming “over 95% accuracy and traceability”, while LexisNexis promoted its Lexis+ AI as delivering '100% hallucination-free' legal citations.


However, an empirical evaluation by Stanford University's RegLab and the Human-Centered Artificial Intelligence Institute found these claims significantly overstated. While specially trained LLMs do “reduce errors compared to general-purpose AI models like GPT-4”, the LLMs nevertheless “hallucinate an alarming amount of the time”. The tested hallucination rates for Lexis+AI (17%) and Westlaw AI (34%) are consistent with a growing body of computational research suggesting that the problem may be structural rather than developmental: a feature of how language models function, not solely a flaw awaiting eventual perfection. As Xu et al. argue, “hallucination is inevitable for any computable LLM, regardless of model architecture, learning algorithms, prompting techniques, or training data”. While architectural innovations such as retrieval-augmented generation (RAG) aim to reduce this risk by constraining LLMs to a verified external database, these systems remain susceptible to interpretation errors and limited database coverage, which perpetuate risks incompatible with judicial standards. Indeed, OpenAI, the parent company of the most widely used LLM, ChatGPT, has itself validated this assessment, conceding in an internal research paper that “even if the training data were error-free, the objectives optimized during language model training would lead to errors being generated”.


Yet the premise of error-free data has been called into question: computer scientists and sociologists alike argue that real-world training data is inherently incomplete, biased, and unevenly representative. OpenAI researchers further acknowledge systems built “with realistic training data… [should] expect even higher error rates” than those observed under experimental conditions. This presents a discouraging implication for the future stability of the legal system. Incremental improvements to general and specialized LLMs, however beneficial to other fields, cannot eliminate the root risk that fabricated and inaccurate AI outputs pose to the public trust in judicial institutions. And even as prioritized legal AI tools grow in popularity among American BigLaw firms seeking to augment productivity, leading services like Harvey and Legora, which together serve 70% of top legal firm customers, face the same fundamental limitation: no degree of improved user-end accessibility, workflow integration, or marketing can compensate for an inability to guarantee fully credible outputs.  


The success of legal systems as conflict-resolution mechanisms rests on their capacity to deliver fair, faithful, and accurate outcomes for litigants. Even as private companies pour billions into minimizing the risk of hallucinations, the frequency of AI-generated errors continues to rise in step with the technology’s growing adoption. Thus, a flood of AI-generated content can reasonably be expected to continue undermining public confidence in the judicial system. 


Problem Statement


The institutional response to generative AI hallucinations within the American legal system, though considerable in volume, has been narrow in scope, with near-exclusive attention focused on the user end of the equation. Despite mounting empirical evidence that AI-generated content is appearing at an accelerating rate in legal proceedings (Charlotin database), no unified framework exists to address this threat at the level of the courts themselves.


Absent a Unified Institutional Response


At the federal level, meaningful guidance across all three branches of the U.S. government has been noticeably absent. The U.S. Congress has yet to establish any comprehensive federal guidelines comparable to the European Union’s 2024 ‘AI Act’ to regulate the rapid proliferation of AI systems. Though the U.S. House of Representatives passed a bill in 2025 imposing “a 10-year moratorium on the enforcement of state laws governing artificial intelligence”, the U.S. Senate decisively struck it down with a 99-1 vote, signaling a deep divide over AI policy between the two chambers. 


In an attempt to fill this vacuum, state governments have pursued their own regulatory frameworks. Colorado’s 2024 Artificial Intelligence Act (SB 24-205) was the first comprehensive state-level law intended to “protect consumers from… foreseeable risks of algorithmic discrimination”. California, Texas, Florida, and others followed with their own AI regulatory provisions. However, two structural limitations have hampered their efficacy. First, most state-level AI regulations are oriented toward consumer protection and transparency requirements rather than judicial proceedings. Second, and arguably more consequential, the enforceability of these provisions has been undercut by the Trump Administration's deregulatory posture toward AI. Entering his second term with strong support from major U.S. technology firms, U.S. President Trump shifted regulatory priorities to promote the unfettered development of AI systems. Executive Order 14179 codified this, mandating that  “United States AI companies shall be free without state-by-state regulation” in order “to remove barriers and encourage the adoption of AI applications across sectors”. Shortly thereafter, the United States Department of Justice established an Artificial Intelligence Litigation Task Force to “challenge state AI laws deemed inconsistent with a 'minimally burdensome national policy framework for AI”. 


Taken together, these developments do not merely reflect a gap in federal AI policy: they illustrate an active institutional commitment to preventing its development. The executive branch has signaled that any attempt to regulate AI will face political and legal resistance, while Congress has shown neither consensus nor legislative priority on the matter. Although Congressional elections are approaching in November 2026, even a shift toward legislators who prioritize AI regulation would almost certainly face a presidential veto, requiring a two-thirds supermajority in both chambers to override it.   


The Regulatory Patchwork Regime


Lacking broader institutional guidance, individual federal and state court judges have been left to act on their own. In June 2023, Judge Brantley Starr of the Northern District of Texas issued the first standing order of its kind, requiring attorneys to certify that “no portion of [a] filing was drafted by generative artificial intelligence…or that any language drafted by generative artificial intelligence was checked for accuracy”. Since then, a wave of federal and state judges have issued their own AI directives that “vary widely with respect to applicability, requirements, and repercussions”. This resulting mosaic, ranging from “proactive and detailed...to exploratory and foundational…[is] creating challenges for attorneys practicing across multiple jurisdictions”. Alongside these judicial efforts, the American Bar Association and individual state bars have released guidance on how attorneys should uphold “meritorious claims and contentions and candor toward the tribunal” when using generative AI. 


Despite this proliferation of user-facing directives, the empirical record suggests they have done little to stem the tide. Charlotin's AI hallucination database shows that incidents in legal filings have more than doubled over the past year, with the United States accounting for over 300 of the more than 800 documented cases worldwide. The decentralized, attorney-focused nature of these responses has produced a patchwork regime that relies overwhelmingly on self-policing in the absence of institutional support. There is, consequently, a need for practical court-side safeguards that can effectively deter the rise of hallucinated content before it corrupts judicial proceedings. 


Policy Recommendations 


The following recommendations propose a graduated series of court-side interventions designed to deter, detect, and address AI-generated hallucinations in legal filings. Each provision is calibrated to impose progressively greater institutional commitment, moving from low-cost reinforcement of existing norms to more resource-intensive safeguards. Crucially, these recommendations are deliberately curated to be pragmatically implementable in all judicial environments, regardless of variance in procedures, resources, or digital expertise. The National Center for State Courts and the Judicial Conference of the United States serve as viable institutional channels for translating these policies into coordinated practice across state and federal courts. Taken together, these recommendations constitute a layered framework that empowers courts to mitigate AI hallucinations unilaterally without dependence on regulatory action that the current political environment fails to provide. 


Baseline Measure: Reinforcing Existing Judicial Conduct Standards 


Before advancing new institutional safeguards, courts should ensure that existing professional obligations are clearly communicated in the context of generative AI usage. Many courts have operationalized this: standing orders and local rules are currently the primary tools courts use to deter the submission of inaccurate AI-generated content. Grounded in state and federal rules of procedure, these directives reaffirm an attorney’s professional obligations and the sanctions that may follow from their breach.  


This step is a necessary foundation, but it remains insufficient on its own. While reaffirming conduct standards creates a useful barrier against unintentional errors, the empirical evidence shows that declaratory measures alone have failed to meaningfully curb the rise in hallucinated filings. Nonetheless, its value as a baseline solution, rather than a more resource-intensive safeguard, merits far broader institutional implementation. A 2025 Bloomberg Law study found that only 39 U.S. federal judges have implemented standing AI orders; against a federal bench of approximately 900 judges, considerable work remains.


Recommendation 1: Mandatory Hyperlinking of Cited Authorities 


Courts should adopt new procedural rules requiring that all electronically filed pleadings, motions, memoranda, and briefs include functional hyperlinks to any judicial opinion, statute, or regulation cited for review. 


By tying arguments to a direct authoritative source, whether a government website, recognized legal databases, or public repository, hyperlinking eliminates the most prevalent form of AI hallucination in legal proceedings: the fabrication of nonexistent case law. This measure shifts a court’s capacity to intervene from retrospective to proactive. Rather than discovering fabricated submissions during or after judicial review, the hyperlinking requirement compels litigants to locate and verify every cited authority before filing. With the burden of verification passed to the filer, judges and court staff can review the validity of submitted citations at a glance, without specialized technological or legal expertise. 


Key Advantages:

  • Self-executing verification: Rule begins at the point of filing without necessitating additional judicial resources or post-submission review;

  • Front-end rather than back-end control: Filters for fabricated contents proactively before they enter the record, minimizing the corrupting impact;

  • Resource neutrality: Verifying functionality of hyperlink citations requires no additional technological infrastructure or specialized legal knowledge, reducing burdens on courts and staff;

  • Proven feasibility: The New York Commercial Division’s Administrative Order 133/20 demonstrates that hyperlinking requirements are administrable within the existing judicial infrastructure without additional training or resources.


Key Limitation

  • Inability to detect mischaracterization: A hyperlinked citation confirms an authority exists, not that it has been accurately represented;  


Recommendation 2: Random Citation Verification 


Courts should implement a system of randomized citation auditing in which a selection of briefs and motions is screened against citation-verification databases before reaching judicial review.


The infrastructure for this already exists. The Free Law Project’s CourtListener platform provides a ‘Citation Lookup and Verification API’ that parses a database of over 18 million American case citations. Tested against 50 million citations, the citation verifier can scan 250 citations per request and return a determination of whether a given citation was found, not found, or ambiguous. Implementation is flexible: a court could manually select a number of random filings and submit them to the API or directly integrate the tool into its electronic filing system using CourtListener’s open-source code. 


The deterrent logic underpinning this recommendation is in line with traditional regulatory theory. The Internal Revenue Service audits fewer than 1% of individual tax returns annually, yet this random enforcement mechanism shifts risk calculations: the expected cost of noncompliance rises even without a high probability of being caught. Applied to under-resourced courts, even the most modest rates of citation auditing would credibly signal the institutional consequences of submitting fabricated content. This recommendation complements the hyperlinking rule by adding an institutional filter, independent of litigants’ self-policing and without requiring a significant influx of judicial resources. 


Key Advantages:

  • Independent of filer compliance: Court-side verification operates regardless of litigants' diligence or good-faith behavior;

  • Probabilistic deterrence: The randomized audit model deters hallucinations by signaling heightened judicial enforcement, altering the cost-benefit calculus for filers who might otherwise submit unverified AI content;

  • Existing infrastructure: CourtListener’s API offers the judiciary a tested, operational tool for citation verification; The Free Law Project’s nonprofit mission commits them to maintaining the API as free and accessible to all;

  • Minimal resource burden: The API is navigable without a technological background for manual use, and its open-source status permits straightforward integration into a court’s existing electronic filing system. 


Key Limitations:

  • Database coverage gaps: CourtListener’s database, while extensive, does not capture every reported decision or verify statutes, regulations or secondary sources;

  • Inability to detect mischaracterization: As with the hyperlinking rule, this provision only identifies whether a citation exists, not whether it is distorted in its application;

  • Dependence on deterrent effect: By design, randomized auditing lacks universal oversight; its efficacy is tied to perceived enforcement, meaning that indifferent or uninformed filers may not be reached by the mechanism’s logic. 


Recommendation 3: Specialized Focus Towards Self-Represented Litigants


Courts should implement a two-pronged approach for self-represented (pro se) litigants that combines enhanced screenings of pro se filings with structured educational resources designed to prevent hallucination-related errors.


Of the 916 documented AI hallucination incidents in U.S. legal filings within Charlotin’s public database since 2023, 559, over 60%, originated from pro se litigants. This disproportionate trend reflects how resource and education gaps can lead to reliance on generative AI in legal proceedings. Addressing this puzzle would, by the numbers alone, substantially reduce the volume of hallucinated content entering courts. 

To target this concentration of risk, courts should extend the citation-verification framework outlined in Recommendation 2 to intentional screenings of pro se filings. Where Recommendation 2 suggests randomized auditing as a probabilistic deterrent, this measure applies more systemic review to categories of filing where fabricated content is most likely to appear.  


To address the root cause rather than its symptoms, courts should prioritize educating pro se litigants over sanctioning those who may have acted in good faith. Given its complexity, courts should create a standardized information guide for the use of generative AI in legal proceedings. Optimally offered as a physical handout prior to proceedings, such a guide should include:


  1. A plain-language explanation of what AI hallucinations are and how they frequently fabricate false information under the guise of appearing correct;

  2. A step-by-step verification checklist instructing litigants to confirm every citation using free public-access databases like CourtListener and Google Scholar, with guidance on how to search for a case;

  3. A clear warning that submitting fabricated citations, even if unintentional, may be detected and could result in adverse consequences;

  4. Contact information for local legal aid organizations and the court’s own self-help center.


Courts with the capacity to develop digital resources or deliver more comprehensive programming should do so. However, for the majority of courts, operating with full dockets and limited resources, a simple, well-designed handout would represent a meaningful and immediately deployable intervention.  


Key Advantages: 

  • Targets the highest-risk population: Concentrating efforts on pro se filers addresses the largest source of fabricated content in judicial proceedings;

  • Prevention over punishment: Education aims to avoid sanctioning litigants who may have acted in good faith or been unaware of the risks;

  • Respects the reality of generative AI usage: Rather than discouraging generative AI use entirely for self-representing litigants, this recommendation promotes responsible practices, aware of its beneficial use in helping navigate an otherwise inaccessible system;

  • High accessibility: Providing a physical handout comprising plain-language explanations becomes useful for litigants across all educational backgrounds and does not presuppose digital literacy. 


Key Limitations: 

  • Resource intensity: Mandated screenings of pro se filings demand considerably more judicial attention than random auditing, and education materials require development, translation, and distribution;

  • Reduced uniformity: Variation in local legal aid infrastructure, court budgets, and administrative capacity may produce disparities in quality of educational resources across jurisdictions; 

  • Comprehension barriers: Although simplified, explaining generative AI usage and citation requirements remains inherently complex; litigants may become confused by guidelines or discouraged from AI’s beneficial uses entirely.


Implementation Pathway


The most viable channels for adopting these recommendations run through the National Center for State Courts (NCSC) and the Judicial Conference of the United States. The NCSC already serves as a hub to “drive innovation and advancement in court and justice systems”, hosting conferences and publishing resources to educate judicial staff on artificial intelligence. Its existing focus on guiding state jurists on legal developments and solutions to contemporary problems makes it well-positioned to translate these proposals into adaptable model rules for courts across jurisdictions. At the federal level, the Judicial Conference of the United States, composed of senior justices and judges, can “prescribe general rules of practice and procedure and rules of evidence” for all 94 federal judicial districts. Through its Committee on Rules of Practice and Procedure, the federal Judicial Conference could incorporate hyperlinking and citation-verification standards into amendments to the Federal Rules of Civil and Appellate Procedure or issue advisories encouraging individual districts to adopt them. Either pathway offers a viable means of achieving institutional coordination without dependence on congressional or executive action.


Conclusion 


The proliferation of generative AI in legal practice is an irreversible trend. For millions of litigants, especially those without the resources to retain counsel, large language models have made the legal system more navigable for tasks that previously required professional expertise. Yet, the promise of democratizing legal access is inseparable from the consequences of relying on a technology that routinely fabricates outputs. When nonexistent citations are submitted to a court, even with the best intentions, the integrity of judicial proceedings is compromised. Without an adequate institutional response, which has been largely absent from the U.S. government, this trend will degrade the broader public trust in courts as institutions capable of delivering fair and accurate outcomes. The empirical record has substantiated this concern, as the fragmented regulatory landscape in America has failed to curb the growth of hallucinations in legal filings. 


The recommendations advanced in this paper are designed to address this gap: instead of lobbying to completely restrict the use of LLMs, safeguards should be implemented within the courts themselves. Mandatory hyperlinking transfers the burden of citation verification from an abstract professional obligation into an unavoidable self-executing barrier at the point of filing. Randomized citation auditing, using tools like CourtListener’s API, deploys probabilistic enforcement to deter hallucinations and provides courts with a streamlined means to independently review filings. Targeted screenings of self-representing litigants, along with accessible education resources, address the single largest source of hallucinated content in American courts. 


These measures, individually or collectively, cannot eliminate all AI fabrications in legal proceedings. They can, however, substantially reduce the volume of hallucinated content reaching judicial review through a more pragmatic approach. Each recommendation is calibrated for immediate implementability: low cost, grounded in existing infrastructure, and operable without specialized technical expertise. Through existing institutions, such as the National Center for State Courts and the Judicial Conference of the United States, these policies can be readily translated into a coordinated approach to replace the current patchwork of individual standing orders. In the continued absence of unified regulatory frameworks, this graduated approach offers a necessary and actionable first step.

 
 
 

Comments


bottom of page