Training, not just access, determined whether courtroom AI delivered results

A large randomized field experiment in Pakistan’s judiciary suggests that artificial intelligence can raise public-sector productivity, but only when institutions invest in training people to use it well. The study, conducted by researchers from ETH Zurich, Imperial College London, and the New Economic School, tested an AI assistant called JudgeGPT across 1,559 judges in 118 courts, roughly half of Pakistan’s trial-court judges.

The headline result was not simply that AI helped. It was that access alone did very little. Judges who received targeted training alongside access to the system used it far more often and were associated with measurable gains in case resolution. Districts with moderate exposure to trained judges resolved about 1,848 additional cases per year, equivalent to a 6.3 percent increase, according to the study summary provided in the source material.

That makes the experiment notable beyond Pakistan. Governments around the world are testing large language models for legal research, case management, document drafting, and administrative support. What this trial adds is a clearer operational lesson: deployment strategy matters as much as the model itself.

A judiciary under strain provided a demanding test case

Pakistan’s court system offered a setting where even modest productivity gains could matter. The researchers said the country has fewer than two judges per 100,000 residents, compared with 22 across the European Union and 30 in England and Wales. At the end of 2024, 2.26 million cases were pending, with 82 percent of them sitting in trial courts.

Those figures point to a system working under heavy pressure, with limited staffing and sparse technology support. Before the experiment, only about a quarter of judges had ever used a large language model such as ChatGPT. That matters because it means the study was not testing AI among power users or in a digitally mature bureaucracy. It was testing whether a specialized assistant could make a difference in a real court system where time, personnel, and technical familiarity were all constrained.

JudgeGPT itself was built on OpenAI’s GPT-4 and adapted for Pakistani trial courts through retrieval-augmented generation. According to the source text, it searched a database of 129,235 documents, including 128,292 court rulings and 943 Pakistani laws. For each query, the system selected the 10 most relevant passages and generated a cited answer.

That design is significant because it narrows the model’s task. Rather than asking a general-purpose model to improvise legal analysis from memory, the system anchored responses in a defined corpus of local legal materials. In practice, that means the tool was positioned less as an autonomous judge and more as a research and drafting aid.

The biggest gains came from structured instruction

The researchers divided judges into three groups. One group received JudgeGPT access plus targeted instruction consisting of six 90-minute lectures over three weeks. Those sessions focused on which tasks fit the tool, where its limits were, and how to verify its output. A second group got access to the same AI system but received only a general seminar on technology and law. A control group attended the seminar without JudgeGPT access.

The contrast between those groups is the core finding. Judges who got targeted training used the system about four times as much as those who only received the general seminar. After 40 weeks, the trained group averaged nearly 60 logins and more than 200 prompts. The comparison group averaged around 20 logins and fewer than 50 prompts.

That gap in usage translated into different system-level outcomes. Districts with more trained judges resolved materially more cases. Even districts in the bottom quartile of exposure still cleared about 616 additional cases per year, according to the reported results.

The top panel shows that open legal research was the most common task, followed by text editing and text generation. The bottom panel shows that targeted training shifted use toward editing and summarization and away from broad legal questions, where hallucinations are more likely. | Image: Mehmood, Goessmann, Ash (2026)
The top panel shows that open legal research was the most common task, followed by text editing and text generation. The bottom panel shows that targeted training shifted use toward editing and summarization and away from broad legal questions, where hallucinations are more likely. | Image: Mehmood, Goessmann, Ash (2026)

The implication is straightforward. AI adoption in institutions is not just a software rollout problem. It is a workflow and capability problem. Without practical instruction, users may underuse the system, misunderstand where it helps, or avoid relying on it for meaningful work. With training, the same model can become part of everyday decision support.

Quality did not appear to deteriorate

Productivity gains in courts would mean little if they came at the expense of judgment quality or fairness. The source text says ruling quality held steady or improved modestly. Appeal rates per 1,000 resolved cases fell slightly, while judges’ working hours and work-life balance did not materially change.

Those findings do not settle every concern around judicial AI, but they do address one common fear: that faster output necessarily means sloppier decisions. In this experiment, the reported gains did not come from longer hours or obvious quality erosion. Instead, they appear to have come from more efficient research and case handling.

That distinction matters for policymakers considering similar tools in overburdened courts, agencies, or public defenders’ offices. If AI is used to surface relevant precedents, summarize records, and draft support materials, it may reduce bottlenecks without shifting final legal responsibility away from judges.

The economics make the result harder to ignore

The study estimated savings of about $38.50 for every dollar invested, based on the equivalent cost of hiring additional judges to obtain similar increases in case resolution. That figure should be read as an estimate rather than a guaranteed budget outcome, but it helps explain why public institutions are likely to keep testing similar systems.

In many countries, adding judges, clerks, or legal researchers is expensive and slow. Training existing judges to use a narrowly designed AI assistant may be cheaper and faster, especially when case backlogs are politically salient and institutionally damaging. A tool that improves throughput without extending work hours or worsening appeals is likely to attract attention well beyond Pakistan.

Still, the source material also points to a caution that can be generalized. The intervention worked because it combined model access, curated legal retrieval, and deliberate instruction about limits and verification. Strip away those components and the case for productivity gains becomes much weaker.

Why the study matters beyond the courtroom

The broader significance of the experiment is not that AI can replace legal judgment. It is that well-scoped AI systems may improve the performance of strained institutions when they are introduced with enough operational discipline. The lesson applies to courts, but also to tax agencies, benefits systems, hospitals, schools, and regulators: human adoption is a central variable, not an afterthought.

Much of the public conversation around generative AI still swings between hype and dismissal. This study lands somewhere more useful. It suggests that the technology can produce real gains in a demanding public setting, but only under conditions that make those gains plausible: domain-specific retrieval, explicit safeguards, and training that teaches users both what the tool can do and what it should not be trusted to do alone.

For governments weighing whether AI belongs in high-stakes administrative systems, that may be the most important finding of all. The return did not come from deploying a model and waiting for transformation. It came from treating AI as an institutional capability that had to be taught, integrated, and checked.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com