Skip to content
← All news
3 min read

A Pre-Registered Trial Found an AI Assistant Helped Pakistani Judges Clear 6.3% More Cases

A randomized controlled trial across 1,559 judges found a GPT-4 research assistant raised case throughput without hurting ruling quality.

A pre-registered RCT found an AI legal assistant helped Pakistani judges close 6.3% more cases.

A pre-registered randomized controlled trial across 1,559 trial court judges in Pakistan found that judges trained on an AI legal research assistant closed about 6.3% more cases a year than judges who were not, with ruling quality holding steady or improving.

What was tested

Researchers Elliott Ash of ETH Zurich and Sultan Mehmood of the New Economic School in Moscow, working with collaborators at Imperial College London, built JudgeGPT, a GPT-4 retrieval-augmented assistant that searches a database of 129,235 documents: 128,292 past rulings and 943 laws. They rolled it out across 118 courts, covering roughly half of Pakistan's trial judges.

Access to the tool wasn't the whole intervention. One study arm received six 90-minute lectures over three weeks teaching judges how to query the system and verify its output; another arm got a lighter general seminar; a third was a pure control. That three-way split is what lets the researchers attribute the case-clearance gain to the training, not just to having a chatbot on a desk.

The numbers

Districts with trained judges resolved about 1,848 more cases per year than control districts, a 6.3% increase over the observation window. In pairwise comparisons of ruling quality, 59% of rulings from trained judges were preferred over the equivalent control ruling, against 42% for controls, so throughput went up without an apparent quality trade-off, at least on the measures the researchers used.

The trial has circulated with a striking headline number: a $38.50 return for every dollar spent on the program. That figure comes from the researchers' own back-of-envelope calculation of what it would cost to hire enough extra judges to match the same increase in resolved cases. It is not an independently audited return on investment, and it should be read as the researchers' own cost comparison, not a verified financial result.

Why the trial design holds up

The trial was pre-registered on the AEA Social Science Registry as trial #12906 in January 2024, before Wave 1 launched that February and well before any results existed. Wave 2 expanded the rollout that October, and the full observation window ran 40 weeks. Committing to what would be measured and how, before seeing the answer, is what separates this from a vendor case study.

Why a build studio cares

This is a rare data point on AI tooling measured with a control group instead of a testimonial. Most claims about AI productivity gains in white-collar work come from vendor benchmarks or self-reported surveys; a pre-registered RCT with 1,559 participants and a real backlog to clear is a different category of evidence, and a useful reference point for any client asking what a retrieval-augmented assistant built on their own documents could plausibly do for a team drowning in case files, contracts, or tickets.

Next step: read the trial registration at socialscienceregistry.org or the writeup at The Decoder. If you're weighing a retrieval-augmented assistant against your own backlog, write to hello@gattyworks.com.

AILegal TechResearchPakistanJudgeGPTLegalTechAIJusticeCourtBacklogsRandomizedControlledTrialGPT4RAGJudicialReformAIatWork

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.