Learn · Intermediate
Randomized trials: measuring what an AI tool actually changes
A randomized controlled trial estimates a treatment’s causal effect by assigning otherwise comparable participants to receive it or a control condition by chance. In AI studies, the treatment might be access to an assistant rather than a new medicine. Random assignment helps distinguish what the tool changed from differences that already existed between the people who chose to use it.
Imagine evaluating an AI tutor by comparing students who subscribe with students who do not. Subscribers may have more money, stronger motivation or greater difficulty with the subject. Their later scores mix the tutor’s effect with those preexisting differences. Randomly offering access makes assignment independent of those characteristics in expectation. The groups can still differ by chance, but that uncertainty can be analyzed rather than assumed away.
Donald Rubin’s canonical 1974 paper on estimating causal effects gives a useful framework: potential outcomes. Each person has a hypothetical outcome under treatment and another under control. You observe only one of them. A lawyer cannot complete the same first exposure to a drafting tool both with and without that exposure. The missing outcome is why a simple before-and-after story cannot reveal an individual causal effect.
The experiment estimates an average by comparing groups. Suppose randomly assigned users with an assistant average eight correct tasks and controls average six. The two-task difference estimates the average effect of being assigned access under that experiment’s conditions. These numbers are illustrative. They do not identify which participant benefited, whether an unusually skilled person would gain more or what would happen with a different tool.
The intervention must be defined carefully. Offering access is different from requiring use. Some assigned users may barely open the assistant; some controls may find another tool. An intention-to-treat estimate compares people according to their original assignment. It preserves the benefit of randomization and answers the operational question of what an access policy accomplished. Comparing heavy users with light users afterward introduces selection again, because intensity of use was not itself randomized.
The Google patent-lawyer experiment illustrates the distinction. David Autor and coauthors randomized 133 lawyers to assistant access or control and later assessed drafting and unaided redlining. The full working paper reports better assisted drafting, while the clearer unaided advantage was among experienced attorneys. Random assignment supports a causal interpretation of access within the study, but the measured outcome remains performance on specified tasks.
A causal estimate can be precise about assignment and narrow about meaning. Better patent drafts do not automatically mean more valuable patents, faster ordinary casework or deeper lasting expertise. Those are different outcomes. The researchers’ unaided assessment is a useful attempt to measure carryover, but a single task after three months cannot establish career-long skill development. Readers should inspect the definition of success before interpreting the size of an effect.
Completion also matters. Only 91 randomized lawyers completed the final redlining task. If stronger treatment participants stay while weaker ones drop out, the observed difference can mix the intervention with changed group composition. The paper reports similar completion by assignment and no evidence of differential attrition. That is reassuring, but a smaller final sample still changes precision and warrants transparent accounting. Randomization at the start does not make missing outcomes disappear.
Subgroup effects need their own uncertainty. A result that is statistically detectable for seniors and undetectable for juniors does not, by itself, prove the treatment effects differ. The relevant test compares the effects directly. Similarly, an average near zero can conceal winners and losers. The trial’s junior redlining distribution became more dispersed, showing why a mean and its uncertainty can leave useful distributional information outside the headline.
Preregistration complements randomization by recording questions and analysis choices before outcomes are known. It addresses researcher discretion; randomization addresses comparability of treatment assignment. Holdout sets address a different problem again: separating development data from assessment data. These tools protect distinct links in the evidence chain.
A trial’s internal validity is also distinct from its external reach. Eleven sophisticated patent-law firms are not every workplace, and a custom drafting assistant is not every coding or tutoring product. Model upgrades, different incentives and longer follow-up can change the effect. The related news article keeps those limits beside the result. A randomized study earns confidence in a bounded comparison rather than a universal conclusion.
The practical reading habit is to identify five things: the assignment, the control, the measured outcome, who remained in the analysis and the population to which the conclusion is being extended. A randomized trial can turn a plausible story about AI into stronger causal evidence. Its rigor comes from specifying exactly what changed and exactly what was measured, while leaving broader claims to further experiments.
Donald Rubin, Estimating causal effects of treatments in randomized and nonrandomized studies (1974)
Key questions
What does random assignment remove from an AI-tool comparison?
What is an intention-to-treat effect?
Does a randomized trial prove that an AI improves long-term expertise?
Cite this
APA
Ground Truth. (2026, October 8). Randomized trials: measuring what an AI tool actually changes. Ground Truth. https://groundtruth.day/learn/randomized-trials-and-treatment-effects.html
BibTeX
@misc{groundtruth:randomized-trials-and-treatment-effects,
title = {Randomized trials: measuring what an AI tool actually changes},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/learn/randomized-trials-and-treatment-effects.html}
}