The Standing Order plate for 17 September 2026. Gold D major above the words 10 studies, and beneath them: 1,235 publications screened, 95 full texts assessed, 10 met eligibility; human plus AI, the pooled gain rests on 2 of them. Over a darkened photograph of the California coast at La Jolla.

10 studies: what the npj Digital Medicine meta-analysis of human-LLM collaboration can and cannot carry

The pooled gain of 4.88 points rests on 2 of the 10 studies. The prediction interval runs from minus 31.65 to plus 41.42, and it answers a different question.

A systematic review and meta-analysis of human-large language model collaboration in clinical medicine. PRISMA 2020, four databases searched through 28 June 2025. 1,235 records screened. 95 full texts assessed.

Only 10 studies met eligibility.

The pooled gain for human plus AI on composite diagnostic and management scores was 4.88 percentage points, 95% CI +0.65 to +9.12. It is the figure that will be quoted in every procurement deck this quarter. It is statistically significant but it rests on 2 of the 10 studies.

The prediction interval runs from minus 31.65 to plus 41.42. It is the figure that will not be quoted, and it is not a more pessimistic version of the first one. It answers a different question.

A confidence interval asks how precisely we have located the average of the studies we already have. Add participants and it narrows. A prediction interval asks what a new study, in a new setting, would find. It carries the uncertainty in that average plus the real spread between settings, and that second part does not shrink as studies accumulate. It gets measured better, but it stays, because it is variation in the world rather than noise in the estimate. Different clinicians, different case mix, different workflow, different prompt.

With 2 studies, that spread rests on a single degree of freedom. Which is why the honest reading is not that a hospital risks losing 32 points. It is that 2 studies support no prediction at all about a third setting. The width is a measure of how little is known.

Two findings sit inside the paper that the abstract does not carry, and they cut in opposite directions. In the three-arm studies the synergy ratio was approximately 1: no statistical evidence that human plus AI outperforms the AI working alone. And in electrodiagnostic reporting, AI alone scored 0.70 against 0.94 for both the clinician and the pair. The evidence is task-dependent and it supports no general claim in either direction.

7 of the 9 randomised trials carried some concerns for risk of bias. The one study embedded in a real clinic showed a near-null effect on time where the simulations did not, which the authors read as efficiency gains that may not survive contact with a real workflow.

And one structural fact that deserves stating plainly. 3 of the 10 studies come from a single research group, and both of the 2 behind the headline number are theirs, recruited from the same 3 academic centers under identical eligibility criteria, with no statement anywhere of whether the same physicians appear in both.

The authors’ own recommendation is not a conclusion about AI. They ask for preregistered, pragmatic, multicenter trials embedded in real workflows, with harms reported alongside benefits. That is a description of infrastructure that does not yet exist.

So the evidence base for human-AI collaboration in clinical medicine is 10 studies wide. That is not an argument for or against anything.

It is a statement about how much weight the literature can currently bear, and the answer is less than the discussion around it assumes.

We are arguing about the destination while the evidence is still asking for a road.


Source. Wang, Zhang, Jiang et al. Human-LLM collaboration in clinical medicine: a systematic review and meta-analysis. npj Digital Medicine, 28 January 2026. PROSPERO CRD420251068272. doi:10.1038/s41746-026-02382-2


The Standing Order

Published Tuesday, Wednesday and Thursday on nonalgorithmic.com, and nowhere else. The essays arrive each Friday.

Discover more from Nonalgorithmic

Subscribe now to keep reading and get access to the full archive.

Continue reading