Bizarus.
HomeResearch › 63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins
Evidence Review

63 studies on LLMs in systematic reviews: strong agreement on screening, much weaker where judgment begins

The first systematic review to measure how well large language models perform the actual tasks of a systematic review pooled 63 studies and 148 separate performance assessments. Agreement with human screeners was high, a median positive percent agreement of 0.92 for title and abstract screening. On risk of bias assessment, median accuracy fell to 0.62. The spread around every one of those medians is wide enough to change what a researcher should do with them.

14 August 2026 Bizarus

A systematic review is not one task. It is a sequence of them: formulating the question, building the search, screening titles and abstracts, screening full texts, extracting data, appraising risk of bias, synthesising the result. Asking whether a language model "can do systematic reviews" collapses that sequence into a question with no useful answer. Florian Laignelot, Guillaume Martin and colleagues asked the tractable version instead. Across every study that has tried to measure it, how closely does a language model's output on each of those tasks agree with what human reviewers decided?

Why this matters is a workload problem that has been visible for years and is getting worse. The authors cite a 2016 estimate of 67.3 weeks and roughly five reviewers to complete and publish a single systematic review, against more than 1.2 million publications entering MEDLINE annually. Reviews are the form of evidence clinicians and policymakers are told to trust most, and they are the form least able to keep pace with the literature they summarise. That gap is what makes automation attractive, and it is also what makes an honest measurement of the automation urgent rather than merely interesting.

What the researchers did. They searched PubMed, Embase and the Cochrane Library from inception to 14 January 2025, plus preprint platforms oriented to health research, registered the protocol in advance on PROSPERO (CRD42024594795), and reported against the PRISMA extension for diagnostic test accuracy studies. Pairs of reviewers independently screened and extracted, with a senior reviewer resolving disagreements. Because the purpose-built appraisal tool for AI diagnostic accuracy studies, QUADAS-AI, was still under development, they adapted QUADAS-2 and documented every modification. From 5,408 items and 3,889 unique references they included 63 studies, 11 qualitative and 52 quantitative, among them six preprints and six conference abstracts. The 52 quantitative studies yielded 148 separate performance assessments.

One methodological decision deserves attention before any number does. The authors reported positive and negative percent agreement, PPA and NPA, rather than sensitivity and specificity, following US Food and Drug Administration guidance. The reason is precise: sensitivity and specificity imply a comparison against a reference standard that is taken to be correct. Human reviewers are not that. What these studies measure is concordance between a model and a fallible comparator, and the authors chose a vocabulary that says so. Heterogeneity across metrics and settings was substantial enough that they declined to meta-analyse, reporting medians and interquartile ranges instead.

What they found. Attention has been unevenly distributed across the review workflow. Title and abstract screening accounted for 78 of the 148 assessments (53%), data extraction for 23 (16%), full-text screening for 20 (14%) and risk of bias assessment for 13 (9%). GPT models accounted for 114 assessments (77%). Nobody, in this evidence base, has evaluated protocol drafting or review updating at all.

For title and abstract screening, median PPA was 0.92 (IQR 0.69 to 0.98) and median NPA was 0.89 (IQR 0.72 to 0.95). For full-text screening, median PPA was 0.93 (0.87 to 1.00) and median NPA 0.92 (0.78 to 0.97). Model generation mattered, and it mattered asymmetrically: for title and abstract screening, late-generation models (GPT-4 and later) reached a median NPA of 0.91 against 0.58 for earlier models, while PPA was similar across generations. In full-text screening the same pattern was sharper still, median NPA 0.95 against 0.54. Newer models did not get better at recognising what belonged. They got dramatically better at recognising what did not.

Performance on the interpretive tasks looks different. Data extraction accuracy ranged from 0.36 to 1.00 with a median of 0.95 across 11 assessments. Risk of bias assessment ranged from 0.44 to 0.90, with a median accuracy of 0.62 across 6. Identifying novel research questions ranged from 0.68 to 0.83 across 3. Two assessments of search equation construction returned 0.78 and 0.89.

What this does NOT establish. It does not establish accuracy. It establishes agreement, and the authors are explicit that this was the point of using PPA and NPA. High agreement with a comparator that is itself unreliable is a weaker claim than it sounds, and the review supplies the evidence for that caution in its own discussion: single-reviewer abstract screening has been shown to miss up to 13% of eligible studies, human data extraction error rates as high as 50% have been reported, and agreement between reviewers at the meta-analysis level has been measured as low as 31%. This cuts in both directions. It makes model parity less impressive than it first reads, and it also makes model parity less alarming.

The medians should not be read without the spread. An IQR of 0.69 to 0.98 on screening PPA means a quarter of assessments fell below 0.69 positive agreement. In screening, low PPA is the failure mode that actually damages a review, because it means eligible studies were dropped and will never be seen again. A data extraction median of 0.95 sitting above a floor of 0.36 across 11 assessments describes variance, not reliability. A researcher choosing whether to delegate a task needs the worst plausible case, not the middle one.

The finding that web-based interfaces outperformed API implementations should be treated with particular care, because the authors themselves report a confound. Risk of bias was rated higher in studies using APIs, and they suggest those studies may have been run predominantly by data scientists and engineers less familiar with methods for minimising bias. The comparison is not clean, and it should not be read as evidence that the chat interface is the better tool.

The underlying evidence is also not strong. Risk of bias was high in at least one domain for 13 of the 52 quantitative studies (25%) and unclear in at least one domain for 30 (57%). Only 10 studies (19%) were low risk across all domains. Where a study made repeated assessments of the same model, task and dataset, the authors selected the one with the highest PPA, a choice that introduces optimism bias; sensitivity analyses using random selection instead were reported as consistent, but the headline figures carry that decision. The scope is biomedical, with technical and preclinical preprint servers deliberately excluded, so nothing here transfers automatically to a review in law, education or the social sciences.

Bizarus interpretation. The most useful output of this review is not a number but a gradient. Performance holds up where the task is retrieval and matching against explicit criteria, and degrades where the task requires interpretation and judgment. A median of 0.92 on screening against 0.62 on risk of bias is not a small difference of degree; it is the boundary between a task that can be supervised and a task that cannot safely be delegated at all. That gradient is directly actionable in a way the aggregate figures are not.

The second observation is about timing, and it is unusually self-referential. The review's clearest positive finding is that model generation drives performance. Its search closed on 14 January 2025 and it published in June 2026. By the paper's own logic, its evidence base is dated in precisely the dimension it identifies as decisive. The authors acknowledge this and argue the review works as a baseline reference against which later models can be read, which is a fair defence of what they have produced. It is also a warning about how the figure will be used. Anyone citing 0.92 as the current screening performance of a current model is citing something this study did not measure.

The third is a gap that ought to unsettle anyone currently evaluating research software. The authors found no study assessing tools purpose-built for systematic reviews. The evidence describes general-purpose models being prompted by researchers. It says nothing about the products being marketed to research teams on the strength of exactly this kind of finding.

The authors suggest that late-generation models achieved performance suggesting they could feasibly substitute for one of two human reviewers in screening. That is the operationally significant claim in the paper, and it is offered as a suggestion rather than a result, which is the correct register for it. It is worth being clear why. Dual independent screening exists because a single reviewer is unreliable, and its value comes from two error patterns being partly independent of each other. The review reports assessments of model output against human decisions; it reports no assessment of a model-and-human pair measured against a human-and-human pair, which is the design that would actually test the substitution. Until something does, the substitution is a plausible inference from agreement data, not a validated workflow.

What remains unanswered. Reproducibility is the largest open question and the one the authors press hardest. Language models are not deterministic, the same prompt can return different outputs across sessions and backend updates, and model versions and prompt details were scarcely reported in the included studies. Very few studies examined why models err, whether performance varies by clinical domain, or how sensitive outputs are to ambiguous information, so the error mechanism remains unexamined even where the error rate is known. Retrieval-augmented generation and source-grounded prompting are named as plausible mitigations for hallucination, but the review found no included study reporting their use. There are no shared, well-annotated benchmarking datasets, which is why comparison across studies and across time is currently so difficult. And the two tasks where performance was weakest, risk of bias assessment and data extraction, are also among the least studied, at 13 and 23 assessments respectively. The thinnest evidence sits exactly where the risk is highest.

Source record. Laignelot, F., Martin, G. L., Ossman, M., Pingeon, O., Boubaker, A., Picovschi, E., Kim, J., Tannier, X., Cohen, J. F., Dechartres, A. Large language models show promising performance for some systematic review tasks but call for cautious implementation: a systematic review. Journal of Clinical Epidemiology 194, 112221 (June 2026). DOI: 10.1016/j.jclinepi.2026.112221. Open access under a Creative Commons licence. Protocol registered as PROSPERO CRD42024594795. Journal quartile verified on SCImago: Journal of Clinical Epidemiology, Epidemiology category, Q1 (2025, and Q1 in every year from 1999 to 2025). Reported funding: none. Declared competing interest: one author is an employee of Synapse Medicine, a company developing clinical decision support software and an AI system for synthesising medical guidelines; the remaining authors declared none. Within the included studies, the review reports that no study was funded by an LLM developer, and that one author in one included study was affiliated with an LLM developer and held stock in that company.

AI & ResearchResearch MethodologyEvidence & Evaluationsystematic reviewsscreeningevidence synthesislarge language models
← All Research
© 2026 Bizarus AI