Bizarus.
HomeNews › Thousands of papers certify that the authors had full access to the data. On one popular platform, that is not what happens.
Research World

Thousands of papers certify that the authors had full access to the data. On one popular platform, that is not what happens.

Journals routinely ask authors to certify they had full access to all the data in a study. Three researchers argued in Retraction Watch on 25 August that for the thousands of papers built on the TriNetX dashboard, that certification cannot be honoured: coding, cleaning and analysis decisions are made by the vendor and cannot be inspected by authors, reviewers or readers. Their proposed remedies run from a new disclosure statement to retraction where access was misrepresented.

27 August 2026 Bizarus

Many journals require authors to sign a statement that they "had full access to all the data in the study" and take "responsibility for the integrity of the data and the accuracy of the data analysis". On 25 August, three researchers argued in a guest post for Retraction Watch that for a fast-growing class of papers, that statement cannot be true.

The authors are Stephen Rhodes, a biostatistician at the Urology Institute at University Hospitals Cleveland Medical Center; Fredrick R. Schumacher of Case Western Reserve University; and Jonathan E. Shoag, of University Hospitals Cleveland Medical Center and Case Western Reserve. Rhodes discloses having co-authored papers using the platform before working through the argument.

The claim

TriNetX is a federated platform giving researchers access to de-identified electronic health records. Its dashboard returns aggregate counts, descriptive statistics and estimated quantities such as hazard ratios. Users of that interface never interact with line-level data.

The authors' point is not that the data are bad. It is that the decisions which determine what the data mean are made by the vendor and cannot be inspected:

All of this is proprietary. Neither authors, nor reviewers, nor readers can check it. And because records are continuously added, an analysis run on the platform is often irreproducible even by the person who ran it.

Line-level data can be requested and downloaded, and the authors say so. Their observation is that most of the studies they encounter, including in high-impact journals, use the dashboard.

Scale

The authors report that a PubMed search for "TriNetX" returned over 4,400 articles as of 19 August 2026, and that TriNetX describes itself as "the most cited real-world data source in peer-reviewed research". Science and Retraction Watch reported in a joint investigation that the dashboard is being used by medical trainees to produce fast and unreliable papers.

What they propose

Four things, in escalating order of consequence.

  1. Editors and reviewers should decide whether outsourcing analytic decisions to a dashboard is acceptable at all. The authors think not, on the grounds that line-level electronic health record data are available elsewhere.
  2. If it is acceptable, reporting standards must say so explicitly. They offer draft wording: "The authors did not have access to the data for this study. Cohort creation, descriptive statistics, and analyses were performed on the TriNetX platform... The code to implement patient selection and analysis is not available to the authors, reviewers, or readers as it is the intellectual property of TriNetX LLC."
  3. TriNetX, or the company employees involved, should be credited as non-author contributors under ICMJE criteria. The reasoning is precise: the consensus that a large language model cannot be an author rests on its inability to take responsibility. A company can.
  4. Published papers containing incorrect statements about data access should be corrected. Where editors judge that authors misrepresented their access, the papers should be retracted.

Analysis, and one caution

This is an opinion piece by named researchers, not a study, and the strongest version of its recommendation, retraction, would apply to papers whose findings may be perfectly sound. That is worth stating plainly before agreeing with the diagnosis.

The diagnosis, though, is hard to argue with, and it generalises well beyond one company. An attestation is a control only if it is checkable and only if signing it falsely has a cost. "I had full access to all the data" has become, for a large and growing literature, a sentence nobody verifies and many authors cannot honour. A control that is universally signed and never tested is not a weak control; it is a source of false confidence, which is worse than none, because a reader treats the signed statement as evidence of a check that did not occur.

The parallel the authors draw with large language models is the most useful part. The field settled quickly on the rule that authorship requires accountability, which is why a model cannot hold it. Apply the same test to dashboard-based analysis and the answer falls out: someone made the coding and cleaning decisions, that someone is identifiable, and they are currently named nowhere.

The commenters on the original post extend it usefully. One former TriNetX employee lists specific platform constraints, including that propensity matching is done on three-digit ICD-10 groupings rather than clinical definitions and that users cannot change the matching caliper or method. Another notes the same problem in image databases such as the Human Protein Atlas, where a figure can be lifted into a paper with no independent validation of the antibody or the patient characteristics behind it.

What to watch

Whether any journal or editorial body responds with a reporting requirement. The comparable AI-disclosure norms moved from a Retraction Watch argument to ICMJE guidance in roughly two years; the mechanism here is the same, and the fix is a disclosure field, not new technology.

Research IntegrityResearch MethodologyEvidence & Evaluationreal-world evidenceelectronic health recordsTriNetXdata accessreproducibilityauthorshipICMJEdashboard-based research
← All News
© 2026 Bizarus AI