Guest post: The uncomfortable truth an explosion of papers reveals about authors’ access to data

wildpixel/iStock

The involvement of independent academic researchers in examining data and conducting analyses is central to producing trustworthy research. Because of this, many journals require authors to sign a statement acknowledging they “had full access to all the data in the study,” and take “responsibility for the integrity of the data and the accuracy of the data analysis.” Indeed, the International Committee of Medical Journal Editors (ICMJE) recommendations for authorship note that authors should be “accountable for all aspects of the work.” 

A rapidly increasing number of published articles rely on analyses produced via the TriNetX platform. However, TriNetX users don’t get full access to the information in the platform, muddying the truthfulness of any analyses of the data and compromising the integrity statements offered by authors and accepted by journals.

Instead, users see aggregate counts, descriptive statistics and estimated quantities, such as hazard ratios, without ever interacting with line-level data. And because the underlying data are constantly changing as new records are added, analyses run on the platform are often irreproducible, even by the researchers who originally conducted them.  

A PubMed search for “TriNetX” returns over 4,400 articles (as of August 19, 2026) and TriNetX has recently celebrated reaching over 2,000 citations, describing itself as “the most cited real-world data source in peer-reviewed research.” While it is possible to request line-level data for download from TriNetX, the majority of studies we encounter, including those in high impact journals, use the dashboard interface. A recent article by Retraction Watch and Science highlighted the use of this dashboard by medical trainees to produce fast, but unreliable papers

What has received less attention is what dashboard-based research means for author responsibility and accountability. In our opinion, lack of access to the underlying line-level data and outsourcing of the analysis to a proprietary dashboard is a major deviation from standard research practices. This is because it is not possible for these authors to take “responsibility for the integrity of the data and accuracy of the analysis.” Important decisions, such as mapping different medical coding systems onto one another, cleaning data (e.g., dealing with clinical events that occur outside of their associated encounter, identifying inconsistent diagnoses, etc.) and writing analysis code are outsourced to a company.

While standardizing some aspects of electronic health record data processing might be desirable, the specific decisions the TriNetX platform makes cannot be checked or verified by authors and readers alike, as the software is proprietary.  

Allowing a for-profit enterprise to make important analytic decisions at scale in the literature would typically be a cause for concern and would motivate measures to mitigate bias. TriNetX likely has no stake in the outcome of individual analyses that are run on their platform, other than perhaps an interest in research that generates headlines. Just as new standards and norms have had to be established given the rising use of LLMs in research and writing, we need norms regarding dashboard-based analysis.  

First, editors and reviewers need to decide whether authors should be permitted to outsource important analytic decisions to a dashboard. Given that line-level electronic health records data are available via other sources (e.g., Epic Cosmos, All of Us), we would suggest the answer is no. One could argue that the sheer number of articles using this platform is proof of its acceptability. But it’s not clear that reviewers and editors are generally aware of the extent that the TriNetX platform deviates from normal research practice, especially when articles come with misleading statements regarding data access. It is certainly not apparent to readers. 

If this kind of dashboard-based research is acceptable, we would argue that reporting standards need to be updated to more clearly state how the data were handled and the analysis was performed. This should be transparent to reviewers, readers and editors. 

This would require attestations different from those many journals currently require. We suggest a disclosure along the lines of: “The authors did not have access to the data for this study. Cohort creation, descriptive statistics, and analyses were performed on the TriNetX platform, a graphical user interface that provides access to aggregate data. The code to implement patient selection and analysis is not available to the authors, reviewers, or readers as it is the intellectual property of TriNetX LLC. Exports from the TriNetX platform showing complete cohort definition and results are provided in supplementary material.” 

We further propose that TriNetX, or representatives of the company involved in the analysis, should be acknowledged as non-author contributors to an article, consistent with ICMJE criteria. While the consensus has emerged that LLMs cannot take credit on articles because they cannot take responsibility, a private company or its employees certainly can. 

Finally, published articles containing incorrect statements regarding the level of access that authors had to data should be corrected, perhaps to include a more accurate statement like the example above. If journal editors determine the authors have misrepresented their level of access to data and responsibility for analyses, those papers should be retracted. 

Stephen Rhodes, Ph.D., is a biostatistician for the Urology Institute at University Hospitals Cleveland Medical Center. He co-authored articles using the TriNetX platform prior to thinking through the issues raised here. 

Fredrick R. Schumacher, Ph.D., is an Associate Professor in the Department of Population and Quantitative Health Sciences, School of Medicine at Case Western Reserve University. 

Jonathan E. Shoag, M.D., is an Associate Professor in the Department of Urology at University Hospitals Cleveland Medical Center and the Case Western Reserve University School of Medicine.


Like Retraction Watch? You can make a tax-deductible contribution to support our work, follow us on X or Bluesky, like us on Facebook, follow us on LinkedIn, add us to your RSS reader, or subscribe to our daily digest. If you find a retraction that’s not in our database, you can let us know here. For comments or feedback, email us at [email protected].


One thought on “Guest post: The uncomfortable truth an explosion of papers reveals about authors’ access to data”

  1. Agree 100% this is a hole in current integrity standards.

    A similar example is the use of image databases such as Human Protein Atlas (HPA). It’s very easy for authors studying a particular protein in some type of cancer to just pull up images of protein expression in normal vs. cancer samples and just slap it in their paper and refer to HPA with no independent validation (was the antibody specific, what were the patient characteristics, etc.) It’s essentially a free extra figure in many papers, imparting a veneer of clinical relevance on mostly cell-based data sets.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.