CEO of AI company loses paper on cognitive impact of generative AI

A psychology journal has retracted a paper on how using generative AI can affect confidence in work tasks after its sole author refused to share data with the editors. 

The April paper, published in Technology, Mind, and Behavior, which is published by the American Psychological Association (APA), purported to describe the results of a survey of 1,923 adults recruited online to perform AI-assisted tasks. The paper was featured in an APA press release and articles in TIME and Futurism

As we wrote in June, experts in the field began flagging issues with the paper’s data on PubPeer shortly after it was published. Sandra Grinschgl, who studies human-technology interaction at the University of Bern in Switzerland, was the first to do so, and later enlisted the help of colleagues Ian Hussey and Malte Elson. Their concerns included inconsistencies in graphs and results, study design, “erroneous” references and a lack of ethics approval. 

Continue reading CEO of AI company loses paper on cognitive impact of generative AI

Guest post: The uncomfortable truth an explosion of papers reveals about authors’ access to data

wildpixel/iStock

The involvement of independent academic researchers in examining data and conducting analyses is central to producing trustworthy research. Because of this, many journals require authors to sign a statement acknowledging they “had full access to all the data in the study,” and take “responsibility for the integrity of the data and the accuracy of the data analysis.” Indeed, the International Committee of Medical Journal Editors (ICMJE) recommendations for authorship note that authors should be “accountable for all aspects of the work.” 

A rapidly increasing number of published articles rely on analyses produced via the TriNetX platform. However, TriNetX users don’t get full access to the information in the platform, muddying the truthfulness of any analyses of the data and compromising the integrity statements offered by authors and accepted by journals.

Instead, users see aggregate counts, descriptive statistics and estimated quantities, such as hazard ratios, without ever interacting with line-level data. And because the underlying data are constantly changing as new records are added, analyses run on the platform are often irreproducible, even by the researchers who originally conducted them.  

Continue reading Guest post: The uncomfortable truth an explosion of papers reveals about authors’ access to data

Kaggle removes problematic stroke dataset for copyright infringement 

Images from the now-deleted dataset on Kaggle purportedly showed people who had a stroke, but the folder included stock images and photos of politicians and celebrities, as seen in this screenshot.

The online data repository Kaggle has removed a problematic dataset for violating intellectual property rights months after it was flagged by sleuths for containing celebrity photos and lacking information on data provenance or confirmed medical diagnoses. 

Researchers had used the images to train machine learning algorithms to detect stroke, but the dataset instead contained images of people with Bell’s palsy — as well as pictures of celebrities from movies and on the red carpet, as we reported in May. Adrian Barnett, a medical statistician at Queensland University of Technology in Australia, flagged the dataset after discovering two others he and his colleagues investigated in a February preprint that showed signs of being fabricated. 

Using a reverse image search, Barnett matched four of the photos in the dataset, called “Facial droop and facial paralysis image,” to a paper published years earlier in the Journal of the Canadian Dental Association (JCDA). Only after prompting the journal to complain to Kaggle did the company take down the dataset. 

Continue reading Kaggle removes problematic stroke dataset for copyright infringement 

Critics of birdsong study fight to be named in Nature’s retraction

A zebra finch in New South Wales, Australia. Source: JJ Harrison/Wikimedia Commons (CC BY-SA 4.0)

Researchers who flagged methodological issues in a paper on birdsong a year and a half before Nature retracted it say they should be credited in the editorial notice. But the editors have refused, with one telling the critics the paper was retracted for unrelated reasons.

The March 2024 study at the center of the dispute looked at how sexual selection may drive song patterns in male zebra finches. Nature retracted the paper last month because two of the synthetic song pairs used in the study were found to be unreliable, according to the notice. All three authors agreed to the retraction. 

Todd Roberts, the paper’s corresponding author, told Retraction Watch the critics now asking for credit “prompted us to check the synthetic song pairs used in our paper.” He said his team did not do the reliability analysis of the pairs until after publication.

Continue reading Critics of birdsong study fight to be named in Nature’s retraction

Widely criticized keto diet study retracted

Aamulya/iStock

A 2025 paper claiming the keto diet does not promote the formation of arterial plaques has been retracted after widespread criticism of the study’s methods and claims. The journal found “the identified errors are too great to be corrected with a corrigendum,” according to the March 11 retraction notice.

In April 2025, JACC: Advances published the study, which looked at plaque build-up in 100 otherwise generally healthy people who had experienced an increase in their cholesterol levels while being on a keto diet. The study claimed scans performed one year apart by the company Cleerly showed the diet was not associated with the development of arterial plaques. 

This finding went against what previous studies had found, and it led to what Wired called “a new war in the nutrition world.” 

Continue reading Widely criticized keto diet study retracted

How the media hypes “research that is absurd on its face”

Aaron Brown says his new book, Wrong Number: How to Extract Truth from a Blizzard of Quantitative Disinformation, “isn’t an exposé of fraud—Retraction Watch covers that ground. It’s about legitimate-looking research that is absurd on its face.” 

Published this month by Wiley, Brown uses dozens of case studies to show “why widely reported and influential studies in top journals are not just wrong, but obviously and egregiously illogical or contrary to simple fact. My focus is less on the policy and statistical errors than on why no one seems to care,” he says.

Brown is a risk manager working in hedge fund management. He also teaches statistics at New York University and the University of California San Diego and writes columns for Reason and Bloomberg, among other outlets. We asked him to tell us more about how he thinks about the nexus of science, journalism and the publish-or-perish system that also pushes researchers to engage with non-experts to promote their work.

Continue reading How the media hypes “research that is absurd on its face”

‘Comically bad’ datasets used to train clinical models for stroke and diabetes 

A dataset on Kaggle purportedly showing people who have had a stroke includes images of Sylvester Stallone from Rambo and other celebrities. Source

Scrolling through an online image dataset, Adrian Barnett, a statistician at the Queensland University of Technology in Australia, pointed out a few familiar faces. Sylvester Stallone as Rambo, and then again on the red carpet. “This is just ridiculous,” Barnett said. George Clooney, Angelina Jolie and Daniel Craig all appear more than once, often with the same image. “You can see,” Barnett said, “this is just a comically bad dataset.”

This particular dataset, collected in a folder titled “droopy” and hosted on an open-source repository called Kaggle, underpins a paper published in Scientific Reports – not as a find-the-celebrity game, but as a training set for a predictive clinical model for early detection of strokes. 

The paper is the most recent example of a much wider problem that Barnett and his Ph.D. student Alexander Gibson have documented with Kaggle, which is owned by Google and hosts datasets uploaded by users that researchers and machine learning practitioners can use to build predictive models. By examining two other Kaggle datasets on stroke and diabetes, both of which included tabular patient data, Gibson and Barnett traced how the data move through the scientific literature and in some cases, into clinical use. Their work, described in a preprint posted to medRxiv in February, already has led to several retractions of the papers using these dubious datasets. 

Continue reading ‘Comically bad’ datasets used to train clinical models for stroke and diabetes 

Are AI chatbots infiltrating online survey data? Not yet, says new study

Mohamed Nohassi/Unsplash

Despite concerns some have raised about potentially compromised data, AI chatbots aren’t yet completing online research surveys widely, according to a new preprint. 

The authors of the study, posted earlier this month on PsyArXiv, found that fewer than 1% of around 4,800 survey responses collected by 12 different companies contained text that was likely not written by a human. Among 400 responses from a 13th company, however, around 16% were flagged for possibly being completed by a chatbot. 

The study used a novel detection tool created by the survey research company Prolific, which funded the project.

Continue reading Are AI chatbots infiltrating online survey data? Not yet, says new study

Bloodhound code sniffs out copied-and-pasted numerical data

Pexels

Markus Englund, a software developer and sleuth based in the Netherlands, first hit paydirt with invasive plant species in China. After having scanned 12 other published scientific datasets with his novel detection software with no results, he came across one showing something suspicious: rows and rows of measurements of plant roots repeated across entirely different species. 

“I was really excited,” he said in a recent call with Retraction Watch. “I couldn’t think of any innocent explanation for why that would be the case.” 

Englund had built a tool dedicated to “purging” fabricated data by identifying “impossible” data in spreadsheets available on open repositories, according to Science Detective, his site about the initiative. From his initial review, he has found 18 datasets containing duplicated values that are possibly serious enough to need correcting — including one from an influential paper on Parkinson’s disease, as The Transmitter recently reported. (Retraction Watch’s cofounder Ivan Oransky is that publication’s editor-in-chief.)

Continue reading Bloodhound code sniffs out copied-and-pasted numerical data

BMJ retracts cardiac stem cell paper, removes authors months after sleuths flag data ‘mismatch’

The BMJ has retracted a paper on stem cell therapy for heart failure after sleuths flagged the work for “serious” inconsistencies in data.

Published in October, the paper reported the results of a phase III clinical trial of more than 400 patients in Shiraz, Iran, looking at whether stem cell therapy lowers the risk of heart failure after a heart attack. The journal announced the results in a press release, and news of the findings appeared in several outlets. New Scientist called the study the “strongest evidence yet that stem cells can help the heart repair itself.”

A week after the study was published, sleuths took to PubPeer to point out inconsistencies between the data reported in the article and the dataset uploaded with it. The concerns included a “curious repeating pattern” of records in the dataset and a high number of integers for the height and weight of patients. 

Continue reading BMJ retracts cardiac stem cell paper, removes authors months after sleuths flag data ‘mismatch’