Kaggle removes problematic stroke dataset for copyright infringement 

Images from the now-deleted dataset on Kaggle purportedly showed people who had a stroke, but the folder included stock images and photos of politicians and celebrities, as seen in this screenshot.

The online data repository Kaggle has removed a problematic dataset for violating intellectual property rights months after it was flagged by sleuths for containing celebrity photos and lacking information on data provenance or confirmed medical diagnoses. 

Researchers had used the images to train machine learning algorithms to detect stroke, but the dataset instead contained images of people with Bell’s palsy — as well as pictures of celebrities from movies and on the red carpet, as we reported in May. Adrian Barnett, a medical statistician at Queensland University of Technology in Australia, flagged the dataset after discovering two others he and his colleagues investigated in a February preprint that showed signs of being fabricated. 

Using a reverse image search, Barnett matched four of the photos in the dataset, called “Facial droop and facial paralysis image,” to a paper published years earlier in the Journal of the Canadian Dental Association (JCDA). Only after prompting the journal to complain to Kaggle did the company take down the dataset. 

The dataset page on Kaggle now shows a 404 error but no explanation for why it was removed. 

“When an enforcement action occurs, the page is made unavailable on Kaggle and the uploading user receives a direct notification detailing the reasoning alongside instructions on how to appeal the decision,” a representative for the company told us by email. They did not address a question about how users of the database are informed of a removal. 

The removal came months after Barnett tried unsuccessfully to convince Kaggle the widely used collection of stroke images was “scientific nonsense.”

Kaggle has removed datasets before for breaching their terms of service, but Barnett didn’t have any luck with his earlier attempts to flag the collection.  “I think it’s funny,” he said, “because we approached Kaggle and said, ‘Hey, this dataset is rubbish, scientific nonsense,’ and that didn’t work.” But breaching copyright seems to be “the quickest way to get some of this rubbish taken down,” Barnett said. “I am going to be using this approach now that I’ve seen that it’s successful.”

Kaggle removes datasets that violate community guidelines, including intellectual property rights, the representative said. “Our approach to platform integrity involves active moderation and reviewing daily content reports from the community to ensure adherence to our guidelines, including the protection of intellectual property rights,” they said.

Several of the images taken from the JCDA paper appear to be of a minor.  The authors “consented this young girl, and didn’t anticipate that the image was then going to be just shared all over the internet,” Barnett said. People who agree to share their medical images place trust in the process, and do it to inform other clinicians, “not to help somebody who wants to boost their ratings on Kaggle,” he said. 

Barnett alerted both the journal and the author of the paper about the use of the images, since they cannot be used without written permission from the Canadian Dental Association, according to the journal’s policy. Michelle Bergeron, a publications specialist at the journal, told Barnett in June JCDA was looking into their unauthorized use, according to emails seen by Retraction Watch. By July, the journal had contacted Kaggle directly about the copyright violation, and Kaggle had removed the dataset. 

Since we last reported on the problematic datasets hosted on Kaggle, several papers using them have been retracted, but more have since been published. The Scientific Reports paper we wrote about in May was retracted on June 29. The authors had used Kaggle images to train a stroke prediction model, but many of them were of people with Bell’s palsy, including those pulled from the JCDA paper. 

The retraction notice states the dataset underpinning the model “can no longer be verified through its original source” and appeared to contain identifiable faces without information about ethical oversight, nor whether their diagnosis was confirmed. The content of the article is no longer accessible “to protect the privacy of the individuals,” the editors wrote. But more researchers continue to use the dataset, including in at least two papers published this year. 

Since our reporting, two other papers in Scientific Reports  that had been under investigation for using the flagged datasets have also been retracted. The notices state the works relied on publicly available datasets whose “provenance and validity cannot be confirmed.” One of these, the “Stroke Prediction Dataset,” which includes quantitative health parameters such as BMI and smoking status, remains available online and continues to be used to train clinical prediction models. 

In early July, authors of another preprint justified using it to train an AI for stroke prediction because it is “widely adopted” in machine learning research, saying it had been used in more than 100 research articles — a figure they cited from Barnett’s work calling out the dataset for its unreliability. 

Barnett and his team are now planning a broader audit of 100 datasets from Kaggle to assess their quality and whether they are suitable for medical research. The goal is to determine: “is this an enormous problem, or were we just unlucky and came across a few odd ones?” he said. 

“It never occurred to me that I’d have to ask the question: Is this real?” Barnett said, “But now, I think that’s the question I’m going to be asking for the rest of my career.”


Like Retraction Watch? You can make a tax-deductible contribution to support our work, follow us on X or Bluesky, like us on Facebook, follow us on LinkedIn, add us to your RSS reader, or subscribe to our daily digest. If you find a retraction that’s not in our database, you can let us know here. For comments or feedback, email us at [email protected].


Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.