When open data are misused

Credit: HALUSTD from Popspixel via Canva

When Dorothy Bishop was asked to review a paper for a small neuroscience journal in August, she could tell from the abstract the authors had used a dataset she had first published four years ago on the public repository OpenNeuro. The dataset contained raw functional MRI scans from 50 participants, taken while they completed language-related tasks such as thinking of words and matching words to pictures.

“I was curious that somebody had already analyzed it, because it’s a hefty dataset that would have taken a lot of work to analyze,” Bishop, an emeritus professor of developmental neuropsychology at the University of Oxford, told Retraction Watch. 

She quickly found problems in the manuscript. The description of the dataset didn’t match what she and her colleagues had actually done, and several references didn’t support the claims they were cited for. “There was clearly something seriously wrong,” she said. The editor was also suspicious, Bishop said, and the paper was eventually rejected. But Bishop believed the authors would try to publish the manuscript elsewhere, so she did a broader search of their names. 

Her search pulled up two papers from this year in the Journal of Neurolinguistics, an Elsevier publication. Both of them purport to use brain imaging data from OpenNeuro, but Bishop and Kathy Rastle, a psychologist at University of London, whose data appear in one of the papers, believe both were generated by AI. 

“I was flabbergasted that someone could take the time to generate a paper that I believe is fake,” Rastle told us. 

Although the papers look convincing, the authors don’t appear to have analyzed the data “beyond perhaps downloading a little bit of it to produce a figure,” Bishop said, and they appear to contain fundamental errors. Bishop has asked the journal to retract the articles, and Elsevier says it is investigating. 

The first paper, published in May and coauthored by Afzal Khan and Dawood Khan of King Saud University in Saudi Arabia and Iqra University in Pakistan, purportedly used brain scans from a public dataset created by Rastle and her colleagues. “When we looked at the paper, it soon became clear that they probably hadn’t used our data at all,” Rastle told us. 

The paper stated behavioral data were only available for some of the participants, but they were available for everyone, Rastle said. It also reported on reading comprehension scores that the dataset doesn’t contain. “We’re not sure how they generated the statistics for these variables that didn’t even exist,” she said. A diagram purporting to show the brain regions involved in reading was also wrong. “It just bore no resemblance to what we know about reading in the brain,” Rastle said. “Anybody with expertise would have seen that.” 

The paper’s acknowledgements thanked Rastle and her colleagues for sharing their data. “I don’t believe that they actually used our data,” Rastle said. “I believe that they said they did in order to create a veneer of respectability.”

The second paper, published in August by Afzal Khan alone, claimed to analyze four OpenNeuro datasets. Bishop raised concerns about it on PubPeer, and told us Khan “had a total fiasco attempting to respond.” The number of participants included in the paper did not match the datasets — Khan responded that he had removed participants with incomplete scans, although the number he claims to have removed still doesn’t match the original data, per Bishop’s response on PubPeer.

One graph also shows a participant had used a second language for more than 50 years, despite the maximum value being 37 in the data he claimed to have analyzed. In his responses, Khan insisted the graph does not show a value above 50. “The reported results remain valid as published,” he wrote on PubPeer. Khan eventually acknowledged that he would redo some of the analyses and said a correction was being processed by the journal. 

“This is clearly some sort of AI that knows enough about a dataset to produce a plausible paper that purports to use it,” Bishop said. “It looks superficially OK.” She said the slight inaccuracies in the numbers, or data that looked vaguely right but didn’t match up with the dataset, were signs AI was used to produce the paper. In the manuscript she reviewed, the code didn’t actually run or link up to the data, and the text contained subtle miscitations. In a blog post, Bishop suggests checking the data, code and references, the plausibility of the research, and whether the article makes sense, though concedes “the frauds are getting so good that these methods now seem inadequate.” 

In an email, Khan told us he was aware of the concerns and was “taking them seriously,” and “engaging directly and in good faith with the journal.” He said he could not comment on specifics while a review was ongoing. 

Bishop asked the editor of the journal to retract the paper, and in response, Elsevier told her that it would investigate, but that it “may take considerable time.” An Elsevier spokesperson confirmed to us that the papers are under investigation. Nai Ding, the editor of the journal, told us they had submitted the findings of the journal’s own inquiry to Elsevier, which is conducting further review. 

“When a paper is just so obviously deeply flawed, and probably a fake paper, it really shouldn’t take months to get it out of the scientific literature,” Rastle said. “It’s just polluting the scientific record and creating confusion amongst readers and other scientists.” 

Khan has a Ph.D. in linguistics and is affiliated with the English Language Skills Department at King Saud University in Riyadh. Before 2026, his publications were largely about linguistics and teaching English, as well as essays on literary criticism, not brain imaging.  “These are huge brain imaging data sets that take a lot of computing capacity and expertise to get out anything sensible from them,” Bishop said, noting there was “no evidence that this author had that expertise.”

The case is “very upsetting for those of us who are into open data,” Bishop said. Rastle pointed to the volunteers who took part in the original studies. “Participants allow their data to be shared openly to further scientific progress,” she said. 

“They don’t allow it to be shared openly to perpetuate fraud.”


Like Retraction Watch? You can make a tax-deductible contribution to support our work, follow us on X or Bluesky, like us on Facebook, follow us on LinkedIn, add us to your RSS reader, or subscribe to our daily digest. If you find a retraction that’s not in our database, you can let us know here. For comments or feedback, email us at [email protected].


Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.