Kaggle removes flawed stroke-image dataset after copyright complaint, five months after integrity flags were ignored
A widely used Kaggle dataset marketed for stroke-detection research contained celebrity photos and images lifted from a dental journal. Kaggle acted only after a copyright complaint, not the original research-integrity report.
The online data repository Kaggle has removed a dataset called "Facial droop and facial paralysis image," used by multiple research teams to train machine-learning models for stroke detection, after the Journal of the Canadian Dental Association (JCDA) complained that images in the set had been taken from one of its own papers without permission, Retraction Watch reported on 29 July 2026.
Adrian Barnett, a medical statistician at Queensland University of Technology, first flagged the dataset in February 2026 after a separate audit found signs of fabrication in two related Kaggle datasets. Using reverse image search, Barnett matched four photos in the set to a JCDA paper published years earlier; several depicted a minor whose guardians had consented to the images' use in that original clinical publication, not to redistribution as open training data. Many of the "stroke" images were later found to depict people with Bell's palsy, an unrelated condition, along with stock photos and pictures of celebrities and politicians.
Kaggle told Barnett in earlier exchanges that the dataset did not violate its terms of service. It was the journal's copyright complaint in July, not the research-integrity concern raised in February, that led to the takedown. "I think it's funny," Barnett told Retraction Watch, "because we approached Kaggle and said, 'Hey, this dataset is rubbish, scientific nonsense,' and that didn't work... breaching copyright seems to be the quickest way to get some of this rubbish taken down."
At least three Scientific Reports papers built on the dataset have since been retracted, citing an inability to verify the data's provenance. Retraction Watch also found the dataset, and a related "Stroke Prediction Dataset" that remains online, still in use in newer papers and preprints, including one that cited Barnett's own critique as evidence the dataset is "widely adopted."
Analysis: this is a clean case study in how platform moderation actually responds to different categories of complaint. A specific, documented research-integrity concern, raised directly to the platform for five months, produced no action; a single copyright claim from a rights-holder resolved the same problem within weeks. For anyone building or evaluating machine-learning models on public datasets, provenance verification cannot be outsourced to the hosting platform's own moderation, since the incentives that get bad data removed are not the same incentives that would keep it out in the first place.
What to watch: Barnett's team is planning a broader audit of 100 Kaggle datasets used in medical research, to determine, in his words, "is this an enormous problem, or were we just unlucky."
Source: Retraction Watch, https://retractionwatch.com/2026/07/29/kaggle-removes-problematic-stroke-dataset-for-copyright-infringement/