New approach developed to detect prostate cancer with fewer labeled images
A team of iHEALTH researchers tested self-supervised learning strategies that improve the performance of artificial intelligence (AI) models for detecting and segmenting prostate cancer in magnetic resonance imaging, even when very little labeled data is available.
The study, titled "Dissecting self-supervised learning strategies for transfer learning in MRI prostate cancer diagnosis," was led by Jorge Facuse and Diego Campanini, together with Denis Parra, Cecilia Besa, Rodrigo Salas, Pablo A. Estévez and Domingo Mery, iHEALTH researchers affiliated with the Pontificia Universidad Católica de Chile, the Universidad de Chile and the Universidad de Valparaíso.
Published in Scientific Reports (Nature group), the work focuses on a central problem for the application of medical AI in Chile and Latin America: the underrepresentation of these populations in the large datasets that feed these models. This is because most of the available prostate data comes from developed countries in the northern hemisphere —Germany, the Netherlands, Norway— and existing AI models have been trained and validated on that data.
"Chile and Latin America in general are an underrepresented region when it comes to medical data. This raises concerns about the ability of these models to generalize, their robustness, and the possible biases they may have when applied to populations different from those they were trained on," the team explains.
The challenge is twofold: although efforts have been made to create and label local datasets, the clinical workload of medical specialists means these tend to be small, on the order of hundreds of labeled scans. The question that motivated the research was whether such small datasets can serve to adapt and improve existing AI models for our population.
The proposal: learning from "unlabeled" images
To address the problem, the researchers could have created a dataset annotated with severity levels and segmentations of Chilean patients, which would have taken several years and budget to label. Instead, the team used the PICAI dataset —public, unlabeled data from patients in hospitals in the Netherlands, Norway and Germany— to pre-train the model using SSL techniques.
Then, to test the ability of these techniques to generalize to smaller-volume data, they evaluated their performance on labeled datasets of much smaller size, such as Prostate-158 and ChiPCa. The latter dataset was collected from 111 studies of Chilean patients from the UC CHRISTUS Health Network.
As the team explains: "One of the main advantages of learning from unlabeled data is that you don't need a specialist labeling hundreds or thousands of data points. In general, obtaining good medical labels requires a specialist with years of experience, who is not always available. With self-supervised learning, the model learns from the data itself without labels, which allows the AI model development workflow to be accelerated and enables better transfer of learning to the target dataset, even with a small amount of labeled data."
The proposal consists of a three-stage training strategy: self-supervised pre-training on unlabeled images, supervised training on a medium-sized dataset, and a final fine-tuning on the smaller target dataset. The researchers systematically tested different architectures (UNet and UNETR) and SSL methods, evaluating how each decision influences performance.
The most relevant result is that, even with a very small labeled target dataset —around 100 studies—, if it is of good quality, it is possible to improve the performance of models previously trained with populations very different from those of Latin America.
"This research is an important step in the sense that it shows that even with a very small labeled dataset, if it is of good quality, it is possible to improve the performance of existing models trained with populations very different from those of Latin America or other underrepresented regions," the authors note.
This opens the door for hospitals without large annotated databases to use images from other medical centers for pre-training and then carry out the final training on labeled databases on the order of a hundred studies.
How close is this technology to becoming a real support tool in a Chilean hospital or clinic? The team is cautious on this point:
"It is difficult to determine how close the adoption of these algorithms is in Chilean hospitals or clinics, since it depends not only on the accuracy of the models but also on regulatory aspects. In terms of accuracy, these models must be validated on other labeled, multi-center datasets. For their application in real-world settings, more testing on larger labeled datasets is required; moreover, the data distribution may vary depending on the equipment used and the medical centers where they were acquired."
Among future lines of work, the team especially highlights the potential of combining images with other sources of information: "Obtaining a large amount of data is not a trivial task, since it requires authorization from the medical centers, patients and universities involved; in addition, the computational infrastructure of medical centers is not always designed to support a massive data extraction. However, something we can explore is how to align existing data with other modalities, such as clinical text. Our current and future work seeks to exploit multimodality with existing and local data."
The approach could also be extended to other health problems. The lesion segmentation strategy evaluated in the paper can be applied to other medical imaging modalities where a relevant finding needs to be located, although the main limitation remains access to quality labels produced by specialists.
Facuse, J., Campanini, D., Parra, D. et al. Dissecting self-supervised learning strategies for transfer learning in MRI prostate cancer diagnosis. Sci Rep 16, 21561 (2026). https://doi.org/10.1038/s41598-026-50355-y