Back to search

Article

Leveraging a Vision-Language Model with Natural Text Supervision for MRI Retrieval, Captioning, Classification, and Visual Question Answering

2025-02-20

Abstract excerpt

Large multimodal models are now extensively used worldwide, with the most powerful ones trained on massive, general-purpose datasets. Despite their rapid deployment, concerns persist regarding the quality and domain relevance of the training data, especially in radiology, medical research, and neuroscience. Additionally, healthcare data privacy is paramount when querying models trained on medical data, as is trans...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
01bbcc3c-ccef-5f43-b1c5-d9ef95d3685a
DOI
10.1101/2025.02.15.638446
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Leveraging a Vision-Language Model with Natural Text Supervision for MRI Retrieval, Captioning, Classification, and Visual Question AnsweringDOI 10.1101/2025.02.15.638446
Select a neighboring publication to make it the new centre.