Back to search

Article

A flaw in using pre-trained pLLMs in protein-protein interaction inference models

2025-04-23

Abstract excerpt

<h4>ABSTRACT</h4> With the growing pervasiveness of pre-trained protein large language models (pLLMs), pLLM-based methods are increasingly being put forward for the protein-protein interaction (PPI) inference task. Here, we identify and confirm that existing pre-trained pLLMs are a source of data leakage for the downstream PPI task. We characterize the extent of the data leakage problem by training and comparing...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
edabd56a-88b9-5688-9337-63a39ed6c638
DOI
10.1101/2025.04.21.649858
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A flaw in using pre-trained pLLMs in protein-protein interaction inference modelsDOI 10.1101/2025.04.21.649858
Select a neighboring publication to make it the new centre.