What happened
Protein-protein interactions (PPIs) have been detected and reported in the millions, but while they are used in many different contexts for better understanding cellular processes in health and disease, the knowledge of the human PPI network is far from complete, containing many false positive measurements and being highly biased. Both to chart the extent of those problems and to solve them requires not just a knowledge of high-confidence positive interactions, but also likely non-interacting protein pairs. However, this information is typically not reported in PPI studies. We developed a methodology to reconstruct this knowledge from existing PPI data. We reconstruct the experimental search space in which PPI screens have been performed and then create a model that informs how likely a PPI is real given its testing and observation frequency. We argue that negative protein pairs allow us to estimate the error rates of experimental and computational screens. We show how this knowledge could be incorporated for calibration. Finally, we evaluate a simple machine learning approach to PPI prediction and propose how such negative data can be used for training instead of random protein pairs. Together, our results show that reconstructing the experimental search space recovers a largely overlooked layer of information from existing PPI data that can help guide a more accurate and complete mapping of the human interactome.
