What happened
Breast cancer remains one of the leading causes of cancer-related deaths among women worldwide. Drug resistance, toxicity, and limited target specificity are the major challenges in the development of effective therapeutics. Anti-breast cancer peptides (ABCPs) have emerged as effective drug candidate due to its low toxicity, high selectivity, and ability to target cancer. However, it is very time consuming and expensive to identify novel ABCPs only through experiment methods. To address this challenge, we developed ABCP_finder, the first dedicated computational framework specifically designed for the prediction of ABCPs using transformer-based protein language model embeddings. Positive and negative datasets were carefully constructed to ensure a biologically meaningful classification task. To prevent the data leakage and realistic evaluation, homology aware train-test split strategy was utilized by using CD-HIT at 30% sequence identity with 80% coverage. Peptide representations were generated using pretrained transformer models, ProtBERT and ESM2, followed by classification using a multilayer perceptron (MLP). Among the tested models, ProtBERT showed superior performance, achieving 93.82% accuracy, 86.88% recall, 90.59% F1-score, 0.8618 MCC, 96.67% AUC, and a Brier Score of 0.0633, demonstrating strong predictive capability under imbalanced conditions. Calibration analysis supported the selection of a 0.7 probability threshold for identifying high confidence ABCPs. Further external validation using xDeep-AcPEP demonstrated that unknown peptide sequences that has been predicted as ABCPs by ABCP_finder are exhibiting favourable IC values. This supports the biological relevance of these unknown peptides. Overall, ABCP_finder provides a reliable and practical platform for large-scale ABCP screening and can significantly accelerate the discovery of novel peptide therapeutics for breast cancer treatment.
