Net Good IndexSubmit a correction

Benefit Ledger · provisional · Biology

Mechanistic Interpretability of Protein Language Models Reveals Encoded Structural and Functional Properties of Intrinsically Disordered Proteins

Protein language models (PLMs) such as ESM-2 encode protein sequences as embeddings for downstream tasks. PLMs are trained on a masked learning objective that leverages evolutionary constraints. While interpretability studies of ESM-2 have focused on folded proteins, their behavior on intrinsically disordered proteins (IDPs), which constitute a substantial fraction of the human proteome and are implicated in numerous diseases, remains understudied. Because IDPs experience different types of evolutionary constraints on their amino acid sequences, we hypothesized that PLMs would behave different

Tier 1 UsefulMethodology 0.1
Submit a correction

Current score

+0.00

1 base · Useful (tier 1 of 5, 1 pts)
× 0.1000 attribution · Minor documented assistance
× 0.1000 evidence · Firsthand or social claim
× 0.2000 realization · Proposed
× 0.5000 durability · Medium-term
Event-level product before credit split: 0.00

Auto-published from news ingest as a provisional placeholder. Score is conservative until a named release is identified and the record is rescored.

What happened

Protein language models (PLMs) such as ESM-2 encode protein sequences as embeddings for downstream tasks. PLMs are trained on a masked learning objective that leverages evolutionary constraints. While interpretability studies of ESM-2 have focused on folded proteins, their behavior on intrinsically disordered proteins (IDPs), which constitute a substantial fraction of the human proteome and are implicated in numerous diseases, remains understudied. Because IDPs experience different types of evolutionary constraints on their amino acid sequences, we hypothesized that PLMs would behave differently on disordered versus folded regions. Here we show that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals. The model assigns heightened attention to disease-relevant residues even at high levels of disorder. Moreover, we show that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings. These findings suggest PLMs capture valuable information relevant to IDP biology despite their bias toward structured residues.

Model attribution

Unspecified AI system
Version unspecified
+0.00

Unspecified

Unspecified system mentioned or implied by a news item. Remap to a named release when identified.

News ingest does not infer a named model from the publisher alone. Attribution stays unspecified until a release is identified.

Attribution 0.1000 · Credit share 100% · Unknown

Claims

  • Mechanistic Interpretability of Protein Language Models Reveals Encoded Structural and Functional Properties of Intrinsically Disordered Proteins

    outcome · supported

Sources

Revision history

  • 14 Sep 2026 · 0.00 0.00

    Auto-published from news ingest.