What happened
Ensuring the availability and accessibility of research data is fundamental to advancing knowledge, as codified in the FAIR principles (Findable, Accessible, Interoperable, and Reusable). Accurate metadata documentation is indispensable for meeting these principles; however, entries in deposition databases often contain inadequate, repetitive, or incomplete descriptions. Much of this metadata is captured in free-text fields, motivating the need for scalable, repository-agnostic methods to quantify metadata richness. Here, we quantify free-text metadata richness across three repositories using Natural Language Processing (NLP) methods: BioDare2, an experimental circadian rhythm database; DataShare, a domain-agnostic University of Edinburgh database; and Image Data Resource (IDR), a public repository of biological image datasets from published studies. In general, repositories exhibit distinct distributions of word count and information density, consistent with differences in scope. Named-entity recognition and information-density metrics detected significant category-level differences in BioDare2 (species) and DataShare (communities), while identifying greater consistency in the more curated IDR. The most frequent entities reflected each repository's focus: circadian terminology in BioDare2, microscopy-related entities in IDR, and community-driven terms in DataShare. We develop a scalable framework utilising word counts, named-entity recognition, and entity-derived information density to assess metadata quality across repositories, offering a broadly applicable evaluation tool.
