{"data":{"event":{"id":"96efef85-f405-49c3-8dd1-ddb599c45596","slug":"iterative-gene-enrichment-analysis-interpretable-networks-for-human-and--dbac36cc48","title":"\n\nIterative Gene Enrichment Analysis: interpretable networks for human and AI-assisted biological insights \n\n","short_summary":"Background: Over-representation analysis (ORA) is widely used to interpret gene lists from high-throughput biological experiments. However, ORA often produces long and fragmented lists of enriched terms that are difficult to translate into coherent, testable biological hypotheses. This challenge is amplified by the rapid growth of gene set collections and the need to integrate results across multiple gene-set collections. Large language models (LLMs) can assist in summarizing enrichment outputs, but their effectiveness is limited by the lack of reproducibility and statistically supported repre","full_description":"Background: Over-representation analysis (ORA) is widely used to interpret gene lists from high-throughput biological experiments. However, ORA often produces long and fragmented lists of enriched terms that are difficult to translate into coherent, testable biological hypotheses. This challenge is amplified by the rapid growth of gene set collections and the need to integrate results across multiple gene-set collections. Large language models (LLMs) can assist in summarizing enrichment outputs, but their effectiveness is limited by the lack of reproducibility and statistically supported representations. Results: We introduce iterative Gene Enrichment Analysis (iGEA), a software framework that transforms ORA-based enrichment outputs into structured, interpretable networks, supporting human interpretation and AI-assisted exploration. iGEA iteratively selects the most significant enriched term, removes its overlapping genes from the input list, and repeats enrichment until no significant terms remain. Applied independently across gene-set collections, this procedure yields a compact set of non-overlapping enriched terms within each collection. Integration of collection-specific results yields a cross-collection gene-term network in which genes connect terms from different collections. Using a published set of HIV dependency factors, iGEA identified five compact modules spanning secretory trafficking, nuclear transport, transcription elongation, proteostasis, and innate immune signaling, enabling rapid hypothesis generation and interactive exploration. The resulting network structure supports standardized prompting and provides a structured representation for LLM-assisted summarization and exploration of enrichment results. Conclusions: iGEA provides a software framework for gene enrichment analysis that addresses the interpretation challenge of ORA by generating compact, empirically benchmarked cross-collection gene-term networks. This network representation reduces within-collection redundancy, exposes relationships across gene-set collections, and provides a structure that makes modules and hub genes easier to identify through visualization, network-based analysis, and LLM-assisted exploration. Availability: Source code is available at: https://github.com/aion-labs/Gene-Enrichment-Analysis A web-based version of the application is available at: https://iterative-gene-enrichment.streamlit.app","ledger_type":"benefit","primary_domain_id":"8f1af1b9-7302-4b51-aae9-fdfac04d158a","event_status":"provisional","event_date":"2026-09-13T00:00:00.000Z","discovery_date":"2026-09-13T00:00:00.000Z","first_published_date":"2026-09-13T00:00:00.000Z","last_reviewed_date":"2026-09-13T00:00:00.000Z","geographic_scope":"International","affected_population":null,"base_impact_tier":1,"base_score":"1.00","attribution_multiplier":"0.1000","evidence_multiplier":"0.1000","realization_multiplier":"0.2000","durability_multiplier":"0.5000","current_event_score":"0.001000","confidence_level":"low","score_explanation":"Auto-published from news ingest as a provisional placeholder. Score is conservative until a named release is identified and the record is rescored.","methodology_version_id":"d7881163-fb23-4e71-8625-bac1a8662c0f","original_methodology_version_id":"d7881163-fb23-4e71-8625-bac1a8662c0f","published_at":"2026-09-14T04:00:48.960Z","created_at":"2026-09-14T04:00:48.960Z","updated_at":"2026-09-14T04:00:48.960Z","flags":[],"domain_name":"Biology","domain_slug":"biology","methodology_version":"0.1"},"contributions":[{"id":"32276a57-0fd7-4f5d-b8be-31dd6d306181","event_id":"96efef85-f405-49c3-8dd1-ddb599c45596","model_id":"535014f7-c941-4d39-aa9b-8f9c1d3d07b7","role_description":"Unspecified system mentioned or implied by a news item. Remap to a named release when identified.","attribution_multiplier":"0.1000","credit_share":"1.0000","contribution_score":"0.001000","attribution_rationale":"News ingest does not infer a named model from the publisher alone. Attribution stays unspecified until a release is identified.","attribution_confidence":"medium","first_used_date":"2026-09-13T00:00:00.000Z","model_version_if_known":null,"review_status":"approved","created_at":"2026-09-14T04:00:49.003Z","model_slug":"unspecified-ai-system","model_name":"Unspecified AI system","identity_class":"unknown","is_internal":false,"family_name":"Unspecified","family_slug":"unknown-unspecified","organization_name":"Unknown","organization_slug":"unknown"}],"sources":[{"id":"f6a00bce-3a05-4348-8e6d-cb6f618285c6","event_id":"96efef85-f405-49c3-8dd1-ddb599c45596","url":"\nhttps://www.biorxiv.org/content/10.64898/2026.09.06.749570v1?rss=1\n","canonical_url":"\nhttps://www.biorxiv.org/content/10.64898/2026.09.06.749570v1?rss=1\n","source_type":"preprint","publisher":"bioRxiv","author":null,"publication_date":"2026-09-13T00:00:00.000Z","retrieved_at":"2026-09-14T04:00:49.062Z","title":"\n\nIterative Gene Enrichment Analysis: interpretable networks for human and AI-assisted biological insights \n\n","excerpt":"Background: Over-representation analysis (ORA) is widely used to interpret gene lists from high-throughput biological experiments. However, ORA often produces long and fragmented lists of enriched terms that are difficult to translate into coherent, testable biological hypotheses. This challenge is amplified by the rapid growth of gene set collections and the need to integrate results across multiple gene-set collections. Large language models (LLMs) can assist in summarizing enrichment outputs,","content_hash":null,"source_reliability_class":"medium","is_primary_source":true,"is_independent":true,"is_peer_reviewed":false,"archived_url":null,"created_at":"2026-09-14T04:00:49.062Z"}],"claims":[{"id":"98b50a44-5931-45b4-a744-9be68a9df1c0","event_id":"96efef85-f405-49c3-8dd1-ddb599c45596","claim_text":"\n\nIterative Gene Enrichment Analysis: interpretable networks for human and AI-assisted biological insights \n\n","claim_type":"outcome","claim_status":"supported","confidence_score":null,"created_at":"2026-09-14T04:00:49.074Z","updated_at":"2026-09-14T04:00:49.074Z"}],"revisions":[{"id":"d8f6eb68-0046-4949-91cb-9f583fe9616d","event_id":"96efef85-f405-49c3-8dd1-ddb599c45596","model_id":null,"previous_score":"0.000000","new_score":"0.001000","previous_factors":{},"new_factors":{"evidence":0.1,"base_score":1,"durability":0.5,"attribution":0.1,"realization":0.2},"change_reason":"Auto-published from news ingest.","trigger_type":"news_ingest","trigger_source_ids":null,"reviewer_id":null,"review_status":"published","created_at":"2026-09-14T04:00:49.089Z"}],"secondary":[]},"methodology_version":"0.1"}