{"data":{"event":{"id":"7f2c7e61-2d43-4dca-9990-dc8abb410f91","slug":"gta-graph-theory-agent-and-benchmark-for-algorithmic-graph-reasoning-wit-c3e8878ab0","title":"GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs","short_summary":"arXiv:2609.12265v1 Announce Type: new \nAbstract: Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural langu","full_description":"arXiv:2609.12265v1 Announce Type: new \nAbstract: Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and adjacency matrix. Evaluating eight LLMs on GT Bench shows that accuracy is strongly tied to the input representation, that the best representation shifts with graph density, size, and topology as well as with the model, and that this sensitivity persists, attenuated, in the strongest reasoning models. Building on these observations, we propose the Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM. GTA lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split, outperforming eight prompting and agent baselines, and transfers without retraining to GraCoRe and NLGraph. Code for benchmark generation and evaluation: https://github.com/xzx34/GTA. The project homepage is available at https://xzx34.github.io/gta/.","ledger_type":"benefit","primary_domain_id":"8f1af1b9-7302-4b51-aae9-fdfac04d158a","event_status":"provisional","event_date":"2026-09-14T00:00:00.000Z","discovery_date":"2026-09-14T00:00:00.000Z","first_published_date":"2026-09-14T00:00:00.000Z","last_reviewed_date":"2026-09-14T00:00:00.000Z","geographic_scope":"International","affected_population":null,"base_impact_tier":1,"base_score":"1.00","attribution_multiplier":"0.1000","evidence_multiplier":"0.1000","realization_multiplier":"0.2000","durability_multiplier":"0.5000","current_event_score":"0.001000","confidence_level":"low","score_explanation":"Auto-published from news ingest as a provisional placeholder. Score is conservative until a named release is identified and the record is rescored.","methodology_version_id":"d7881163-fb23-4e71-8625-bac1a8662c0f","original_methodology_version_id":"d7881163-fb23-4e71-8625-bac1a8662c0f","published_at":"2026-09-14T04:00:51.446Z","created_at":"2026-09-14T04:00:51.446Z","updated_at":"2026-09-14T04:00:51.446Z","flags":[],"domain_name":"Biology","domain_slug":"biology","methodology_version":"0.1"},"contributions":[{"id":"18154397-6623-462b-813b-9c190935fcb0","event_id":"7f2c7e61-2d43-4dca-9990-dc8abb410f91","model_id":"535014f7-c941-4d39-aa9b-8f9c1d3d07b7","role_description":"Unspecified system mentioned or implied by a news item. Remap to a named release when identified.","attribution_multiplier":"0.1000","credit_share":"1.0000","contribution_score":"0.001000","attribution_rationale":"News ingest does not infer a named model from the publisher alone. Attribution stays unspecified until a release is identified.","attribution_confidence":"medium","first_used_date":"2026-09-14T00:00:00.000Z","model_version_if_known":null,"review_status":"approved","created_at":"2026-09-14T04:00:51.457Z","model_slug":"unspecified-ai-system","model_name":"Unspecified AI system","identity_class":"unknown","is_internal":false,"family_name":"Unspecified","family_slug":"unknown-unspecified","organization_name":"Unknown","organization_slug":"unknown"}],"sources":[{"id":"1220b94f-49ff-4265-a8de-9611bc9ef204","event_id":"7f2c7e61-2d43-4dca-9990-dc8abb410f91","url":"https://arxiv.org/abs/2609.12265","canonical_url":"https://arxiv.org/abs/2609.12265","source_type":"preprint","publisher":"arXiv cs.AI","author":null,"publication_date":"2026-09-14T00:00:00.000Z","retrieved_at":"2026-09-14T04:00:51.465Z","title":"GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs","excerpt":"arXiv:2609.12265v1 Announce Type: new \nAbstract: Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems ","content_hash":null,"source_reliability_class":"medium","is_primary_source":true,"is_independent":true,"is_peer_reviewed":false,"archived_url":null,"created_at":"2026-09-14T04:00:51.465Z"}],"claims":[{"id":"e84f9e07-3511-4c51-80c3-9278d40f9c88","event_id":"7f2c7e61-2d43-4dca-9990-dc8abb410f91","claim_text":"GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs","claim_type":"outcome","claim_status":"supported","confidence_score":null,"created_at":"2026-09-14T04:00:51.472Z","updated_at":"2026-09-14T04:00:51.472Z"}],"revisions":[{"id":"65ed9ee8-b0e4-4514-9bf8-49731c5a4eec","event_id":"7f2c7e61-2d43-4dca-9990-dc8abb410f91","model_id":null,"previous_score":"0.000000","new_score":"0.001000","previous_factors":{},"new_factors":{"evidence":0.1,"base_score":1,"durability":0.5,"attribution":0.1,"realization":0.2},"change_reason":"Auto-published from news ingest.","trigger_type":"news_ingest","trigger_source_ids":null,"reviewer_id":null,"review_status":"published","created_at":"2026-09-14T04:00:51.479Z"}],"secondary":[]},"methodology_version":"0.1"}