{"data":{"event":{"id":"9f6e7df9-6e42-4427-97bf-350e16c75793","slug":"winsyn-an-automated-pipeline-for-realistic-enterprise-question-answering-52123371dd","title":"WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation","short_summary":"arXiv:2609.12171v1 Announce Type: new \nAbstract: Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise se","full_description":"arXiv:2609.12171v1 Announce Type: new \nAbstract: Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise settings. In this work, we introduce an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. Our method simulates long-running enterprise projects spanning several months and involving up to 25 interacting employees across multiple roles. The data emphasizes ambiguity, distributed information, and naturally occurring queries. To validate the pipeline, we evaluate few standard agentic baselines on our datasets using the latest frontier models. We find that aggregate scores averaged over all queries remain below 80% for each dataset, indicating significant room for improvement. These findings suggest that more work remains to be done for enterprise deployment and underscore the importance of realistic, high-complexity evaluation data for developing stronger real-world enterprise DR systems.","ledger_type":"benefit","primary_domain_id":"8f1af1b9-7302-4b51-aae9-fdfac04d158a","event_status":"provisional","event_date":"2026-09-14T00:00:00.000Z","discovery_date":"2026-09-14T00:00:00.000Z","first_published_date":"2026-09-14T00:00:00.000Z","last_reviewed_date":"2026-09-14T00:00:00.000Z","geographic_scope":"International","affected_population":null,"base_impact_tier":1,"base_score":"1.00","attribution_multiplier":"0.1000","evidence_multiplier":"0.1000","realization_multiplier":"0.2000","durability_multiplier":"0.5000","current_event_score":"0.001000","confidence_level":"low","score_explanation":"Auto-published from news ingest as a provisional placeholder. Score is conservative until a named release is identified and the record is rescored.","methodology_version_id":"d7881163-fb23-4e71-8625-bac1a8662c0f","original_methodology_version_id":"d7881163-fb23-4e71-8625-bac1a8662c0f","published_at":"2026-09-14T04:00:51.186Z","created_at":"2026-09-14T04:00:51.186Z","updated_at":"2026-09-14T04:00:51.186Z","flags":[],"domain_name":"Biology","domain_slug":"biology","methodology_version":"0.1"},"contributions":[{"id":"08df023e-3549-49f4-93eb-404d04196706","event_id":"9f6e7df9-6e42-4427-97bf-350e16c75793","model_id":"535014f7-c941-4d39-aa9b-8f9c1d3d07b7","role_description":"Unspecified system mentioned or implied by a news item. Remap to a named release when identified.","attribution_multiplier":"0.1000","credit_share":"1.0000","contribution_score":"0.001000","attribution_rationale":"News ingest does not infer a named model from the publisher alone. Attribution stays unspecified until a release is identified.","attribution_confidence":"medium","first_used_date":"2026-09-14T00:00:00.000Z","model_version_if_known":null,"review_status":"approved","created_at":"2026-09-14T04:00:51.197Z","model_slug":"unspecified-ai-system","model_name":"Unspecified AI system","identity_class":"unknown","is_internal":false,"family_name":"Unspecified","family_slug":"unknown-unspecified","organization_name":"Unknown","organization_slug":"unknown"}],"sources":[{"id":"254c6dd9-87a2-45b8-840a-b32d24333f85","event_id":"9f6e7df9-6e42-4427-97bf-350e16c75793","url":"https://arxiv.org/abs/2609.12171","canonical_url":"https://arxiv.org/abs/2609.12171","source_type":"preprint","publisher":"arXiv cs.AI","author":null,"publication_date":"2026-09-14T00:00:00.000Z","retrieved_at":"2026-09-14T04:00:51.203Z","title":"WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation","excerpt":"arXiv:2609.12171v1 Announce Type: new \nAbstract: Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-","content_hash":null,"source_reliability_class":"medium","is_primary_source":true,"is_independent":true,"is_peer_reviewed":false,"archived_url":null,"created_at":"2026-09-14T04:00:51.203Z"}],"claims":[{"id":"3f8f1605-5c54-4db2-9bad-18436663a21f","event_id":"9f6e7df9-6e42-4427-97bf-350e16c75793","claim_text":"WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation","claim_type":"outcome","claim_status":"supported","confidence_score":null,"created_at":"2026-09-14T04:00:51.211Z","updated_at":"2026-09-14T04:00:51.211Z"}],"revisions":[{"id":"71988d4a-f738-4650-a299-a3c75a3d2ff3","event_id":"9f6e7df9-6e42-4427-97bf-350e16c75793","model_id":null,"previous_score":"0.000000","new_score":"0.001000","previous_factors":{},"new_factors":{"evidence":0.1,"base_score":1,"durability":0.5,"attribution":0.1,"realization":0.2},"change_reason":"Auto-published from news ingest.","trigger_type":"news_ingest","trigger_source_ids":null,"reviewer_id":null,"review_status":"published","created_at":"2026-09-14T04:00:51.218Z"}],"secondary":[]},"methodology_version":"0.1"}