New evaluation methods and benchmarks
Methods and benchmarks for evaluating human-agent systems in deployed settings.
2026 NeurIPS Workshop Human-AI Coevolution
The HAIC @ NeurIPS 2026 workshop focuses on developing rigorous empirical evaluation methods for human-agent teams in the agentic era, particularly in high-stakes domains like healthcare, mental health, aviation, and finance. It builds on the inaugural ICLR 2025 workshop by narrowing its scope to three interrelated themes centered on validity, expert disagreement, and adaptive evaluation as systems and users coevolve.
Full paper
August 30, 2026
AoE
Workshop timeline
Full paperKey deadline
August 30, 2026 · AoE
Paper fit
A strong submission should clearly identify its contribution and evaluate it appropriately.
Methods and benchmarks for evaluating human-agent systems in deployed settings.
Datasets capturing human-agent interaction in deployed or longitudinal settings.
Case studies from high-stakes domains such as healthcare, mental health, aviation, and finance.
Reproducible critiques or replications of existing agentic benchmarks.
Position papers advancing validity-centered or coevolutionary evaluation frameworks.
Compiled from the official call for papers. The organizers’ pages remain authoritative.
Last verified September 9, 2026