New evaluation methods and benchmarks
Methods and benchmarks for evaluating human-agent systems in deployed settings.
2026 NeurIPS Workshop Human-AI Coevolution
The HAIC @ NeurIPS 2026 workshop focuses on developing rigorous empirical methods to evaluate human-agent teams as they coevolve in real-world deployment, particularly in high-stakes domains like healthcare, mental health, aviation, and finance. Building on the inaugural ICLR 2025 workshop, this second edition centers on three interlinked themes: validity in deployed contexts, expert disagreement in human feedback, and adaptive evaluation for coevolving systems.
Full paper
September 3, 2026
AoE
Workshop timeline
Full paperKey deadline
September 3, 2026 · AoE
Paper fit
A strong submission should clearly identify its contribution and evaluate it appropriately.
Methods and benchmarks for evaluating human-agent systems in deployed settings.
Datasets capturing human-agent interaction in deployed or longitudinal settings.
Case studies from high-stakes domains such as healthcare, mental health, aviation, and finance.
Reproducible critiques or replications of existing agentic benchmarks.
Position papers advancing validity-centered or coevolutionary evaluation frameworks.
Compiled from the official call for papers. The organizers’ pages remain authoritative.
Last verified September 9, 2026