Full Papers
6 pages + references · Oral presentation
First Workshop on Reliable Evaluation for Language Models (JUDGe)
JUDGe 2026 is a full-day NeurIPS workshop focused on the reliability and validity of LLM-based evaluators, addressing how evaluation failures cascade within AI systems. It brings together researchers and practitioners to examine systemic issues in LLM judges, including bias, criteria drift, and deployment risks, with the goal of establishing community standards for evaluator transparency.
Paper fit
A strong submission should clearly identify its contribution and evaluate it appropriately.
6 pages + references · Oral presentation
4 pages + references · Poster presentation
Compiled from the official call for papers. The organizers’ pages remain authoritative.
Last verified September 9, 2026