GEM 2026

Deadlines
NLP/CORE Unranked

GEM 2026

Fifth Workshop on Natural Language Generation, Evaluation, and Metrics

Jul 03 2026San Diego, California, United StatesOfficial workshop site Site reachable

The GEM 2026 Workshop, co-located with ACL 2026 in San Diego, focuses on advancing the evaluation and metrics of natural language generation systems, particularly in the context of large language models. It emphasizes innovative, reproducible, and sociotechnical evaluation practices, including human-in-the-loop methods, LLM-as-judge frameworks, and benchmarking challenges, with a special Comic-Con-themed edition encouraging creative presentation formats.

Official CFP Back to deadlines Verified September 9, 2026

Key deadlines

Verified September 9, 2026

Full paper

May 22, 2026

AoE

Workshop timeline

Submission and decisions

Full paperKey deadline

May 22, 2026 · AoE

Paper fit

Contribution paths

A strong submission should clearly identify its contribution and evaluate it appropriately.

Archival Papers

Original and unpublished work for Main, ReproNLP, and Opinion/Statement tracks; published on the ACL Anthology; subject to full review; must not be dual submitted.

Non-Archival Extended Abstracts

Direct submissions of work already presented or under review at another peer-reviewed venue; not published in proceedings; intended for sharing ongoing or recent work.

Findings Papers

Presentations of papers accepted to ACL Findings; authors must submit via a dedicated form to present at GEM.

Opinion and Statement Papers

Position papers that challenge conventional wisdom or propose new directions; must be titled with 'Position:'; do not require new empirical results but must be supported by scientific evidence; reviewed separately and presented in panel discussions.

Research areas in scope

01

Topics of Interest

Automatic evaluation of generation systems, including the use of LMs as evaluatorsCreating evaluation corpora, challenge sets, and living benchmarksCritiques of benchmarking efforts, including contamination, memorization, and validityEvaluation of cutting-edge topics in LM development, including long-context understanding, agentic capabilities, reasoning, and moreEvaluation as measurement beyond raw capability, including ideas such as robustness, reliability, and moreMultimodal evaluation across text, vision, and other modalitiesCost-aware and efficient evaluation methods applicable across languages and scenariosHuman evaluation and its role in the era of powerful LMsEvaluation of sociotechnical systems employing large language modelsSurveys and meta-assessments of evaluation methods, metrics, and benchmarksBest practices for dataset and benchmark documentationIndustry applications of the above-mentioned topics, especially internal benchmarking or navigating the gap between academic metrics and real-world impact

Policies worth checking twice

  • Dual submission of archival papers is not allowed.
  • ARR-reviewed papers (archival only) are accepted via special ARR Commitments OpenReview and must not have dual commitments.
  • Authors of direct submissions may be asked to provide two reviews (either one author doing two, or two authors each doing one).
  • All papers must include an Ethics statement and a Limitations section, which do not count toward the page limit.
  • References and Appendices do not count toward the page limit.
  • Archival papers must be 4–8 pages; Opinion/Statement papers must be 2–4 pages; Extended abstracts must be 1–2 pages.
  • Accepted papers may receive an additional page to address reviewer comments.

Official sources

Compiled from the official call for papers. The organizers’ pages remain authoritative.

Last verified September 9, 2026

GEM 2026: deadlines, venue, and submission guide | COREXA