GEM 2026

Deadlines
NLP/CORE Unranked

GEM 2026

Fifth Workshop on Natural Language Generation, Evaluation, and Metrics

Jul 03 2026San Diego, California, United StatesOfficial workshop site Site reachable

The GEM 2026 Workshop, co-located with ACL 2026 in San Diego, focuses on advancing the evaluation and metrics of language models, addressing challenges in automatic and human evaluation, benchmark validity, and emerging areas like agentic capabilities and multimodal systems. This year's theme is a 'Comic-Con Edition,' encouraging creative presentations, and features special tracks for opinion papers and the ReproNLP shared task on evaluation reproducibility.

Official CFP Back to deadlines Verified September 9, 2026

Key deadlines

Verified September 9, 2026

Full paper

May 22, 2026

AoE

Workshop timeline

Submission and decisions

Full paperKey deadline

May 22, 2026 · AoE

Paper fit

Contribution paths

A strong submission should clearly identify its contribution and evaluate it appropriately.

Archival Papers

Original and unpublished work for Main, ReproNLP, and Opinion/Statement tracks; published on the ACL Anthology. Can be direct submissions or ARR-reviewed papers.

Non-Archival Extended Abstracts

Direct submissions of work already presented/committed or under review at a peer-reviewed venue; not published in proceedings.

Findings Papers

Presentation of papers accepted to the ACL Findings; authors must fill out a form to present at GEM.

Opinion and Statement Papers

Position papers that challenge conventional wisdom or propose new directions; must be titled with 'Position:' and are presented in curated panel discussions.

Research areas in scope

01

Topics of Interest

Automatic evaluation of generation systems, including the use of LMs as evaluatorsCreating evaluation corpora, challenge sets, and living benchmarksCritiques of benchmarking efforts, including contamination, memorization, and validityEvaluation of cutting-edge topics in LM development, including long-context understanding, agentic capabilities, reasoning, and moreEvaluation as measurement beyond raw capability, including ideas such as robustness, reliability, and moreMultimodal evaluation across text, vision, and other modalitiesCost-aware and efficient evaluation methods applicable across languages and scenariosHuman evaluation and its role in the era of powerful LMsEvaluation of sociotechnical systems employing large language modelsSurveys and meta-assessments of evaluation methods, metrics, and benchmarksBest practices for dataset and benchmark documentationIndustry applications of the above-mentioned topics, especially internal benchmarking or navigating the gap between academic metrics and real-world impact

Policies worth checking twice

  • Dual submissions are not allowed for direct archival submissions.
  • ARR-reviewed papers (archival only) are allowed with a meta-review; dual commitments not allowed.
  • All papers must include an Ethics statement and a Limitations section, which do not count towards the page limit.
  • References and Appendices do not count towards the page limit.
  • Archival papers: 4–8 pages; Opinion/Statement papers: 2–4 pages; Extended abstracts: 1–2 pages.
  • Accepted papers may receive an additional page to address reviewer comments.
  • For non-ARR submissions, authors may be asked to provide 2 reviews (either one author doing 2 reviews or two authors each doing one review).
  • Submissions must conform to ACL 2026 style guidelines.

Official sources

Compiled from the official call for papers. The organizers’ pages remain authoritative.

Last verified September 9, 2026

GEM 2026: deadlines, venue, and submission guide | COREXA