EMM-QA 2026

Deadlines
Machine Learning/CORE Unranked

EMM-QA 2026

ICML Workshop - Efficient Multimodal Question Answering

Jul 11 2026Official workshop site Site reachable

QANTA 2026 is the world’s first multimodal Quizbowl competition that challenges AI systems and human teams to answer questions combining text and visual clues such as photographs, diagrams, and artworks. It promotes human-AI collaboration and advances research in multimodal reasoning, with a parallel workshop at ICML 2026 inviting academic papers on efficient, incremental multimodal question answering.

Official CFP Back to deadlines Verified September 9, 2026

Key deadlines

Verified September 9, 2026

Full paper

May 30, 2026

AoE

Workshop timeline

Submission and decisions

Full paperKey deadline

May 30, 2026 · AoE

Paper fit

Contribution paths

A strong submission should clearly identify its contribution and evaluate it appropriately.

Multimodal AI System

Design a system that interprets both text and image clues to answer quiz bowl questions, handling visual reasoning, OCR, diagram interpretation, and cross-modal fusion.

Human Team Participation

Compete as a human player teaming up with an AI agent to answer multimodal questions and study human-AI complementarity.

Multimodal Question Author

Write pyramid-style tossup questions that integrate images with text clues, designed to be adversarial for AI while solvable by expert humans.

Research areas in scope

01

Multimodal Question Answering

Visual reasoningOCR and diagram interpretationCross-modal fusionAdversarial question designHuman-AI collaborationCalibration under incomplete cluesExplainability and robustness of multimodal systems

Policies worth checking twice

  • Questions must be original ('clean') — not previously published
  • No communication between question writers and computer competition teams
  • Questions must follow PACE guidelines
  • Images must be legally usable (public domain, CC-licensed, or fair use in an educational context)
  • Images must relate directly to the answer and add genuine information beyond the text
  • Text clues preceding an image must be resolvable without the image
  • Text clues following an image must build on — not repeat — the image content
  • All clues must be factually accurate and verifiable

Official sources

Compiled from the official call for papers. The organizers’ pages remain authoritative.

Last verified September 9, 2026