Multimodal AI System
Design a system that interprets both text and image clues to answer quiz bowl questions, handling visual reasoning, OCR, diagram interpretation, and cross-modal fusion.
ICML Workshop - Efficient Multimodal Question Answering
QANTA 2026 is the world’s first multimodal Quizbowl competition that challenges AI systems and human teams to answer questions combining text and visual clues such as photographs, diagrams, and artworks. It promotes human-AI collaboration and advances research in multimodal reasoning, with a parallel workshop at ICML 2026 inviting academic papers on efficient, incremental multimodal question answering.
Full paper
May 30, 2026
AoE
Workshop timeline
Full paperKey deadline
May 30, 2026 · AoE
Paper fit
A strong submission should clearly identify its contribution and evaluate it appropriately.
Design a system that interprets both text and image clues to answer quiz bowl questions, handling visual reasoning, OCR, diagram interpretation, and cross-modal fusion.
Compete as a human player teaming up with an AI agent to answer multimodal questions and study human-AI complementarity.
Write pyramid-style tossup questions that integrate images with text clues, designed to be adversarial for AI while solvable by expert humans.
Compiled from the official call for papers. The organizers’ pages remain authoritative.
Last verified September 9, 2026