What happens when artificial intelligence is asked not only to read a question, but also to understand an image, interpret a diagram, recognise a map, or make sense of a scientific figure?
That is the challenge at the heart of QANTA 2026: Efficient, Incremental Multimodal Question Answering, the world’s first multimodal Quizbowl computer competition. Building on the human–AI collaboration format introduced in 2025, QANTA 2026 takes question answering into new territory by combining text and images in a live competition setting.
In other words, this is not just about whether AI can “know” the right answer. It is about whether AI can reason with incomplete information, combine visual and textual clues, and decide when it knows enough to commit.
A competition where humans and AI think together
Quizbowl is a fast-paced question answering game built around clues. Questions are often written in a “pyramid” style: they begin with harder, more obscure clues and gradually move toward more recognisable information. Players can interrupt or “buzz” as soon as they think they know the answer.
This makes Quizbowl a fascinating test of intelligence. Success depends not only on knowledge, but also on timing, confidence, and judgement.
QANTA 2026 adds a new layer to this format: multimodality. Questions may include photographs, artworks, maps, diagrams, charts, or scientific figures alongside written clues. Human players often interpret such visual material naturally. We look at an image, connect it with context, and draw on prior knowledge almost instantly. For AI systems, however, this is still a complex challenge.
Can a system understand what is important in an image? Can it connect visual evidence with textual clues? Can it decide when the evidence is strong enough to answer and when it should wait?
These are exactly the questions QANTA 2026 sets out to explore.
Three ways to take part
QANTA 2026 invites participation from several communities: AI researchers, Quizbowl players, question writers, and anyone interested in the future of human–AI collaboration.
Participants can join in three main ways.
First, computer teams can build a multimodal AI teammate. These systems need to process both language and visual inputs, including image understanding, OCR, diagram interpretation, and cross-modal reasoning. The goal is not simply to produce an answer, but to do so efficiently and at the right moment.
Second, human teams can play in the competition and collaborate with AI agents. This makes the event more than a technical benchmark. It becomes a real-world test of how people and AI systems complement each other: where humans are stronger, where AI can assist, and how teams decide whether to trust a machine-generated suggestion.
Third, authors can contribute by writing multimodal questions. These questions should be challenging for AI systems while remaining solvable by expert humans using both the image and the text. Accepted questions are paid, and standout questions and packets may also receive recognition.
Why multimodal question answering matters
Multimodal question answering is becoming increasingly important because real-world information rarely comes in only one form. We read text, but we also rely on images, tables, screenshots, charts, videos, maps, and diagrams.
A doctor may need to combine clinical notes with a scan. A researcher may interpret a graph alongside a paper. A student may ask a question about a diagram in a textbook. A professional may need to extract meaning from a screenshot, a report, or a visual dashboard.
For AI to be useful in these situations, it must go beyond text-only reasoning. It must learn to “look” carefully, connect evidence across formats, and explain or justify its answers in ways people can understand.
QANTA 2026 turns this broad challenge into an exciting and measurable competition.
Efficiency is part of the challenge
The competition is closely connected to the EMM-QA Workshop on Efficient Multimodal Question Answering, taking place at ICML 2026 in Seoul. The workshop focuses on systems that balance accuracy, efficiency, and adaptability across different types of input.
This is an important point. The future of AI is not only about building larger models. It is also about developing systems that can work under real-world constraints: limited resources, time pressure, incomplete evidence, and the need to decide when not to answer.
That is why topics such as retrieval-augmented generation, compact models, efficient inference, visual token compression, confidence calibration, and human-in-the-loop evaluation are central to the workshop.
In the context of QANTA, efficiency becomes very concrete. In a live Quizbowl match, a system cannot wait forever. It must process clues incrementally and decide whether to buzz, abstain, or wait for more evidence. This makes the competition a practical test of both reasoning and self-assessment.
From live matches to research insights
QANTA 2026 includes both human and computer tracks, with live competition formats based on standard Quizbowl rules.
Tossups reward fast, confident answering under incomplete information. Bonuses, on the other hand, are team questions with multiple parts, where collaboration and explanation matter more. Together, these formats provide a rich environment for studying how humans and AI systems reason, communicate, and make decisions.
The competition also connects directly to the ICML 2026 EMM-QA Workshop. Computer winners will be announced in July, and the workshop will provide a space for system builders, researchers, and participants to discuss results, methods, and lessons learned.
Beyond the leaderboard, the real value lies in the insights generated: Which systems can use visual evidence effectively? When do they fail? How well do they know their own uncertainty? And how much can they genuinely help human players?
A glimpse into the future of human–AI collaboration
QANTA 2026 is more than a game. It is a window into the future of AI systems that can support people in complex, knowledge-intensive tasks.
By bringing together multimodal reasoning, efficient system design, live competition, and human–AI teamwork, QANTA creates a setting where AI is tested not in isolation, but as part of a collaborative process.
That makes the competition especially relevant at a time when AI tools are increasingly entering education, research, healthcare, public services, and everyday work. The question is no longer only whether AI can answer. The bigger question is whether it can help people think better, faster, and more carefully while knowing when to stay silent.
QANTA 2026 invites researchers, players, writers, and curious minds to explore that question together.
Whether you build a system, write a question, join as a human team, or follow the results through the EMM-QA Workshop at ICML 2026, this competition offers a compelling look at where multimodal AI is heading next.
The EMM-QA Workshop is also partially supported by the ELOQUENCE project, reinforcing ELOQUENCE’s commitment to advancing trustworthy, efficient and human-centred multimodal AI.
