ICLR 2026

Deadlines
Machine Learning/CORE Unranked

ICLR 2026

ICLR 2026 Workshop on Principled Design for Trustworthy AI - Interpretability, Robustness, and Safety across Modalities

Apr 26 2026Rio de Janeiro, BrazilOfficial workshop site Site reachable

The ICLR 2026 Workshop on Principled Design for Trustworthy AI focuses on developing interpretable, robust, and safe AI systems across modalities such as language, vision, audio, and time series. It addresses critical challenges in high-stakes applications like healthcare and autonomous driving by unifying research on interpretability, robustness, uncertainty, and safety throughout the AI lifecycle.

Official CFP Back to deadlines Verified September 9, 2026

Key deadlines

Verified September 9, 2026

Abstract registration

February 3, 2026

AoE

Full paper

February 3, 2026

AoE

Workshop timeline

Submission and decisions

Abstract registrationKey deadline

February 3, 2026 · AoE

Full paperKey deadline

February 3, 2026 · AoE

Paper fit

Contribution paths

A strong submission should clearly identify its contribution and evaluate it appropriately.

Short paper

Max 4 pages, excluding references; can include unlimited appendix within the same PDF as long as main text stays within limit.

Long paper

Max 9 pages, excluding references; can include unlimited appendix within the same PDF as long as main text stays within limit.

Research areas in scope

01

Interpretable and Intervenable Models

concept bottlenecks and modular architecturesmechanistic interpretability and concept-based reasoninginterpretability for control and real-time intervention
02

Inference-Time Safety and Monitoring

reasoning trace auditing in LLMs and VLMsinference-time safeguards and safety mechanismschain-of-thought consistency and hallucination detectionreal-time monitoring and failure intervention mechanisms
03

Multimodal Trust Challenges

grounding failures and cross-modal misalignmentsafety in vision-language and deep vision systemscross-modal alignment and robust multimodal reasoningtrust and uncertainty in video, audio, and time-series models
04

Robustness and Threat Models

adversarial attacks and defensesrobustness to distributional, conceptual, and cascading shiftsformal verification methods and safety guaranteesrobustness under streaming, online, or low-resource conditions
05

Trust Evaluation and Responsible Deployment

human-AI trust calibrationconfidence estimationuncertainty quantificationmetrics for interpretability/alignment/robustnesstransparent and accountable deployment pipelinessafety alignment
06

Safety and Trustworthiness in LLM Agents

safety and failures in planning and action executionemergent behaviors in multi-agent interactionsintervention and control in agent loopsalignment of long-horizon goals with user intentauditing and debugging LLM agents in real-world deployment

Policies worth checking twice

  • Reviews are double-blind.
  • Accepted papers are non-archival.
  • Submissions must be original and not accepted in other archival venues by the submission deadline; violation results in desk rejection.
  • New OpenReview profiles without institutional email require up to two weeks for moderation.
  • New OpenReview profiles with institutional email are activated automatically.

Official sources

Compiled from the official call for papers. The organizers’ pages remain authoritative.

Last verified September 9, 2026