What is it about?

Artificial Intelligence (AI) can produce answers that sound convincing and reasonable. But sounding reasonable is not the same as reasoning well. We compared 60 AI models with citizens who deliberated on complex public issues by exchanging reasons, considering different viewpoints, and reflecting on their positions. We asked whether AI connects reasons to decisions in the same structured way people do after deliberation. Most models did not. Their answers often sounded persuasive, but the reasons they gave were not consistently connected to their conclusions. This gap between sounding reasonable and reasoning coherently matters as AI is increasingly used to support decisions on complex public issues. Our findings show that a plausible answer is not necessarily a well-reasoned one.

Featured Image

Why is it important?

AI is increasingly used to support decisions on complex public issues, where there is often no single correct answer. In these situations, good decisions depend on considering different perspectives and connecting conclusions to reasons that others can understand and challenge. Our findings show that AI can appear to reason well without reliably connecting its reasons to its conclusions. This means that, for complex problems with no single correct answer, evaluating AI only on whether it produces the ‘right’ answer is not enough. We also need to evaluate how it forms and justifies its judgments.

Perspectives

I hope this article encourages people to think differently about how we evaluate and trust AI. Many of the most important decisions we face do not have a single correct or wrong answer. They involve competing values, perspectives, and reasons. This means that benchmarking AI only on whether it produces the “right” answer is not enough. We also need to evaluate its broader capacity to form and justify judgments when no clear answer exists. For me, this also raises a fundamental question: how much human judgment are we willing to delegate to AI? AI can produce remarkably convincing answers, but that does not mean we should trust it to make these judgments for us. I hope this article contributes to a broader discussion about how we evaluate AI before deciding where—and whether—we should rely on its judgment.

Francesco Veri
University of Zurich

Read the Original

This page is a summary of: Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment, Proceedings of the National Academy of Sciences, September 2026, Proceedings of the National Academy of Sciences,
DOI: 10.1073/pnas.2600126123.
You can read the full text:

Read

Contributors

The following have contributed to this page