What is it about?
Humans use common sense constantly without thinking about it. We know that a dropped glass will break, that someone carrying an umbrella expects rain, and that pushing a cup off a table makes it fall. For artificial intelligence, this kind of everyday reasoning has been surprisingly hard to master. This paper surveys how well today's large language models (the technology behind tools like ChatGPT and Claude) can handle commonsense reasoning. We review the datasets used to test these models, compare how different models perform, and examine techniques meant to improve their reasoning, such as prompting them to "think step by step." We find that while the best models now approach human-level scores on some tests, they still struggle with tasks involving social situations, cause-and-effect, and unfamiliar scenarios, often relying on pattern-matching rather than genuine understanding. We also find that newer "reasoning" models, despite excelling at maths and coding, do not show the same leap in everyday common sense. The survey highlights ongoing problems like made-up information and bias, and points to promising directions such as combining language models with structured knowledge sources to make their reasoning more reliable and trustworthy.
Featured Image
Photo by Andres Siimon on Unsplash
Why is it important?
This survey arrives at a pivotal moment. Reasoning-focused models such as OpenAI's o1/o3 and DeepSeek-R1 have recently made dramatic gains on maths and coding, prompting a widespread assumption that the same techniques will improve everyday reasoning too. Our work is among the first to show this assumption does not hold: the extra "thinking" that helps with logic and arithmetic does not translate into comparable gains in common sense, suggesting everyday reasoning is a fundamentally different challenge. Unlike earlier surveys that focus narrowly on specific model families or datasets, we bring together datasets, models, benchmarks, and enhancement techniques into one comprehensive picture, and add an original meta-analysis that groups current models into three performance tiers and reanalyses when step-by-step prompting actually helps rather than hurts. This gives researchers and practitioners a clearer, evidence-based guide to what today's models can and cannot do, and where effort is best directed next, making it a timely resource as the field increasingly relies on language models for real-world decisions in areas like healthcare, education, and content moderation.
Perspectives
What drew me to this work was a simple curiosity: why do machines that can ace a maths olympiad still stumble over things a child finds obvious? Surveying this field made me appreciate just how much of human intelligence is invisible to us precisely because it feels effortless. I hope this article encourages readers to look past the headline benchmark scores and ask a deeper question about what these models actually understand, and I hope it is useful to anyone trying to build AI systems that people can genuinely rely on.
Nicole Teo
Singapore Management University
Read the Original
This page is a summary of: A Survey of Commonsense Reasoning in LLMs, ACM Computing Surveys, July 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3832753.
You can read the full text:
Contributors
The following have contributed to this page







