What is it about?
This study tested whether artificial intelligence (AI) could pass a challenging medical licensing exam given entirely in Spanish. The researchers used GPT-4o, an advanced AI language model, to answer over 1,000 multiple-choice questions from Chile's medical licensing exam (EUNACOM). They tested two approaches: having a single AI work alone, and having multiple AI agents work together as a team. The team-based approach, where different AI agents collaborated like doctors consulting each other, achieved nearly 90% accuracy - significantly better than the 85-87% scored by single AI agents working alone. The AI performed exceptionally well in psychiatry, neurology, and surgery (over 95% accuracy), but struggled more with specialties like neonatology and otolaryngology (around 77%). Importantly, the study found that many medical exam questions could be answered correctly without complex reasoning, suggesting that only a fraction of standardized tests truly require sophisticated problem-solving skills.
Featured Image
Photo by Ant Rozetsky on Unsplash
Why is it important?
This research is groundbreaking because it's the first comprehensive evaluation of AI performance on Spanish-language medical exams using multiple collaborative strategies. With Spanish being the second most spoken language globally, this work addresses a critical gap in making AI-powered medical education accessible to millions of Spanish-speaking healthcare professionals and students. The timing is particularly relevant as medical education increasingly adopts AI tools. This study demonstrates that AI can effectively support medical training in non-English contexts, potentially helping address physician shortages in Spanish-speaking regions. The finding that team-based AI approaches mirror real-world medical collaboration also suggests these tools could enhance clinical decision-making beyond just education.
Perspectives
As someone involved in this research, I'm struck by how the collaborative AI approach mirrors the way medical teams actually work. When the AI agents discussed cases together, they achieved accuracy levels that rival or exceed many human test-takers. What surprised me most was discovering that simpler strategies worked well for many questions - reminding us that not every medical decision requires complex reasoning. This has important implications for how we design both AI tools and medical exams. The lower performance in certain specialties like neonatology highlights that we still have work to do. These gaps likely reflect biases in AI training data, emphasizing the need for more diverse, specialty-specific datasets. Overall, this research convinces me that AI can democratize access to high-quality medical education, but we must ensure it works equitably across all medical domains and languages.
Nicolas Sumonte
Pontificia Universidad Catolica de Chile
Read the Original
This page is a summary of: Performance of single-agent and multi-agent language models in Spanish language medical competency exams, BMC Medical Education, May 2025, Springer Science + Business Media,
DOI: 10.1186/s12909-025-07250-3.
You can read the full text:
Contributors
The following have contributed to this page







