What is it about?
Thirteen mental health experts rated 30 suicide-related questions by how likely they were to indicate self-harm (from “very low” to “very high” risk). Three AI chatbots—ChatGPT, Claude, and Gemini—were then asked each question 100 times, and researchers noted whether they gave a direct answer or not. For the easiest (very low risk) and most serious (very high risk) questions, ChatGPT and Claude behaved much like the human experts expected, always replying directly to very low risk questions and never directly replying to very high risk ones but at the “in-between” risk levels—low, medium, and high—their answers were less consistent and didn’t match clinician guidance.
Featured Image
Read the Original
This page is a summary of: Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment, Psychiatric Services, August 2025, American Psychiatric Association,
DOI: 10.1176/appi.ps.20250086.
You can read the full text:
Contributors
Be the first to contribute to this page







