What is it about?

People use TikTok as a search engine and often ask health-related questions. Furthermore, LLM providers advertise their product by encouraging users to fact-check through them. In this work, we first collect and analyze the accuracy and harmfulness of reproductive health content people view when they search for common questions on TikTok. About 30% of the health claims surfaced by search in this dataset are inaccurate, and roughly half of those are clinically harmful. Then, we analyze how LLMs perform when instructed to fact-check. Gemini-2.5-Flash performs best overall, but accuracy drops by about 15% when a model moves from verifying a single clinician-highlighted claim to judging an entire video, and most models are better at flagging harmful information than at catching factual inaccuracies.

Featured Image

Why is it important?

As people use and are encouraged to use LLMs to fact-check, it is important we evaluate the data sources that are likely to be contributing to use of LLMs. Information seen on social media is a crucial source of information which can lead to magnification of existing harms when models that are being deployed at a billion scale population can lead are not capable of detecting inaccuracies and harms.

Perspectives

I hope this article highlights how complex the health information space is, and how jaggedness of LLMs can really impact downstream use cases like fact-checking where careful deployment and more evals are needed.

Vaibhav Balloli
University of Michigan

Read the Original

This page is a summary of: RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3770855.3817447.
You can read the full text:

Read

Resources

Contributors

The following have contributed to this page