What is it about?
Scientific papers contain large amounts of valuable information, but turning that information into structured datasets usually requires experts to decide in advance what fields to extract. AutoSchema uses large language models to automatically discover useful data fields from scientific literature and then extract information according to the resulting schema. It also links extracted values back to supporting evidence in the papers, making the results easier to verify and use.
Featured Image
Photo by Igor Omilaev on Unsplash
Why is it important?
Building structured datasets from scientific literature is often slow and labor-intensive because researchers must manually design extraction schemas and review large numbers of papers. AutoSchema reduces this burden by allowing the schema to emerge from the literature itself rather than requiring it to be fully predefined. By combining automatic schema discovery with evidence-grounded extraction, the approach can make scientific literature mining more scalable, adaptable to new research areas, and easier to audit.
Perspectives
This work grew from a practical question: can we extract structured scientific knowledge without already knowing exactly what the final schema should look like? We found that letting the system learn the schema from the literature creates a more flexible workflow, especially when entering a new or rapidly evolving research area. For me, an important part of the project was also making the extracted information traceable to evidence, so that automation does not come at the cost of scientific reliability.
Mingfang Zhu
Washington University in Saint Louis
Read the Original
This page is a summary of: AutoSchema: Self-Prompted Schema Induction and Evidence-Grounded Extraction for Materials Science Literature, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3770855.3817812.
You can read the full text:
Contributors
The following have contributed to this page







