What is it about?
Researchers have published more than 1,600 studies on transistor-based chemical sensors, but the data are difficult to reuse. The useful details are spread across prose, figures, and tables, and different research groups often describe their devices in different ways. We built Text-Twin-Translation, or T3, to turn that literature into a design tool. One language model extracted data while another reviewed the extraction errors and improved the instructions. The resulting system organized 28 types of information from the papers. We then represented each sensor as a network of connected parts, allowing the model to learn how the channel, electrodes, dielectric, sensing molecule, fabrication steps, and test conditions work together. Across three performance-prediction tasks, it reached accuracies from 85.1% to 92.3% and outperformed 17 tabular and neural baselines. For the final step, we kept the virtual sensor fixed and changed only the sensing molecule. This allowed us to screen more than 123 million PubChem molecules under the same device conditions. Quantum-chemistry calculations highlighted one candidate with better predicted selectivity for PFOS than the beta-cyclodextrin reference molecule. It gives the future wet-lab experiment a clear starting point.
Featured Image
Photo by Denes Kozma on Unsplash
Why is it important?
Most scientific papers are written to communicate a result, not to become machine learning data. A field can therefore contain years of useful experiments without having a dataset that preserves the details needed for comparison or prediction. T3 addresses this problem at the level of the whole device. Instead of treating a sensor as a flat list of materials, it keeps the connections among the channel, electrodes, dielectric, sensing molecule, target, processing history, and test environment. That structural information was not cosmetic. When graph message passing was removed and only the material fingerprints were retained, prediction performance dropped substantially. The trained model could then be used as a controlled screening tool, with one device component changed at a time. In the PFOS case study, the framework screened 123,239,643 molecules, selected five leading candidates for quantum-chemistry validation, and identified a previously unexplored molecule with improved predicted selectivity over the reference molecule. The immediate outcome is a concrete experimental lead. The broader contribution is a reusable route from scientific literature to a physically grounded, testable hypothesis for complex material-device systems.
Perspectives
As an experimental materials researcher, I know the difference between reading a promising paper and being able to use it. The details that determine whether a device works are often reported in different places, with different terminology, and under different test conditions. After reading hundreds of papers, you may still not have a dataset that can guide the next experiment. That frustration was the starting point for T3. What surprised me most was how strongly the device layout affected the model. When we removed the graph connections and kept only the material fingerprints, performance dropped sharply. In hindsight, this is exactly what an experimentalist would expect. A sensing material does not work alone. Its behavior depends on what it contacts, where it sits in the device, how it was processed, and how the measurement was performed. The PFOS molecule identified here is still a computational hypothesis. It now needs to be immobilized, integrated into a device, and tested in the laboratory. That is the role I want AI to play: narrow an enormous search space, preserve the physics that matters, and give experiments a better place to start.
Rui Ding
University of Chicago
Read the Original
This page is a summary of: Text-Twin-Translation (T
3
): A Full-Stack Machine Learning Framework for Functional Material-Device Systems Discovery, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3770855.3819013.
You can read the full text:
Contributors
The following have contributed to this page







