What is it about?

Here we propose strategies that couple natural language processing with deep learning to enhance machine capability for corrosion resistant alloy design. The accuracy of machine learning models for materials datasets is often limited by their inability to incorporate textual data. Manual extraction of numerical parameters from descriptions of alloy processing or experimental methodology inevitably leads to a reduction in information density. To overcome this, we developed a fully automated natural language processing approach to transform textual data into a form compatible for feeding into a deep neural network. Additionally, we implemented a deep learning model with a transformed input feature space, consisting of elemental physical and chemical property based numerical descriptors replacing raw alloy compositions. We trained a process aware deep neural network model capable of simultaneously processing numerical inputs and textual excerpts detailing alloy processing history and electrochemical test methodology. By utilizing tokenization, word embedding, and a long short term memory layer, the model captures the semantic context of experimental protocols. This enriched information density resulted in a pitting potential prediction accuracy substantially beyond the state of the art. Our sixfold cross validation yielded an average validation loss of 150 mV and an R2 coefficient of 0.78 on test data, a marked improvement over the 170 mV loss and 0.61 R2 coefficient of previous simplistic models. Using this process aware model, we performed composition optimizations that accurately discerned the beneficial roles of specific elements. We observed a notably enhanced contribution of molybdenum in ferritic stainless steels and nickel chromium molybdenum alloys, as well as the positive impact of interstitial nitrogen and carbon. The model also correctly identified dissolved copper as a positive contributor in aluminum alloys, resolving apparent data contradictions that arose from differing test methodologies in the raw dataset. Feature Transformed Deep Learning Model To derive mechanistic insights independent of specific elemental identities, we transformed the alloy composition inputs into a set of numerical descriptors based on atomic, physical, and chemical properties. Optimization of this feature space revealed that configurational entropy, atomic packing efficiency, local electronegativity differences, and atomic radii differences are the most critical parameters for enhancing pitting resistance. Transition metal alloys favored increased packing density and electronegativity mismatch to hinder metal dissolution, whereas aluminum alloys showed a tendency toward pure constituent behavior to maintain passive film stability. Predictive Capability for Unseen Alloys A unique advantage of the feature transformed model is its ability to evaluate alloy systems completely absent from the training data. We demonstrated this by predicting the pitting potentials of an Al Cu Sc Zr alloy system. Despite scandium and zirconium being absent in our training records, the model correctly predicted a steady increase in pitting potential with increasing scandium and zirconium contents, learning the effect of these alloying elements directly from their intrinsic physical and chemical properties.

Featured Image

Why is it important?

OVERCOMING THE DATA BOTTLENECK IN MATERIALS DESIGN Corrosion remains a pervasive challenge, contributing to annual global economic losses estimated at 2.5 trillion USD. Traditional machine learning approaches in materials science have struggled to address this effectively because they rely almost exclusively on structured numerical data. This forces researchers to manually extract parameters from textual descriptions of alloy processing and experimental methodologies, a process that is unscalable and inevitably strips away critical contextual information. By integrating natural language processing with deep learning, we preserve the full information density of experimental records. This allows our models to resolve apparent contradictions in raw datasets, such as conflicting pitting potential measurements for the same alloy composition tested under different electrochemical protocols, ultimately yielding substantially higher prediction accuracy. BRIDGING THE GAP BETWEEN CORRELATION AND CAUSALITY Beyond mere prediction, we must understand the underlying mechanisms driving material performance. Standard composition based models operate as black boxes, identifying statistical correlations without revealing the physical reasons behind them. By transforming raw alloy compositions into fundamental atomic, physical, and chemical descriptors, we shift the paradigm toward explainable artificial intelligence. Our optimization trajectories explicitly highlight that parameters such as configurational entropy, atomic packing efficiency, and local electronegativity differences are the primary drivers of pitting resistance. This mechanistic insight allows metallurgists to move beyond trial and error and intentionally engineer microstructures and solid solutions that intrinsically hinder metal dissolution. ACCELERATING DISCOVERY OF UNSEEN ALLOY SYSTEMS The most transformative advantage of this descriptor based approach is its capacity for true generalization. A model trained solely on elemental compositions cannot evaluate an alloy containing elements absent from its training data. In contrast, our feature transformed network calculates intrinsic properties for any given combination of elements. We demonstrated this by accurately predicting the beneficial role of scandium and zirconium in aluminum copper alloys, despite these elements being entirely absent from our training records. This capability drastically reduces the experimental burden required to explore novel compositional spaces, providing a powerful, quantitative tool to accelerate the design of next generation corrosion resistant materials.

Perspectives

INTEGRATING NLP AND FEATURE TRANSFORMATION We recognize that the true potential of this framework lies in unifying our two primary methodologies. By coupling the natural language processing module, which captures nuanced processing histories and test protocols, with the feature-transformed descriptor model, we can quantitatively evaluate entirely uninvestigated alloy compositions. This hybrid approach will allow us to predict the corrosion behavior of novel systems containing elements absent from our training data, while still accounting for the critical microstructural context dictated by thermal and mechanical processing. SCALING DATA AND PHYSICS-INFORMED CONSTRAINTS To overcome the current tendency of the model to predict unrealistic alloying concentrations, such as excessive molybdenum or interstitial contents, we must integrate physics-informed constraints into the optimization algorithms. Imposing thermodynamic solubility limits and phase stability boundaries will naturally cap the optimization trajectories and prevent unphysical overshooting. Furthermore, we plan to scale our datasets through automated text mining of scientific literature. Expanding the training data volume will specifically resolve current outlier predictions in complex iron-based systems and improve the model's sensitivity to processing variations in aluminum alloys. CLOSING THE LOOP WITH HIGH-THROUGHPUT EXPERIMENTATION Ultimately, computational predictions must be validated and refined through targeted experimentation. We envision coupling our predictive deep learning framework with automated high-throughput experimental workflows, such as scanning droplet cells and rapid additive manufacturing. This closed-loop system will generate high-fidelity, standardized electrochemical data at scale, continuously feeding back into our models to iteratively reduce prediction errors and accelerate the targeted discovery of next-generation corrosion-resistant materials.

Professor Dierk Raabe
Max-Planck-Gesellschaft zur Forderung der Wissenschaften

Read the Original

This page is a summary of: Enhancing corrosion-resistant alloy design through natural language processing and deep learning, Science Advances, August 2023, American Association for the Advancement of Science,
DOI: 10.1126/sciadv.adg7992.
You can read the full text:

Read

Contributors

The following have contributed to this page