3rd World Congress on Artificial Intelligence in Materials & Manufacturing (AIM 2025): Large Language Models - Information Extraction and Materials Design
Program Organizers: Remi Dingreville, Sandia National Laboratories; Ali Riza Durmaz, Fraunhofer Institute Iwm
Thursday 8:30 AM
June 19, 2025
Room: Elite Ballroom 1 & 2
Location: Anaheim Marriott
Session Chair: Remi Dingreville, Sandia National Laboratories
8:30 AM Cancelled
Physics-Aware Language Models for the Design and Processing of Lightweight Alloys: Avik Mahata1; 1Merrimack College
The integration of large language models (LLMs) into materials science offers new opportunities for lightweight alloy discovery and process-aware design. We present a physics-informed LLM framework that extracts structured knowledge, such as alloy composition, mechanical properties, and processing methods from unstructured scientific PDFs. The extracted data is formatted into a standardized alloy-specific language, which is then used to train an autoregressive model capable of learning composition–structure–property relationships. Unlike prior approaches limited to narrow datasets, our system generalizes across aluminum, magnesium, and titanium-based lightweight alloys. Physics-based rules and domain logic are incorporated to enhance interpretability and ensure valid outputs, such as identifying strengthening mechanisms or inferring processability via additive manufacturing (AM) methods like LPBF. Case studies demonstrate inverse querying e.g., retrieving printable Al-Zn alloys with target yield strength and composition suggestion for AM-critical issues like hot cracking. This framework highlights how LLMs, informed by materials knowledge, can drive scalable alloy innovation with built-in manufacturability guidance.
8:50 AM
Enhancing Data Acquisition in Manufacturing: Leveraging LLMs for Effective Material Property Databases: Inés Pérez Couñago1; Lara Suárez Casabiell1; Gabriel Novas Domínguez1; Santiago Muíños Landín1; Pilar Rey Rodriguez1; Félix Vidal Vilariño1; 1AIMEN Technology Center
Large Language Models (LLMs) enable the creation of up-to-date material property databases, offering significant potential for identifying materials with desired properties, proposing alternatives to costly options, resolving discrepancies in reported values, and automating inputs for tools like Finite Element Analysis (FEA). This work explores Multimodal Retrieval-Augmented Generation (RAG) for extracting information from Wire Arc Additive Manufacturing (WAAM) and Sheet Moulding Compound (SMC) documents. Given that material property data is often tabular, a benchmark was developed for parsing tools capable of handling diverse PDF layouts. Tables, images, and text were extracted using multimodal embeddings and table2text techniques, enabling the retrieval of relevant material information. The extracted data was processed into structured formats, including automatic detection and conversion into SI units, and LLM outputs were evaluated using standard metrics, as well as a custom metric. These advancements enhance workflow efficiency and decision-making by providing rapid, traceable access to manufacturing information.
9:10 AM
LLM-Assisted Data Curation in Starrydata: An Open Database of Material Properties Extracted From Published Plots: Yukari Katsura1; Tomoya Mato1; Yu Takada1; Dewi Yana1; Eiji Koyama1; Erina Fujita2; Yoshihiro Sakamoto3; Naoto Saito1; Fumikazu Hosono1; Atsumi Tanaka1; Masaya Kumagai4; 1National Institute for Materials Science; 2Institute of Statistical Mathematics; 3RIKEN; 4SAKURA Internet Inc.
We developed the Starrydata web system (https://www.starrydata2.org) as an open database of experimental materials data collected by tracing plot images from scientific literature. This platform enables users, including our data curators, to share experimental data extracted from published papers. The data hosted on Starrydata are publicly available and can be freely downloaded and used for both commercial and non-commercial purposes, provided our paper is cited. Starrydata includes datasets for functional materials such as thermoelectric, magnetic, and battery materials. The thermoelectric materials project comprises data for approximately 50,000 physical samples reported in around 10,000 papers, including more than 130,000 curve data on the temperature dependence of thermoelectric properties. We will present our efforts to accelerate data curation in Starrydata by developing an automated retrieval system using commercial Large Language Models (LLMs). Tasks such as figure labeling and the extraction of experimental processes from text have proven effective in assisting data curators.
9:30 AM Break
9:50 AM
Harnessing Large Language Models for Information Extraction and Material Design: Shuozhi Xu1; Xin Wang1; Kun Lu1; Haiming Wen2; 1University of Oklahoma; 2Missouri University of Science and Technology
Since the advent of ChatGPT in November 2022, the field of large language models (LLMs) has seen an explosive development. In the scientific community, LLMs have been employed to automate literature reviews, assist data analysis, and generate hypotheses. In the materials science community, there is a rapid surge in the interest in LLMs because the majority of information concerning materials exists as text, aligning closely with the text-centric nature of LLMs. In this work, we explore the abilities of LLMs for two tasks: information extraction and material design. In the first task, we employed an LLM to extract information including chemical compositions, processing conditions, microstructures, and properties. The LLM was shown to significantly outperform a conventional rule-based method. In the second task, we utilized an LLM to establish the relationship between structures and properties in a couple cases. The LLM-based relationship was compared favorably with machine-learning based ones.
10:10 AM
Utilizing Large Language Models for Interpreting & Interfacing With Instance Segmentation Results of Metallic Powder Micrographs: Stephen Price1; Kiran Judd1; Caroline Dowling1; Danielle Cote1; Kyle Tsaknopoulos1; 1Worcester Polytechnic Institute
Powder characterization is key to the efficient optimization of additive manufacturing techniques such as cold spray. To accomplish this, computer vision has proven to be effective at predicting and segmenting cold spray powder from SEM micrographs. However, while good at identifying and analyzing powders, outputs are large, complex, and difficult to analyze. This work presents a novel post-processing step utilizing a large language model (LLM) to interface with these outputs, making them more interpretable for real-world applications. This framework can process and analyze these results, answering practical questions about various morphological characteristics, such as area, eccentricity, sphericity, etc. Additionally, this approach can be used to compare multiple powders, enabling the user to query which characteristics are most similar or most different between samples. By embedding computer vision results in a readable format, an LLM improves the accessibility and interpretability of computer vision results for powder characterization.
10:30 AM Cancelled
Application of LLMs in Understanding Advanced Materials Properties and Manufacturing Processes: Jamiu Odusote1; Salihu Tanimowo1; Kamardeen Abdulrahman1; 1University of Ilorin
Traditional large language models (LLMs) have limited capabilities in comprehending specialized datasets related to materials or manufacturing fundamentals, such as microstructures, phase changes, and thermomechanical properties. This study aims to integrate multimodal data; images, sensor readings, and numerical data, with textual descriptions for better understanding of Advanced Materials Properties and Manufacturing Processes (AMPMP). A complete picture of the domain was constructed when multiple modalities is combined, enabling the development of rich representations that capture the complexities inherent in AMPMP. The results showed that multimodal approaches have potential to significantly improve the modeling and analysis of these specialized domains compared with LLMs. Integration of diverse data sources could lead to models with lower variance and improved generalization than those that rely solely on a single modality. This means that a multimodal approach can provide a more comprehensive understanding of AMPMP, with practical implications for various applications in materials science and engineering.
10:50 AM
Salvaging of Materials Fatigue Data from Literature Using Language Model Systems: Ali Riza Durmaz1; Jyoti Mohanty2; Akhil Thomas3; 1University of California - Santa Barbara, Fraunhofer IWM; 2Fraunhofer IWM; 3Fraunhofer Iwm
Materials fatigue has been researched since the 18th century which has led to a wealth of data and information contained in scientific literature. Associated cyclic testing of a single material often requires months to years. Thus, design of safety-critical components is often constrained to few thoroughly characterized materials or relies on rather crude estimates. Comprehensive extraction of fatigue information from publications into a structured and harmonized format could support building predictive models in the future.