| Abstract Scope |
Previously, we have trained our machine-learning models on the vast starrydata2 database, to predict the key thermoelectric properties based on composition only. To fine-tune the predictions, we have downloaded ~10,000 publications in .pdf format, and extracted a number of fabrication details for each material discussed therein using large language models. Via feature engineering, we then selected the most important ones to be included in our machine-learning algorithm. During this talk, I will present the importance of these various parameters, and how their inclusion led to which improvements of the predictive power using different unrelated test sets of truly unseen data. |