| Abstract Scope |
When measurements are scarce and costly, materials development advances through new experiments or educated guesses from existing data. Semisupervised pseudo-labeling and active learning formalize these routes. This work compares them within one calibrated Gaussian process regression framework, taking nepheline crystallization in high-level waste glasses as the case study, with 154 quantified compositions and simulated oracle experiments. Starting from a large labeled base, self-training absorbed 68 unquantified glasses and cut uncertainty from 18.5 to 3.2 vol% without hurting accuracy. Starting from 30 glasses, it pushed the error 1.3 vol% above the seed-only baseline and left the model overconfident. Active learning, ranking candidates by uncertainty in the space where the model is trained, held a mid-budget advantage near 1 vol% over random sampling, and thirty targeted measurements beat the free semisupervised route by 3.5 vol%. Which route to trust depends on the strength of the labeled base. |