| Abstract Scope |
A central problem in science is to use noisy samples of an unknown function to predict function values for unseen inputs. In classical statistics, the predictive error is understood as a trade-off between the bias and the variance. However, overparametrized models exhibit counterintuitive behaviors, such as “double descent” in which models of increasing complexity exhibit decreasing generalization error. These behaviors are not explained by the dogma of bias-variance. We introduce a decomposition that we call the generalized aliasing decomposition (GAD) to explain the relationship between predictive performance and model complexity. The GAD decomposes the predictive error into three parts: (1) model insufficiency, which dominates when the number of parameters is much smaller than the number of data points, (2) data insufficiency, which dominates when the number of parameters is much greater than the number of data points, and (3) generalized aliasing. We demonstrate GAD in the context of a typical cluster expansion. Because key components of the generalized aliasing decomposition can be explicitly calculated from the relationship between model class and samples without seeing any data labels, it can answer questions related to experimental design and model selection before collecting data or performing experiments. |