Significant advancements in machine learning have accelerated and improved structure–property predictions for materials discovery. However, data are often scarce due to large parameter spaces consisting of chemistry, structure, synthesis, and processing variables. At small data limits, theory can be leveraged to inform machine learning (ML) models with domain knowledge and improve generalization. Here, we determine how the accuracy of first-principles calculations affects theory-informed ML predictions of the experimental solvation free energy ΔGsolvexp in both “small” (101–102) and “large” (103–104) data size limits. We compare several existing theory-informed techniques to a baseline (no theory) model: feature-informed, difference, and ratio. At small data limits, all theory-informed models exhibit lower RMSE, reducing training data size by more than 65% compared to the baseline model. With larger training set sizes and as theory prediction accuracy declines, the difference and ratio models exhibit larger errors than the baseline model, indicating negative transfer occurs. No negative transfer is observed in the feature-informed model; however, the model is unable to extrapolate
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً