Data drift refers to the change in the statistical properties of input data; this data is collected for the training of a mathematical model. We investigated the impact of drift on data linkage models (with further interest in machine-learning models), by researching how to determine drift existence and severity, impact on linkage, and how to mitigate against it. A literature review was conducted of scientific papers and journals focusing on these three key areas. This review was done to assess the true potential impact of drift on a linkage model, and about what steps should be taken if it can be properly mitigated against. The key findings are as follows: There are different types of drift – such as data, concept and performance. Research suggests it is not only necessary to determine the existence of drift, but also its type. Drift can critically damage model performance, and conditions that generate drift can be found in both linkage and non-linkage contexts. In particular, due to the potentially high level of mess, complexity and error that can come with it, moving towards big data could prove challenging and can result in poor linkage results; its non-stationary natur
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً