inklap

Creating a Data Cleaning and Pre-Processing Module for Generalisable Data Linkage

Josie Plachta, Mary Cleaton, Leah Quinn, Alex Mackay, Zoe White · International Journal of Population Data Science · 2024

ObjectiveThe Office for National Statistics (ONS) are developing a generalisable tool to facilitate the linkage of various datasets to its population-spine. However, a generalisable process requires that a variety of input datasets can be adaptively pre-processed – which is a problem for bespoke methodologies. Key requirements of the cleaning pipeline include minimal input from the user, and scalability and efficiency to work on Big Data. ApproachThe pipeline must recognise and adjust the pre-processing steps applied based on the variables present and user requirements. It must accept, preprocess, and derive consistent standardised variables from a variety of input variables and formats, including complex data characteristics. ResultsThe MVP pipeline successfully met the requirements. It is based on a three-level hierarchy of functions, allowing flexibility and complexity in data preparation. With minimal user input, a variety of important linkage variables are cleaned, and additional variables derived consistently. Conclusions & ImplicationsThe module has shown promising results at scale, successfully pre-processing datasets of over 91 million records. It will be a valuable

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً