inklap

Automating production of de-identified linkable data

Madalina Iova, Rachel Huck · International Journal of Population Data Science · 2024

ObjectiveAn innovative large-scale automated method has been developed to produce de-identified linkable data. The objective is to create a wide pool of ready-to-use data to enable faster and wider collaborative analysis for the public good. ApproachA configurable automated pipeline prepares data for onward linkage at location, person, business and classification level through: Big data profiling pre- and post-processing, for overview of variables and characteristics Flagging potentially sensitive/identifiable variables Generalisable linkage methods for large-scale data, to enable the addition of unique IDs for onward linkage of de-identified data De-identification, hashing and redaction mechanisms, to remove and/or obscure sensitive/identifiable variables Automated production of metadata, capturing linkage quality and transformations across the data journey Quality assurance checks, including measure of linkage quality, assurance of variable derivations and redactions, and consistency checks on remaining data. ResultsThe pipeline enables a configurable automated approach to producing de-identified, linkable, ready-to-use data in a traceable and fully documented manner.

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً