inklap

Towards High Performance Data Curation Statistical Disclosure Control Tooling

Deirdre Lungley, Simon George · International Journal of Population Data Science · 2025

ObjectivesSurveys are a widely used and important research resource, whose creation and curation involve skilled, labour-intensive tasks. This abstract details an initiative to improve the tooling available to the data community, to control the risk of unintended disclosure, in line with the Anonymisation Decision Making Framework. MethodsAn initial step in assessing dataset disclosivity is the identification of key variables (KVs) (variables which, when combined, can indicate individual units) and the subsequent computation of frequency counts for combinations of these variables. These counts are a prerequisite for achieving k-anonymity, but can also be used in further risk calculations. Their centrality to our processes prompted us to improve the performance of this algorithm. We achieved a significant improvement over the original sdcMicro R package. We do this by decomposing KV values into "bitmasks" (0s and 1s) that are then easily manipulable by native CPU instructions. ResultsIn the sdcMicro R package these calculations use the data.table library which, while performant, can be improved upon by our algorithm especially in the common case of the dataset containing missing v

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً