inklap

Using SQL for accessing and working with large administrative data in TREs

Tudor Vilcan · International Journal of Population Data Science · 2025

ObjectivesThe objective of this presentation is two-fold. First, it will explain why SQL is needed for large admin data provision and how this is operationalised in practice in TREs. Second, it will provide researchers with guidance and instructions for manipulating, cleaning and analysing data in SQL. MethodAdministrative datasets such as the invaluable Longitudinal Education Outcomes (LEO) are characterized by very large size and number of variables, as well as deep row counts. This entails it is not suitable to be provisioned via means such as flat files and software more familiar to researchers, such as STATA, SPSS or R. SQL is a valuable and sometimes indispensable tool for the provision of such data, as it can provide adequate storage, selective access to prevent disclosure risk and tools for data manipulation. The presentation will also cover weaknesses of SQL and when other software is more appropriate to use. ResultsThrough this presentation, we hope to ensure researchers will have a much better understanding of how and why data is provisioned through SQL, especially large administrative datasets. They will also learn techniques and principles of SQL usage, alongside use

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً