inklap

Talk data to me! Evaluating the potential for large language models to enhance data discoverability across UKRI’s federated data services

Mark Green, Maura Halstead, Cillian Berragan, Caroline Jay, David Topping, Richard Kingston · International Journal of Population Data Science · 2025

ObjectivesThe talk will: (i) Describe the development of a large language model (LLM) powered semantic search tool for UKRI data catalogues, and (ii) Examine the concerns and opportunities of using this tool among researchers for data discovery. MethodsA semantic search tool was developed integrating the data catalogues of Administrative Data Research UK, Consumer Data Research Centre, and UK Data Service. We used OpenAI’s vector embedding service to convert these metadata into embeddings, allowing natural language search to be used rather than keywords only. We assessed the acceptability and suitability of this tool using four focus groups. Participants were recruited across academic researchers, PhD researchers, data services staff, and local government / third sector analysts (n=36). Data collected from focus groups were analysed using thematic analysis. ResultsThe key themes identified in focus groups were: (i) Current data discovery techniques are dependent on keyword strategies for searching (including the dominance of using Google). There is need to support training for using any LLM based resources. (ii) There was low trust of LLMs, especially in academic researchers. Pa

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً