inklap

AUGMENTED AND SYNTHETIC DATA IN ARTIFICIAL INTELLIGENCE

Philip de Melo · International Journal of Artificial Intelligence & Applications · 2025

High-quality data is essential for hospitals, public health agencies, and governments to improve services, train AI models, and boost efficiency. However, real data comes with challenges: strict privacy laws, high storage costs, legal constraints, and issues like bias or incompleteness. These can reduce the reliability of AI systems. As a result, artificial datasets are gaining importance. Synthetic and augmented data offer alternatives, yet their differences and potential are not fully understood. This paper examines how both types of data are generated and used, showcasing their characteristics through practical examples. Data generation techniques—such as Gaussian Mixture Models (GMM), Generative Adversarial Networks (GANs), and Gibbs sampling—enable the creation of realistic, privacy-preserving patient records that mimic the statistical properties of real data. Data augmentation, commonly used in image and signal analysis, is increasingly applied to structured electronic health records (EHRs), laboratory values, and time-series data to enhance model robustness and generalizability. This paper explores mathematical foundations, methodological frameworks, and real-world applicati

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً