inklap

An Asynchronous Parallel Data Loading Optimization Algorithm for Deep Learning Applications

Xingjun Lin · Journal of Data Science and Intelligent Systems · 2026

This study aims to address the concern of inefficient data loading, which often leads to computationally inefficient deep learning workflows and becomes a bottleneck for scalability, especially under resource constraints. An asynchronous parallel data loading optimization algorithm is proposed that will revolutionize the data-training pipeline by enabling multi-threaded concurrent data loading and model training on multiple devices. The two-dimensional array structure and special hash table used ensure the invariance of the data distribution and concurrency safety, which is independent of loading and training processes and supported by rigorous mathematical proof. Experimental results from the CIFAR-10 dataset vividly demonstrate that this method represents a significant improvement over state-of-the-art baselines, achieving a throughput of approximately 3,250 samples/second, 87% GPU utilization, and a 40% reduction in training time, while maintaining the statistical integrity of the original dataset. This paper proposes a solution that does not bind users to a specific framework and increases efficiency without requiring the purchase of expensive hardware. This resolution makes de

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً