inklap

APPLICATION OF NEURAL NETWORKS FOR AUTOMATED TEXT AND SYMBOL RECOGNITION IN CYBERSECURITY TASKS

Nataliia Cherniiashchuk · Cybersecurity: Education, Science, Technique · 2025

This work presents the development and investigation of an Optical Character Recognition (OCR) system for low-quality text using machine learning methods. To address the task, two grayscale image datasets were created: the first consisting of isolated English alphabet characters and digits (4,960 images of 250×50 pixels), and the second containing fragments of meaningful text from the book The Hunger Games (4,010 images of 680×50 pixels). To enhance model robustness, the images were distorted using blur and digital noise functions. The OCR model was built based on a combination of Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) with a Connectionist Temporal Classification (CTC) layer for sequence correction. Training was conducted over 70–80 epochs with a 9:1 split for training and validation datasets. A comparative analysis was carried out between the developed system and Tesseract OCR. Experimental results demonstrated that the proposed model achieves superior recognition performance on low-quality images, particularly those affected by digital noise, whereas Tesseract OCR significantly loses accuracy under such conditions. The results confirm the effectiv

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً