inklap

An Integrated Machine Learning Pipeline for Genome-Level Bacterial Identification and In Silico Prioritization of Candidate Antibacterial and Antifungal Proteins from Public Microbial Sequence Data Resources

Ravi Dayani · International Journal of Artificial Intelligence, Data Science, and Machine Learning · 2023

This paper introduces an integrated machine learning workflow that first classifies bacterial nucleotide sequences and then performs in silico prioritization of proteins that may warrant follow-up as antibacterial or antifungal candidates. Public nucleotide records are collected automatically from NCBI and organized into labeled classes for model development. Sequence fragments are encoded with compositional descriptors and used to train a supervised bacterial identifier with probability-based confidence reporting and a simple novelty-screening step. The trained classifier achieved 0.9904 accuracy on held-out fragments, with weighted precision, recall, and F1-scores of 0.9905, 0.9904, and 0.9887, respectively. The framework was further applied to the NCBI case-study genome LJCX01000023.1, which was assigned to the Mycobacterium smegmatis class with mean confidence of 0.7789 across 450 fragments. To extend the analysis beyond organism labeling, predicted open reading frames from the case-study genome are translated and ranked using sequence-informed heuristics for downstream antimicrobial screening. The study offers a reproducible computational route from public genome retrieval to

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً