This research develops a voice-based Quranic verse detection application using an R-CNN machine learning model adapted for audio data. Focusing on Surah An-Naba (40 verses), the system converts audio recordings into visual representations in the form of spectrograms and Mel Frequency Cepstral Coefficients (MFCC), which are then processed using Conv1D and BiLSTM architectures. The dataset includes 40 verses in various sound qualities and intonations. The implementation results show a training accuracy of 85.12%, validation of 51.16%, and testing of 40.86% for Top-1 Prediction. However, Top-3 Accuracy reached 72.09%, indicating that 30 of the 40 verses were successfully recognized in the top three predictions. The system is integrated with a web interface and API to enable audio uploads, automatic tajwid analysis, and verse identification. This research shows that the R-CNN approach adapted to audio data has significant potential to improve the accessibility of the Quran, especially for users with visual or literacy limitations.
You may also start an advanced similarity search for this article.