Skip to main navigation menu Skip to main content Skip to site footer

Articles

Vol. 10 No. 6 (2025): Kohesi: Jurnal Sains dan Teknologi, ISSN 3025-1311

KLASIFIKASI KATEGORI WEBSITE DENGAN MENGGUNAKAN NAÏVE BAYES DAN SUPPORT VECTOR MACHINES

Submitted
November 15, 2025
Published
2025-11-16

Abstract

Seiring ledakan informasi pada abad ke-21, kelimpahan informasi menjadi tantangan utama di internet. Meskipun search engine membantu pengguna menilai nilai situs web berdasarkan topiknya, menemukan informasi spesifik tetap menantang. Penelitian ini mengusulkan pendekatan praktis dengan membuat catatan data otomatis yang merangkum konten setiap situs web untuk tujuan kategorisasi. Tugas klasifikasi teks (TC) secara otomatis menetapkan dokumen ke kategori tertentu. Penelitian ini fokus pada metode klasifikasi teks, termasuk Naïve Bayes dan Support Vector Machines (SVM), untuk mencapai akurasi yang baik. Selain dua metode tersebut, penelitian merinci penggunaan metode lain seperti K-Nearest Neighbors (KNN) dan Random Forest dalam konteks klasifikasi web phishing dan analisis sentimen. Dalam eksperimen ini, Naïve Bayes dan SVM dievaluasi untuk mengklasifikasikan kategori situs web. Data dibagi menjadi pelatihan (70%) dan pengujian (30%), dengan hasil akurasi sekitar 88% untuk Naïve Bayes dan 85% untuk SVM. Studi ini memberikan pemahaman lebih dalam tentang kinerja dua metode klasifikasi berbeda dalam konteks pengkategorian situs web.

References

  1. [1] S. Lugeon, T. Piccardi, and R. West, “Homepage2Vec: Language-Agnostic Website Embedding and Classification,” Proceedings of the International AAAI Conference on Web and Social Media, vol. 16, 2022, doi: 10.1609/icwsm.v16i1.19380.
  2. [2] G. Gan, F. Songyuan, and Y. Junwen, “Development of a Website Classification Model for Quiet Media,” Saint Petersburg State University, Saint Petersburg, 2022.
  3. [3] O. W. Kwon and J. H. Lee, “Text categorization based on k-nearest neighbor approach for Web site classification,” Inf Process Manag, vol. 39, no. 1, pp. 25–44, Jan. 2003, doi: 10.1016/S0306-4573(02)00022-5.
  4. [4] R. Bruni and G. Bianchi, “Robustness analysis of a Website categorization procedure based on Machine Learning,” Dep. of Computer Control and Management Engineering, Sapienza University of Rome, Italy, pp. 1–25, 2018, [Online]. Available: http://www.dis.uniroma1.it/%7B~%7Dbruni/files/bruni18robustness.pdf
  5. [5] R. Bruni and G. Bianchi, “Website categorization: A formal approach and robustness analysis in the case of e-commerce detection,” Expert Syst Appl, vol. 142, Mar. 2020, doi: 10.1016/j.eswa.2019.113001.
  6. [6] Y. HaCohen-Kerner, D. Miller, and Y. Yigal, “The influence of preprocessing on text classification using a bag-of-words representation,” PLoS One, vol. 15, no. 5, 2020, doi: 10.1371/journal.pone.0232525.
  7. [7] A. Gasparetto, M. Marcuzzo, A. Zangari, and A. Albarelli, “Survey on Text Classification Algorithms: From Text to Predictions,” Information (Switzerland), vol. 13, no. 2, Feb. 2022, doi: 10.3390/info13020083.
  8. [8] P. J. Uppalapati, B. K. Gontla, P. Gundu, S. M. Hussain, and K. Narasimharo, “A Machine Learning Approach to Identifying Phishing Websites: A Comparative Study of Classification Models and Ensemble Learning Techniques,” ICST Transactions on Scalable Information Systems, Jun. 2023, doi: 10.4108/eetsis.vi.3300.
  9. [9] A. S. Neogi, K. A. Garg, R. K. Mishra, and Y. K. Dwivedi, “Sentiment analysis and classification of Indian farmers’ protest using twitter data,” International Journal of Information Management Data Insights, vol. 1, no. 2, Nov. 2021, doi: 10.1016/j.jjimei.2021.100019.
  10. [10] L. Chen, L. Jiang, and C. Li, “Modified DFS-based term weighting scheme for text classification,” Expert Syst Appl, vol. 168, 2021, doi: 10.1016/j.eswa.2020.114438.
  11. [11] S. Gan, S. Shao, L. Chen, L. Yu, and L. Jiang, “Adapting hidden naive bayes for text classification,” Mathematics, vol. 9, no. 19, Oct. 2021, doi: 10.3390/math9192378.
  12. [12] H. Zhang, L. Jiang, and L. Yu, “Attribute and instance weighted naive Bayes,” Pattern Recognit, vol. 111, 2021, doi: 10.1016/j.patcog.2020.107674.
  13. [13] P. S. Zalukhu, T. Handhayani, and M. Sitorus, “ANALISIS SENTIMEN TERHADAP KENAIKAN BBM DI INDONESIA PADA MEDIA SOSIAL TWITTER MENGGUNAKAN METODE NAÏVE BAYES,” Simtek : jurnal sistem informasi dan teknik komputer, vol. 8, no. 1, 2023, doi: 10.51876/simtek.v8i1.177.
  14. [14] H. Kim, P. Rowland, and H. Park, “Dimension reduction in text classification with support vector machines,” Journal of Machine Learning Research, vol. 6, 2005.
  15. [15] A. Wibowo Haryanto, E. Kholid Mawardi, and Muljono, “Influence of Word Normalization and Chi-Squared Feature Selection on Support Vector Machine (SVM) Text Classification,” in Proceedings - 2018 International Seminar on Application for Technology of Information and Communication: Creative Technology for Human Life, iSemantic 2018, Institute of Electrical and Electronics Engineers Inc., Nov. 2018, pp. 229–233. doi: 10.1109/ISEMANTIC.2018.8549748.
  16. [16] L. Rosiana, “Analisis Kemungkinan Keterlambatan Pembayaran SPP Menggunakan AlgoritmaSupport Vector Machine(Studi Kasus: Smp Perintis 2 Bandar Lampung),” Jurnal Ilmu Data, vol. 2, no. 9, 2022.
  17. [17] A. Dhar, H. Mukherjee, N. S. Dash, and K. Roy, “Text categorization: past and present,” Artif Intell Rev, vol. 54, no. 4, 2021, doi: 10.1007/s10462-020-09919-1.
  18. [18] B. P. Yadav, S. Ghate, A. Harshavardhan, G. Jhansi, K. S. Kumar, and E. Sudarshan, “Text categorization Performance examination Using Machine Learning Algorithms,” in IOP Conference Series: Materials Science and Engineering, 2020. doi: 10.1088/1757-899X/981/2/022044.
  19. [19] A. Shahzad et al., “COVID-19 Vaccines Related User’s Response Categorization Using Machine Learning Techniques,” Computation, vol. 10, no. 8, Aug. 2022, doi: 10.3390/computation10080141.
  20. [20] R. Alshammari, “Arabic Text categorization using machine learning approaches,” International Journal of Advanced Computer Science and Applications, vol. 9, no. 3, 2018, doi: 10.14569/IJACSA.2018.090332.
  21. [21] S. M. H. Mahmud, M. A. Hossin, H. Jahan, S. R. H. Noori, and T. Bhuiyan, “CSV-ANNOTATE: Generate annotated tables from CSV file,” in 2018 International Conference on Artificial Intelligence and Big Data, ICAIBD 2018, 2018. doi: 10.1109/ICAIBD.2018.8396169.
  22. [22] J. Koo, S. Baek, and S. Kim, “The effect of personal value on CSV (Creating Shared Value),” Journal of Open Innovation: Technology, Market, and Complexity, vol. 5, no. 2, 2019, doi: 10.3390/JOITMC5020034.
  23. [23] A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. Hanna, “Data and its (dis)contents: A survey of dataset development and use in machine learning research,” Patterns, vol. 2, no. 11. Cell Press, Nov. 12, 2021. doi: 10.1016/j.patter.2021.100336.
  24. [24] Y. Gong, G. Liu, Y. Xue, R. Li, and L. Meng, “A survey on dataset quality in machine learning,” Inf Softw Technol, vol. 162, Oct. 2023, doi: 10.1016/j.infsof.2023.107268.

Similar Articles

21-30 of 145

You may also start an advanced similarity search for this article.