Articles
Vol. 10 No. 8 (2026): Kohesi: Jurnal Sains dan Teknologi, ISSN 3025-1311
DATA CRAWLING TECHNIQUES IN SOCIAL MEDIA AS AN INITIAL STAGE IN UNDERSTANDING BASIC MACHINE LEARNING TECHNIQUES (CASE STUDY: RAM PRICE INCREASE 2025)
-
Submitted
-
January 7, 2026
-
Published
-
2026-01-07
Abstract
Social media is an important data source for understanding public responses to technological issues. This study aims to apply data crawling techniques as an initial stage of basic machine learning with a case study of the increase in Random Access Memory (RAM) prices in 2025. Data was collected from the Twitter (X) platform using the Twitter API, then went through pre-processing, tokenization, and token frequency analysis stages. The results of the study show that pre-processing effectively improves the quality of text data, while token frequency analysis is able to identify dominant words that represent the focus of public discussion. This study confirms that data crawling and token frequency analysis are important foundations in the initial pipeline of text-based machine learning.
References
- [1] N. Van Robert Jhosefhin et al., “Analisis Sentimen Crawling Data dari Sosial Media X tentang Gaza Menggunakan Metode SVM dan Decision Tree,” 2025. [Online]. Available: https://journal.stmiki.ac.id
- [2] M. A. Khder, “Web scraping or web crawling: State of art, techniques, approaches and application,” International Journal of Advances in Soft Computing and its Applications, vol. 13, no. 3, pp. 144–168, 2021, doi: 10.15849/ijasca.211128.11.
- [3] S. S. Sohail et al., “Crawling Twitter data through API: A technical/legal perspective,” May 2021, [Online]. Available: http://arxiv.org/abs/2105.10724
- [4] T. Agustiranti, A. Khalfani Izzati Kurdiana, B. Al Ghiffari, E. Dwi Juniar, and D. Gita Purnama, “Penerapan Naive Bayes Terhadap Sentimen Analisis Media Sosial Twitter Pengguna Kereta Cepat Jakarta-Bandung (Whoosh),” Jurnal Ilmu Komputer dan Sistem Informasi (JIKOMSI, vol. 7, no. 1, pp. 297–305, 2024.
- [5] S. Suhendra and F. Selly Pratiwi, “Peran Komunikasi Digital dalam Pembentukan Opini Publik: Studi Kasus Media Sosial,” Iapa Proceedings Conference, p. 293, Oct. 2024, doi: 10.30589/proceedings.2024.1059.
- [6] “Perancangan Aplikasi Web Crawler untuk Menghasilkan Dokumen Teks pada Domain Tertentu”.
- [7] “ANALISIS SENTIMEN DATA TWITTER TERHADAP PELAKSANAAN PEMBELAJARAN ONLINE DI INDONESIA PADA MASA PANDEMI COVID-19 MENGGUNAKAN METODE NATURAL LANGUAGE PROCESSING”.
- [8] P. Deski Manalu et al., “Implementasi Algoritma Klasifikasi untuk Analisis Sentimen Media Sosial Tiktok Tahun 2025,” Jurnal Teknik Informatika dan Teknologi Informasi, doi: 10.55606/jutiti.v5i1.5644.
- [9] U. Khairani, V. Mutiawani, and H. Ahmadian, “Pengaruh Tahapan Preprocessing Terhadap Model Indobert Dan Indobertweet Untuk Mendeteksi Emosi Pada Komentar Akun Berita Instagram,” Jurnal Teknologi Informasi dan Ilmu Komputer, vol. 11, no. 4, pp. 887–894, Aug. 2024, doi: 10.25126/jtiik.1148315.
- [10] “ANALISIS SENTIMEN PADA KOMENTAR INSTAGRAM MENGGUNAKAN GAUSSIAN NAÏVE BAYES”.
- [11] A. Liken Anggoro, L. V. P. Ken, and M. G. Setiawan, “Analisis Media Text Clustering pada Twitter Akan Kasus Selebriti Menggunakan Orange Data Mining,” remik, vol. 7, no. 1, pp. 189–195, Jan. 2023, doi: 10.33395/remik.v7i1.12001.