International Journal of Innovative Research in Computer and Communication Engineering

ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines

| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |


TITLE Detection of Online Spread Terrorism using Web Data Mining
ABSTRACT The unchecked growth of internet-based platforms has made it far easier for extremist groups to publish propaganda, spread radical ideology, and recruit sympathisers online. Manual monitoring of this content is no longer practical because of the sheer scale and speed at which new web pages appear, and traditional keyword filters tend to miss context, sarcasm, and coded language. This paper presents a lightweight, automated pipeline that combines web mining, Natural Language Processing (NLP) and a Multinomial Naive Bayes classifier to examine the text of a given webpage and estimate how likely it is to contain terrorism-related material. A user simply submits a URL the system fetches the page, strips away scripts, navigation bars and other non-content elements using BeautifulSoup, cleans the remaining text (lower-casing, stop-word removal, tokenisation), converts it into a TF-IDF feature vector, and passes that vector to a Naive Bayes model trained on a labelled dataset of six content categories — Terrorism, Riot, Disaster, Protest, Political and Positive/Safe. The predicted category is then mapped onto a four-tier risk scale (Critical, High, Medium, Low) along with a confidence score, a short list of dominant keywords, and basic page statistics. Testing on a small set of live URLs shows that the pipeline behaves consistently and produces interpretable results, achieving an overall test accuracy of roughly 93% on the held-out dataset. The paper also compares this shallow, resource-efficient approach against heavier deep-learning alternatives reported in recent literature and outlines directions — multilingual support, transformer-based models, and real-time dashboards — for extending the system in future work.
AUTHOR ANUSHA P, YASHASWINI J IV Semester MCA, Dept. of DoS in Computer Sceince, PG Wing of SBRR Mahajana First Grade College (Autonomous), Mysuru, India Assistant Professor, Dept. of DoS in Computer Sceince, PG Wing of SBRR Mahajana First Grade College (Autonomous), Mysuru, India
VOLUME 186
DOI DOI: 10.15680/IJIRCCE.2026.1407022
PDF pdf/22_Detection of Online Spread Terrorism using Web Data Mining.pdf
KEYWORDS
References [1] A. Riabi, V. Mouilleron, M. Mahamdi, W. Antoun, and D. Seddah, "Beyond dataset creation: Critical view of annotation variation and bias probing of a dataset for online radical content detection," arXiv preprint arXiv:2502.01234, 2025.
[2] C. de Kock, A. Riabi, Z. Talat, M. Schlichtkrull, P. Madhyastha, and E. Hovy, "Using language models to decode extremist cryptolects," Proc. 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025.
[3] M. D. Alshehri, F. Iqbal, B. C. M. Fung, and M. Debbabi, "Cybersecurity intelligence through textual data analysis: A framework using machine learning and terrorism datasets," Expert Systems with Applications, vol. 238, 2025.
[4] R. Scrivens, J. D. Freilich, S. M. Chermak, and R. Frank, "Data collection in online terrorism and extremism research: Strengths, limitations, and future directions," Terrorism and Political Violence, vol. 36, no. 5, pp. 601–620, 2024.
[5] A. J. Navnath, B. A. Balu, S. Tejal, S. R. Dilip, and R. M. Dhokane, "To detect the terrorism activity," Int. J. Advanced Research in Computer Science and Software Engineering, vol. 13, no. 4, pp. 45–52, 2023.
[6] S. Ahmad, M. Z. Asghar, F. M. Alotaibi, and I. Awan, "Detection and classification of social media-based extremist affiliations using sentiment analysis techniques," Human-Centric Computing and Information Sciences, vol. 9, no. 1, p. 24, 2019.
[7] S. A. Azizan and I. E. Aziz, "Terrorism detection based on sentiment analysis using machine learning," J. Engineering and Applied Sciences, vol. 12, no. 3, pp. 691–698, 2017.
[8] W. You, L. Khan, and B. Thuraisingham, "Identification of extremism on Twitter," Proc. IEEE Int. Conf. on Intelligence and Security Informatics (ISI), pp. 223–228, 2016.
[9] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," Proc. NAACL-HLT, pp. 4171–4186, 2019.
[10] S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit. Sebastopol, CA: O'Reilly Media, 2009.
[11] C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge, UK: Cambridge University Press, 2008.
[12] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
[13] A. McCallum and K. Nigam, "A comparison of event models for Naive Bayes text classification," Proc. AAAI-98 Workshop on Learning for Text Categorization, vol. 752, pp. 41–48, 1998.
[14] T. Joachims, "Text categorization with Support Vector Machines: Learning with many relevant features," Proc. European Conf. on Machine Learning (ECML), pp. 137–142, 1998.
[15] B. Liu, Web Data Mining: Exploring Hyperlinks, Contents, and Usage Data, 2nd ed. Berlin, Germany: Springer-Verlag, 2011.
[16] R. Richardson, Beautiful Soup Documentation. Crummy.com, 2023.
[17] K. Reitz, Requests: HTTP for Humans. Python Software Foundation, 2023.
[18] Y. LeCun, Y. Bengio, and G. Hinton, "Deep learning," Nature, vol. 521, no. 7553, pp. 436–444, 2015.
[19] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed. Waltham, MA: Morgan Kaufmann, 2011.
[20] Y. Yang and X. Liu, "A re-examination of text categorization methods," Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 42–49, 1999.
Copyright © IJIRCCE 2020.All right reserved