International Journal of Innovative Research in Computer and Communication Engineering
ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines
| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |
| TITLE | Automated Summarization Tool |
|---|---|
| ABSTRACT | As the size of file is increasing in non-structured text format in academic, corporate, law and organisation, there is a severe need of information-extraction and text-summarization systems. Doing summary manually is a time consuming, inconsistent and unscalable process. This paper presents the design realization and evaluation of an Automated Summarization Tool (AST) which is a document intelligence platform based on google gemini 2.5 flash. The platform employs map-reduce summarization for long documents, the use of SHA-256 hash for caching, a RAG-lite chat module grounded in the source document enabling conversational chats, role- and tone-adaptive prompt engineering, and a JSON-based output schema coupling every extracted key point with a verbatim quote and location from the source for traceability, thereby preventing unnecessary API calls. The complete system features a Gradio web interface. It is deployed as a zero-infrastructure Google Colab notebook. Thus, no dedicated server or installation is required. In internal, multi-domain test corpora, our platform achieves a ROUGE-1 score of 58.17. This is the highest we obtain out of six different summarization systems with which we compare against other methods. And it outperforms the strongest fine-tuned transformer baselines (PEGASUS, BART) by ~14 points. And it is clearly ahead of BERTSUM-ext (a strong transformer baseline), Pointer-Generator Network, TextRank. |
| AUTHOR | KESHAV KUMAR, AMANDEEP, DHARMENDER KUMAR, ANKIT, SURAJ Dept. of M.Sc. Computer Science, Artificial Intelligence and Data Science, GJUS&T, Hisar, Haryana, India Assistant Professor, Dept. of Artificial Intelligence and Data Science, GJUS&T, Hisar, Haryana, India Professor, Dept. of Artificial Intelligence and Data Science, GJUS&T, Hisar, Haryana, India |
| VOLUME | 186 |
| DOI | DOI: 10.15680/IJIRCCE.2026.1407067 |
| pdf/67_Automated Summarization Tool.pdf | |
| KEYWORDS | |
| References | [1] H. P. Luhn, “The automatic creation of literature abstracts,” IBM Journal of Research and Development, vol. 2, no. 2, pp. 159–165, 1958. [2] R. Mihalcea and P. Tarau, “TextRank: Bringing order into text,” in Proc. EMNLP 2004, pp. 404–411, 2004. [3] R. Nallapati, F. Zhai, and B. Zhou, “SummaRuNNer: A recurrent neural network based sequence model for extractive summarization,” in Proc. AAAI 2017. [4] A. Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, 2017. [5] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers,” in Proc. NAACL 2019, pp. 4171–4186. [6] Y. Liu, “Fine-tune BERT for extractive summarization,” arXiv preprint arXiv:1903.10318, 2019. [7] L. Ouyang et al., “Training language models to follow instructions with human feedback,” Advances in NeurIPS, vol. 35, 2022. [8] A. M. Rush, S. Chopra, and J. Weston, “A neural attention model for abstractive sentence summarization,” in Proc. EMNLP 2015, pp. 379–389. [9] A. See, P. J. Liu, and C. D. Manning, “Get to the point: Summarization with pointer-generator networks,” in Proc. ACL 2017, pp. 1073–1083. [10] M. Lewis et al., “BART: Denoising sequence-to-sequence pre-training,” in Proc. ACL 2020, pp. 7871–7880. [11] J. Zhang et al., “PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization,” in Proc. ICML 2020. [12] T. B. Brown et al., “Language models are few-shot learners,” Advances in NeurIPS, vol. 33, 2020. [13] A. Nenkova and K. McKeown, “Automatic Summarization,” Foundations and Trends in Information Retrieval, vol. 5, nos. 2–3, pp. 103–233, 2011. [14] N. F. Liu et al., “Lost in the middle: How language models use long contexts,” TACL, vol. 12, pp. 157–173, 2023. |