International Journal of Innovative Research in Computer and Communication Engineering

ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines

| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |


TITLE Automated Summarization Tool
ABSTRACT As the size of file is increasing in non-structured text format in academic, corporate, law and organisation, there is a severe need of information-extraction and text-summarization systems. Doing summary manually is a time consuming, inconsistent and unscalable process. This paper presents the design realization and evaluation of an Automated Summarization Tool (AST) which is a document intelligence platform based on google gemini 2.5 flash. The platform employs map-reduce summarization for long documents, the use of SHA-256 hash for caching, a RAG-lite chat module grounded in the source document enabling conversational chats, role- and tone-adaptive prompt engineering, and a JSON-based output schema coupling every extracted key point with a verbatim quote and location from the source for traceability, thereby preventing unnecessary API calls. The complete system features a Gradio web interface. It is deployed as a zero-infrastructure Google Colab notebook. Thus, no dedicated server or installation is required. In internal, multi-domain test corpora, our platform achieves a ROUGE-1 score of 58.17. This is the highest we obtain out of six different summarization systems with which we compare against other methods. And it outperforms the strongest fine-tuned transformer baselines (PEGASUS, BART) by ~14 points. And it is clearly ahead of BERTSUM-ext (a strong transformer baseline), Pointer-Generator Network, TextRank.
AUTHOR KESHAV KUMAR, AMANDEEP, DHARMENDER KUMAR, ANKIT, SURAJ Dept. of M.Sc. Computer Science, Artificial Intelligence and Data Science, GJUS&T, Hisar, Haryana, India Assistant Professor, Dept. of Artificial Intelligence and Data Science, GJUS&T, Hisar, Haryana, India Professor, Dept. of Artificial Intelligence and Data Science, GJUS&T, Hisar, Haryana, India
VOLUME 186
DOI DOI: 10.15680/IJIRCCE.2026.1407067
PDF pdf/67_Automated Summarization Tool.pdf
KEYWORDS
References [1] H. P. Luhn, “The automatic creation of literature abstracts,” IBM Journal of Research and Development, vol. 2, no. 2, pp. 159–165, 1958.
[2] R. Mihalcea and P. Tarau, “TextRank: Bringing order into text,” in Proc. EMNLP 2004, pp. 404–411, 2004.
[3] R. Nallapati, F. Zhai, and B. Zhou, “SummaRuNNer: A recurrent neural network based sequence model for extractive summarization,” in Proc. AAAI 2017.
[4] A. Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, 2017.
[5] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers,” in Proc. NAACL 2019, pp. 4171–4186.
[6] Y. Liu, “Fine-tune BERT for extractive summarization,” arXiv preprint arXiv:1903.10318, 2019.
[7] L. Ouyang et al., “Training language models to follow instructions with human feedback,” Advances in NeurIPS, vol. 35, 2022.
[8] A. M. Rush, S. Chopra, and J. Weston, “A neural attention model for abstractive sentence summarization,” in Proc. EMNLP 2015, pp. 379–389.
[9] A. See, P. J. Liu, and C. D. Manning, “Get to the point: Summarization with pointer-generator networks,” in Proc. ACL 2017, pp. 1073–1083.
[10] M. Lewis et al., “BART: Denoising sequence-to-sequence pre-training,” in Proc. ACL 2020, pp. 7871–7880.
[11] J. Zhang et al., “PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization,” in Proc. ICML 2020.
[12] T. B. Brown et al., “Language models are few-shot learners,” Advances in NeurIPS, vol. 33, 2020.
[13] A. Nenkova and K. McKeown, “Automatic Summarization,” Foundations and Trends in Information Retrieval, vol. 5, nos. 2–3, pp. 103–233, 2011.
[14] N. F. Liu et al., “Lost in the middle: How language models use long contexts,” TACL, vol. 12, pp. 157–173, 2023.
Copyright © IJIRCCE 2020.All right reserved