Enhancing multi-document summarization through topic–pattern-based sentence selection

(1) * Shaufiah Shaufiah Mail (Queensland University of Technology, Brisbane, Australia)
(2) Yuefeng Li Mail (Queensland University of Technology, Brisbane, Australia)
(3) Richi Nayak Mail (Queensland University of Technology, Brisbane, Australia)
(4) Yutong Wu Mail (Australian e-Health Research Centre, CSIRO, Brisbane, Australia)
*corresponding author

Abstract


The growing volume of digital text content requires automated summarization systems as fundamental tools that enable users to access information at high speed. Automatic Multi-Document Summarization (MDS) requires systems to produce a summary that combines essential information from multiple documents. The extractive methods, which rely on lexical signals and sentence-based rules, yield only repetitive results because they cannot capture complex thematic relationships. This study developed an improved extractive MDS model that combines topic modeling with pattern-based semantic indicators and a method to choose diverse sentences. The model employs LDA to identify concealed thematic structures, retrieves typical word patterns to improve topic models, and selects topics via a greedy algorithm that reduces redundancy to achieve appropriate salience and coverage. The proposed system achieves better results than classical baselines on the DUC 2006 and DUC 2007 datasets, outperforming Lead, CLASSY04, KL-SUM, LexRank, TextRank, and PETMSUM. The system demonstrates superior performance over all baseline methods, achieving better results on the ROUGE-1, ROUGE-2, and ROUGE-SU4 evaluation metrics. The results show that extractive summarization tasks achieve their best performance when topic–pattern representations are combined with diversity-aware scoring methods.

Keywords


Multi-Document Summarization; Extractive Summarization; Topic Modeling Pattern Mining

   

DOI

https://doi.org/10.26555/ijain.v12i2.2352
      

Article metrics

Abstract views : 166 | PDF views : 3

   

Cite

   

Full Text

Download

References


[1] M. Allahyari et al., “Text Summarization Techniques: A Brief Survey,” Int. J. Adv. Comput. Sci. Appl., vol. 8, no. 10, pp. 397–405, Oct. 2017, doi: 10.14569/IJACSA.2017.081052.

[2] Z. Huang, X. Chen, Y. Wang, J. Huang, and X. Zhao, “A survey on biomedical automatic text summarization with large language models,” Inf. Process. Manag., vol. 62, no. 5, p. 104216, Sep. 2025, doi: 10.1016/j.ipm.2025.104216.

[3] A. Nenkova and K. McKeown, “A Survey of Text Summarization Techniques,” in Mining Text Data, vol. 9781461432, Boston, MA: Springer US, 2012, pp. 43–76. doi: 10.1007/978-1-4614-3223-4_3.

[4] R. Barzilay and K. R. McKeown, “Sentence Fusion for Multidocument News Summarization,” Comput. Linguist., vol. 31, no. 3, pp. 297–328, Sep. 2005, doi: 10.1162/089120105774321091.

[5] X. Zhao, T. Wu, X. Zheng, and R. Li, “Discussions on observer design of nonlinear positive systems via T–S fuzzy modeling,” Neurocomputing, vol. 157, no. June, pp. 70–75, Jun. 2015, doi: 10.1016/j.neucom.2015.01.034.

[6] H. P. Luhn, “The Automatic Creation of Literature Abstracts,” IBM J. Res. Dev., vol. 2, no. 2, pp. 159–165, Apr. 1958, doi: 10.1147/rd.22.0159.

[7] H. P. Edmundson, “New Methods in Automatic Extracting,” J. ACM, vol. 16, no. 2, pp. 264–285, Apr. 1969, doi: 10.1145/321510.321519.

[8] D. R. Radev, E. Hovy, and K. McKeown, “Introduction to the Special Issue on Summarization,” Comput. Linguist., vol. 28, no. 4, pp. 399–408, Dec. 2002, doi: 10.1162/089120102762671927.

[9] M. Gambhir and V. Gupta, “Recent automatic text summarization techniques: a survey,” Artif. Intell. Rev., vol. 47, no. 1, pp. 1–66, Jan. 2017, doi: 10.1007/s10462-016-9475-9.

[10] U. Ihsan, H. Ashraf, and N. Jhanjhi, “Survey on Multi-Document Summarization: Systematic Literature Review.” [Online]. Available: https://arxiv.org/abs/2312.12915.

[11] E. Alanzi and S. Alballaa, “Query-Focused Multi-document Summarization Survey,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 6, pp. 822–833, Jun. 2023, doi: 10.14569/IJACSA.2023.0140688.

[12] A. S. Karnyoto, M. M. Henry, and B. Pardamean, “LDA Topic Modeling for Bioinformatics Terms in arXiv Documents,” Procedia Comput. Sci., vol. 245, pp. 229–238, Jan. 2024, doi: 10.1016/j.procs.2024.10.247.

[13] S. S. Mim, D. Logofatu, G. Guerrero-Contreras, and I. Medina-Bulo, “Leveraging Topic Modeling and Extractive Summarization for Unlocking Insights from NeurIPS Papers,” in 2024 International Conference on INnovations in Intelligent SysTems and Applications (INISTA), IEEE, Sep. 2024, pp. 1–6. doi: 10.1109/INISTA62901.2024.10683818.

[14] N. M. AbdelAziz, A. A. Ali, S. M. Naguib, and L. S. Fayed, “Clustering-based topic modeling for biomedical documents extractive text summarization,” J. Supercomput., vol. 81, no. 1, p. 171, Jan. 2025, doi: 10.1007/s11227-024-06640-6.

[15] R. K. Roul, Navpreet, and S. Nalband, “Unified Scoring and Topic Modeling: A Combined Approach for Superior Multi-document Summarization,” in Lecture Notes in Networks and Systems, vol. 1344 LNNS, Springer Science and Business Media Deutschland GmbH, 2025, pp. 481–493. doi: 10.1007/978-981-96-5958-6_42.

[16] D. Wang, S. Zhu, T. Li, and Y. Gong, “Multi-document summarization using sentence-based topic models,” in Proceedings of the ACL-IJCNLP 2009 Conference Short Papers on - ACL-IJCNLP ’09, Morristown, NJ, USA: Association for Computational Linguistics, 2009, p. 297. doi: 10.3115/1667583.1667675.

[17] Y. Wu, Y. Gao, Y. Li, Y. Xu, and M. Chen, “Mining Topical Relevant Patterns for Multi-document Summarization,” in 2015 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), IEEE, Dec. 2015, pp. 114–117. doi: 10.1109/WI-IAT.2015.136.

[18] D. R. Radev, H. Jing, M. Styś, and D. Tam, “Centroid-based summarization of multiple documents,” Inf. Process. Manag., vol. 40, no. 6, pp. 919–938, Nov. 2004, doi: 10.1016/j.ipm.2003.10.006.

[19] Y. Gong and X. Liu, “Generic text summarization using relevance measure and latent semantic analysis,” in Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval, New York, NY, USA: ACM, Sep. 2001, pp. 19–25. doi: 10.1145/383952.383955.

[20] A. Haghighi and L. Vanderwende, “Exploring content models for multi-document summarization,” in Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics on - NAACL ’09, Morristown, NJ, USA: Association for Computational Linguistics, 2009, p. 362. doi: 10.3115/1620754.1620807.

[21] D. H.-T. Asli Celikyilmaz, “A hybrid hierarchical model for multi-document summarization | Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics,” 2010, pp. 815–824. [Online]. Available: https://aclanthology.org/P10-1084/.

[22] D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet Allocation,” J. Mach. Learn. Res., vol. 3, pp. 993–1022, 2003, [Online]. Available: https://dl.acm.org/doi/10.5555/944919.944937.

[23] J.-P. Qiang, P. Chen, W. Ding, F. Xie, and X. Wu, “Multi-document summarization using closed patterns,” Knowledge-Based Syst., vol. 99, no. March, pp. 28–38, May 2016, doi: 10.1016/j.knosys.2016.01.030.

[24] R. Mihalcea and D. Radev, Graph-based Natural Language Processing and Information Retrieval. Cambridge University Press, 2011. doi: 10.1017/CBO9780511976247.

[25] G. Erkan and D. R. Radev, “LexRank: Graph-based Lexical Centrality as Salience in Text Summarization,” J. Artif. Intell. Res., vol. 22, pp. 457–479, Dec. 2004, doi: 10.1613/jair.1523.

[26] R. Mihalcea, “Unsupervised large-vocabulary word sense disambiguation with graph-based algorithms for sequence data labeling,” in Proceedings of the conference on Human Language Technology and Empirical Methods in Natural Language Processing - HLT ’05, Morristown, NJ, USA: Association for Computational Linguistics, 2005, pp. 411–418. doi: 10.3115/1220575.1220627.

[27] A. P. Widyassari et al., “Review of automatic text summarization techniques & methods,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 4, pp. 1029–1046, Apr. 2022, doi: 10.1016/j.jksuci.2020.05.006.

[28] S. Mulla and N. F. Shaikh, “Over comparative study of text summarization techniques based on graph neural networks,” Web Intell., vol. 22, no. 2, pp. 231–248, Apr. 2024, doi: 10.3233/WEB-230014.

[29] Z. Qilin, W. Yu, X. Jian, Z. Qilin, W. Yu, and X. Jian, “Extractive Document Summarization Model Based on Heterogeneous Graph and Keywords,” J. Univ. Electron. Sci. Technol. China, 2024, Vol. 53, Issue 2, Pages 259-270, vol. 53, no. 2, pp. 259–270, Mar. 2024, [Online]. Available: https://www.juestc.uestc.edu.cn/article/doi/10.12178/1001-0548.2023019.

[30] Z. Jalil, M. Nasir, M. Alazab, J. Nasir, T. Amjad, and A. Alqammaz, “Grapharizer: A Graph-Based Technique for Extractive Multi-Document Summarization,” Electronics, vol. 12, no. 8, p. 1895, Apr. 2023, doi: 10.3390/electronics12081895.

[31] C. Zhou et al., “Exploring Contextual Word-level Style Relevance for Unsupervised Style Transfer,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 7135–7144. doi: 10.18653/v1/2020.acl-main.639.

[32] M. Yasunaga, R. Zhang, K. Meelu, A. Pareek, K. Srinivasan, and D. Radev, “Graph-based Neural Multi-Document Summarization,” in Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), Stroudsburg, PA, USA: Association for Computational Linguistics, 2017, pp. 452–462. doi: 10.18653/v1/K17-1045.

[33] E. Davoodijam, N. Ghadiri, M. Lotfi Shahreza, and F. Rinaldi, “MultiGBS: A multi-layer graph approach to biomedical summarization,” J. Biomed. Inform., vol. 116, p. 103706, Apr. 2021, doi: 10.1016/J.JBI.2021.103706.

[34] V. Giri, M. M. Math, S. Kusal, and S. Patil, “Extractive Text Summarization Employing a Graph-Based Ranking System with Context Transfer,” in Communications in Computer and Information Science, vol. 2434 CCIS, Springer Science and Business Media Deutschland GmbH, 2025, pp. 3–19. doi: 10.1007/978-3-031-84602-1_1.

[35] N. Shabani et al., “A Comprehensive Survey on Graph Summarization With Graph Neural Networks,” IEEE Trans. Artif. Intell., vol. 5, no. 8, pp. 3780–3800, Aug. 2024, doi: 10.1109/TAI.2024.3350545.

[36] D. Wang, P. Liu, Y. Zheng, X. Qiu, and X. Huang, “Heterogeneous Graph Neural Networks for Extractive Document Summarization,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 6209–6219. doi: 10.18653/v1/2020.acl-main.553.

[37] Y. Liu and Z. Gong, “Cycling topic graph learning for neural topic modeling,” Knowledge-Based Syst., vol. 310, no. February, p. 112905, Feb. 2025, doi: 10.1016/j.knosys.2024.112905.

[38] Q. Ye, B. Y. Lin, and X. Ren, “CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 7163–7189. doi: 10.18653/v1/2021.emnlp-main.572.

[39] H. ; Zhao et al., “A Multi-Granularity Heterogeneous Graph for Extractive Text Summarization,” Electron. 2023, Vol. 12, Page 2184, vol. 12, no. 10, p. 2184, May 2023, doi: 10.3390/ELECTRONICS12102184.

[40] T. UÇKAN, “A hybrid model for extractive summarization: Leveraging graph entropy to improve large language model performance,” Ain Shams Eng. J., vol. 16, no. 5, p. 103348, Apr. 2025, doi: 10.1016/j.asej.2025.103348.

[41] S. R. Bowman and G. Dahl, “What Will it Take to Fix Benchmarking in Natural Language Understanding?,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 4843–4855. doi: 10.18653/v1/2021.naacl-main.385.

[42] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.

[43] Y. Liu and M. Lapata, “Text Summarization with Pretrained Encoders,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 3728–3738. doi: 10.18653/v1/D19-1387.

[44] A. See, P. J. Liu, and C. D. Manning, “Get To The Point: Summarization with Pointer-Generator Networks,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Stroudsburg, PA, USA: Association for Computational Linguistics, 2017, pp. 1073–1083. doi: 10.18653/v1/P17-1099.

[45] R. Nallapati, B. Zhou, C. dos Santos, C. Gulcehre, and B. Xiang, “Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond,” in Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, Stroudsburg, PA, USA: Association for Computational Linguistics, 2016, pp. 280–290. doi: 10.18653/v1/K16-1028.

[46] J. Zhang, Y. Zhao, M. Saleh, and P. J. Liu, “PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization,” PMLR, Nov. 2020, pp. 11328–11339. [Online]. Available: https://dl.acm.org/doi/abs/10.5555/3524938.3525989.

[47] H. T. Dang, “Overview of DUC 2006,” in Proceedings of the Document Understanding Conference (DUC 2006), National Institute of Standards and Technology (NIST), 2006, pp. 1–10. [Online]. Available: https://duc.nist.gov/pubs/2006papers/duc2006.

[48] H. T. Dang, “Overview of the DUC 2007 Summarization Task,” in Proceedings of the Document Understanding Conference (DUC 2007), National Institute of Standards and Technology (NIST), 2007, pp. 1–38. [Online]. Available: https://duc.nist.gov/duc2007/proceedings/DUC2007_Summarization_Task.

[49] P. Over, H. Dang, and D. Harman, “DUC in context,” Inf. Process. Manag., vol. 43, no. 6, pp. 1506–1520, Nov. 2007, doi: 10.1016/j.ipm.2007.01.019.

[50] A. McCallum, “MALLET: A Machine Learning for Language Toolkit.” [Online]. Available: http://mallet.cs.umass.edu/.

[51] H. Sun, B. Li, and B. Han, “A novel keyphrase extraction method by combining FP-growth and LDA,” in 2017 13th International Conference on Natural Computation, Fuzzy Systems and Knowledge Discovery (ICNC-FSKD), IEEE, Jul. 2017, pp. 1764–1768. doi: 10.1109/FSKD.2017.8393033.

[52] X. Yu, S. Zhou, and A. Liu, “Smart Tourist Attraction Data Mining Based on FP-Growth Algorithm,” in 2024 International Conference on Internet of Things, Robotics and Distributed Computing (ICIRDC), IEEE, Dec. 2024, pp. 127–132. doi: 10.1109/ICIRDC65564.2024.00029.

[53] N. Ramakrishnan, M. Nair J., D. Jayaprakash, H. Ananthakrishnan, and S. Rani S., “Hypergraph based clustering for document similarity using FP growth algorithm,” in 2019 International Conference on Intelligent Computing and Control Systems (ICCS), IEEE, May 2019, pp. 332–336. doi: 10.1109/ICCS45141.2019.9065630.

[54] Y. Wu, Y. Li, Y. Xu, and W. Huang, “Mining Topically Coherent Patterns for Unsupervised Extractive Multi-document Summarization,” in 2016 IEEE/WIC/ACM International Conference on Web Intelligence (WI), IEEE, Oct. 2016, pp. 129–136. doi: 10.1109/WI.2016.0028.

[55] C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” 2004, pp. 74–81. [Online]. Available: https://aclanthology.org/W04-1013/.

[56] K. Ganesan, “ROUGE 2.0: Updated and Improved Measures for Evaluation of Summarization Tasks,” vol. 1, pp. 1–8, Mar. 2018, [Online]. Available: https://github.com/kavgan/ROUGE-2.0.




Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571  (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
 andri.pranolo.id@ieee.org (publication issues)

View IJAIN Stats

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0