Publication

Benchmarking LLMs on semantic overlap summarization

Salvador, John
Bansal, Naman
Akter, Mousumi
Sarkar, Souvika
Das, Anupam
Karmaker, Santu
Citations
Altmetric:
Other Names
Location
Time Period
Advisors
Original Date
Digitization Date
Issue Date
2025-11
Type
Article
Genre
Keywords
Benchmarking,LLMs
Subjects (LCSH)
Research Projects
Organizational Units
Journal Issue
Citation
John Salvador, Naman Bansal, Mousumi Akter, Souvika Sarkar, Anupam Das, and Santu Karmaker. 2025. Benchmarking LLMs on Semantic Overlap Summarization. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 33352–33373, Suzhou, China. Association for Computational Linguistics.
Abstract
Semantic Overlap Summarization (SOS) is a multi-document summarization task focused on extracting the common information shared cross alternative narratives which is a capability that is critical for trustworthy generation in domains such as news, law, and healthcare. We benchmark popular Large Language Models (LLMs) on SOS and introduce PrivacyPolicyPairs (3P), a new dataset of 135 high-quality samples from privacy policy documents, which complements existing resources and broadens domain coverage. Using the TELeR prompting taxonomy, we evaluate nearly one million LLM-generated summaries across two SOS datasets and conduct human evaluation on a curated subset. Our analysis reveals strong prompt sensitivity, identifies which automatic metrics align most closely with human judgments, and provides new baselines for future SOS research
Table of Contents
Description
Click on the DOI link to access this article at the publishers website (may not be free).
Publisher
Association for Computational Linguistics
Journal
Book Title
Series
Digital Collection
Finding Aid URL
Use and Reproduction
Archival Collection
PubMed ID
DOI
ISSN
EISSN
Embedded videos