PerMSum: A Context-Dependent Research Label
- PerMSum is a polysemous research label that varies by context, covering personalized summarization, permutation-based testing, and number-theoretic product-sum problems.
- In personalized summarization, PerMSum provides a structured dataset with document sets and user profiles from news and reviews, enabling tailored summary generation.
- The label underpins evaluation frameworks like ComPSum and AuthorMap, which rigorously assess authorship attribution and personalization quality.
PerMSum is a context-dependent research label rather than a single universally fixed term. In the most explicit usage in the supplied literature, it denotes a personalized multi-document summarization dataset spanning reviews and news domains, introduced to support comparative personalization and reference-free evaluation through the ComPSum and AuthorMap frameworks (Li et al., 25 Sep 2025). In other contexts, the same label has been used for a permutation-based true discovery guarantee framework built around sum tests (Vesely et al., 2021), for broader discussion of personalized summarization driven by dynamic user preference histories (Chatterjee et al., 11 Oct 2025), and, in an interpretive mathematical usage, for finite sequences of positive integers whose product equals their sum (Husarov et al., 13 Aug 2025). This suggests that PerMSum should be read strictly in context.
1. Terminological scope
The current literature represented here uses “PerMSum” for several unrelated objects. The term is therefore polysemous, with meaning determined by the surrounding subfield.
| Usage | Description | Source |
|---|---|---|
| PerMSum dataset | Personalized MDS dataset in news and reviews | (Li et al., 25 Sep 2025) |
| PerMSum framework | Permutation-based true discovery guarantee by sum tests | (Vesely et al., 2021) |
| PerMSum as shorthand | Personalized summarization with dynamic user preferences | (Chatterjee et al., 11 Oct 2025) |
| PerMSum as interpretation | Product-sum equality sequences | (Husarov et al., 13 Aug 2025) |
The personalized summarization usage is the most structurally elaborated in the supplied material. There, PerMSum is the substrate for generating personalized summaries from a document set and a user profile , and for evaluating whether personalization is strong enough to support authorship attribution between summaries produced for different users (Li et al., 25 Sep 2025).
A common misconception is to treat PerMSum as an established singular benchmark or method across arXiv. The record here does not support that reading. A more accurate characterization is that PerMSum is a local label reused across different technical settings, most notably in personalized summarization but not confined to it.
2. PerMSum as a personalized multi-document summarization dataset
In personalized multi-document summarization, the input is a document set containing multiple documents on the same topic, together with a user profile containing multiple profile documents authored by user . The output is a personalized summary of that captures the individual preference of as expressed in 0, with personalization defined along the dimensions of writing style and content focus (Li et al., 25 Sep 2025).
PerMSum was constructed to make this problem operational in a setting where both modeling and evaluation require multiple users per topic. The dataset plays two roles. On the model side, it provides document sets 1 and user profiles 2 for systems that generate personalized summaries. On the evaluation side, it provides the author-labeled corpus from which AuthorMap constructs samples involving two users, their profiles, and two summaries produced for the same document set, enabling authorship-based personalization evaluation (Li et al., 25 Sep 2025).
The ComPSum framework uses PerMSum to instantiate comparative personalization. For a user 3, it first generates a structured analysis
4
where 5 are comparative documents on the same topic written by other users. It then generates the personalized summary
6
This formulation makes user specificity depend not only on isolated profile evidence but also on contrastive differences between users who write about the same event or product (Li et al., 25 Sep 2025).
A plausible implication is that PerMSum was designed not merely as a corpus of authored documents, but as a dataset whose topology supports pairwise and comparative user modeling. That design choice differentiates it from ordinary summarization datasets with user labels but no shared-topic author structure.
3. Construction, preprocessing, and corpus structure
PerMSum has two domains: news and reviews. The news portion is constructed from the All The News 2 dataset, while the review portion is constructed from the Amazon book reviews dataset (Li et al., 25 Sep 2025).
In the news domain, PerMSum removes all sentences containing author names or publishing media, because such direct mentions are undesirable shortcuts for personalized summarization and its evaluation. Articles with more than three authors or media organizations are labeled so that they are not used as profile documents. Articles are clustered into document sets using the POLITICS-style event clustering criteria: publication within 2 days, at least one shared named entity in titles or first three sentences, and cosine similarity over TF–IDF embeddings greater than 7. Maximum cliques are then used as event clusters. Clusters with more than three articles written by the same author are filtered out, clusters with more than 10 articles are split, and each news article is truncated to 300 words (Li et al., 25 Sep 2025).
In the review domain, PerMSum uses the Amazon book category, keeps English reviews between 8 and 9 words, and filters out reviews written by users who write more than 0 reviews. Document sets correspond to sets of reviews about a book or product, following prior review-based MDS construction schemes (Li et al., 25 Sep 2025).
Across both domains, profile documents are defined as documents written by the user. PerMSum only considers users that write at least 1 documents, splits users into training, validation, and test sets, and separately splits document sets into training, validation, and test sets. There is no overlap between users or document sets in the three splits, which is intended to prevent information leakage (Li et al., 25 Sep 2025).
The introduction reports that the combined dataset contains 45K document sets and 5.3K users (Li et al., 25 Sep 2025). The split statistics are as follows:
| Domain | Users (train/val/test) | Document sets (train/val/test) |
|---|---|---|
| News | 828 / 293 / 296 | 10730 / 1393 / 1463 |
| Review | 2400 / 763 / 766 | 27725 / 1878 / 1795 |
The same source reports validation/test AuthorMap sample counts of 2085/2360 for news and 2774/2757 for reviews. Average profile size is 2 in news and 3 in reviews. News document sets contain 3–10 documents, review document sets contain 8 documents, and average document lengths are 4 and 5 words respectively (Li et al., 25 Sep 2025).
These statistics indicate that PerMSum is not a small diagnostic benchmark. It is a comparatively large, author-structured MDS resource designed to expose both within-topic author variation and cross-domain variation.
4. Evaluation with ComPSum and AuthorMap
PerMSum is tightly coupled to two methodological components: ComPSum, which performs comparative personalization, and AuthorMap, which evaluates personalization without requiring human-written personalized reference summaries (Li et al., 25 Sep 2025).
AuthorMap evaluates personalization based on authorship attribution between two personalized summaries generated for different users by the same system. For each document set 6, PerMSum selects pairs of users 7 and 8 only among users who wrote documents belonging to 9. This ensures that the two profiles are topically relevant to the same source set. To avoid shortcut copying, documents written by the target users are removed from the input document set before summary generation (Li et al., 25 Sep 2025).
The evaluation sample therefore consists of a shared document set 0, two users 1, their profiles 2, and two personalized summaries 3 generated for the same 4. AuthorMap retrieves 5 profile documents with BM25, truncates them to 100 words, and uses Llama3.3-70b-Instruct or Gemma-3-27b-it as a judge for style and content attribution (Li et al., 25 Sep 2025).
The paper validates AuthorMap on human-written documents from the PerMSum test sets. Reported accuracies are 76.65% for style and 71.64% for content in news, and 89.00% for style and 82.69% for content in reviews (Li et al., 25 Sep 2025). Those figures matter because the entire evaluation is reference-free: the intended signal is not lexical overlap with a gold summary, but whether the generated summaries preserve enough user-specific trace to be attributable to the correct profile.
PerMSum also supports direct model comparison. Using AuthorMap together with factuality and relevance measures, the ComPSum paper reports that ComPSum outperforms strong baselines on PerMSum, while also showing that over-personalization can be pathological: the Rehearsal baseline achieves extremely high AuthorMap scores but poor factuality and relevance because its summaries drift into copying profile content not in the document set (Li et al., 25 Sep 2025). This is an important corrective against interpreting authorship attribution alone as sufficient evidence of good personalized summarization.
5. Relation to adjacent personalized summarization work
A broader personalized summarization literature uses “PerMSum” more loosely to refer to personalized summarization or personalized multi-document summarization as a task family. In that setting, the emphasis shifts from benchmark construction to the dynamics of user-preference modeling. The PerAugy work treats the central bottleneck as the scarcity and low diversity of dynamic preference histories containing both interaction signals and user-specific summaries, and proposes cross-trajectory shuffling plus summary-content perturbation to augment such data (Chatterjee et al., 11 Oct 2025).
PerAugy reports a best user-encoder improvement of 6 with respect to AUC, and an average improvement of 7 with respect to the PSE-SU4 personalization metric for downstream personalized summarizers (Chatterjee et al., 11 Oct 2025). It also introduces three diversity metrics—8, 9, and DegreeD—and reports that 0 and DegreeD strongly correlate with user-encoder performance. This suggests that, within the broader PerMSum problem setting, dataset topology and trajectory diversity can be as consequential as decoder architecture.
PerMSum should also be distinguished from similarly named summarization constructs. PreSumm defines a document-level task of predicting summarization performance from the source document alone, without generating a summary, using the target
1
where 2 is the ACU-based score of system 3 on document 4 (Koniaev et al., 7 Apr 2025). PERCS, by contrast, is a persona-guided biomedical summarization dataset with four personas—Laypersons, Premedical Students, Non-medical Researchers, and Medical Experts—designed for controllable audience-specific summarization rather than author-specific personalized MDS (Salvi et al., 3 Dec 2025).
The distinction is substantive. PerMSum in (Li et al., 25 Sep 2025) is organized around authored profiles and comparative user differences on the same topic; PreSumm estimates summarizability as a document property; PERCS controls biomedical summaries by persona and medical literacy. Similar names therefore mask different problem definitions, supervision structures, and evaluation criteria.
6. Other technical uses of the label
Outside summarization, “PerMSum” has been used for a permutation-based closed testing framework for sum tests in multiple hypothesis testing. In that usage, the method provides simultaneous lower confidence bounds for the proportion of true discoveries over all subsets of hypotheses, adapts to the unknown joint distribution through permutation testing, and is implemented in the R package sumSome (Vesely et al., 2021). Its local tests use sum-based global statistics inside closed testing, and its branch-and-bound shortcut converges to the full closed testing result while remaining valid even if stopped early (Vesely et al., 2021).
The label has also been used interpretively for a mathematical problem in which finite sequences of positive integers satisfy
5
The corresponding paper itself refers to the “product-sum equality” or “equal-sum-product problem,” not to PerMSum, but the supplied description explicitly notes that the keyword “PerMSum” fits this setting naturally (Husarov et al., 13 Aug 2025). That work studies non-increasing sequences of positive integers, identifies the basic solution 6, proves a prefix inequality
7
and gives a recursive algorithm that finds all solutions of length 8, with time complexity stated to be similar to quicksort (Husarov et al., 13 Aug 2025).
These non-summarization uses are not minor typographic variations of the personalized MDS dataset. They are separate technical objects in different fields. The main encyclopedic consequence is terminological: PerMSum is not a field-wide canonical term, but a reusable label whose meaning depends on whether the surrounding discussion is about personalized summarization, permutation-based inference, or additive-multiplicative number-theoretic structure.