ComPSum: Comparative Personalized Summarization
- ComPSum is a personalized multi-document summarization framework that leverages comparative inference to highlight distinctive writing style and content focus.
- It uses a two-stage prompting pipeline with retrieval of top 5 user documents and dissimilar comparative documents to build a structured analysis.
- Empirical results show that ComPSum improves personalization while maintaining factuality and relevance, outperforming several baselines.
ComPSum is a personalized multi-document summarization framework that treats personalization as a comparative inference problem rather than a purely user-isolated one. In its formulation, the input consists of a document set containing multiple documents on the same topic and a user profile containing documents previously authored by user ; the output is a personalized summary that reflects that user’s writing style and content focus. Its central premise is that fine-grained personalization is easier to identify by comparing a target user’s preferences with other users’ preferences on the same topic, and it operationalizes that premise through a two-stage prompting pipeline with an intermediate structured analysis of user preferences (Li et al., 25 Sep 2025).
1. Conceptual basis and problem formulation
ComPSum is defined in the setting of personalized multi-document summarization (MDS). The framework assumes that a generic summary of a document set is often insufficient when different users prefer different emphases or different modes of expression. The paper therefore frames personalization along two explicit dimensions: writing style and content focus. Writing style concerns how the summary is written, whereas content focus concerns which aspects of the source set are emphasized.
The framework’s key methodological claim is that modeling a user only from their own profile can miss what is distinctive about that user. By contrast, comparing the user’s documents against comparable same-topic documents written by other users can make user-specific preferences more explicit. This comparative perspective is particularly natural in MDS because documents within the same document set are already about the same topic. This suggests that differences across authors in such sets are more plausibly attributable to preference differences than to topic mismatch.
The paper formalizes the task with the following objects. The document set is the source to be summarized, is the profile of user , and is the personalized summary for that user. Internally, ComPSum also constructs a structured intermediate representation , described as a structured analysis of the user’s distinctive writing style and content focus. The paper does not define a learned objective or training loss for this framework, because ComPSum is used as a prompted inference-time system rather than a fine-tuned model.
2. Comparative personalization framework
ComPSum is a two-stage prompting pipeline. Its first stage builds a user representation by retrieval and comparison; its second stage uses that representation to guide personalized summary generation. The intermediate structure is central: the framework does not merely retrieve user documents and pass them directly to a generator, but instead converts comparative evidence into a short, explicit analysis organized into Style analysis and Content analysis.
The first retrieval step selects the top profile documents from 0 that are most relevant to the current summarization input 1. The paper denotes this retrieval with a retrieval model 2 applied to the concatenation of all documents in 3. In the reported experiments, retrieval is performed with BM25, and the number of retrieved profile documents is set to 5.
For each retrieved user profile document 4, ComPSum then identifies comparative evidence from other users. Let 5 denote the documents that belong to the same document set as 6 but are authored by users other than 7. From that set, the framework retrieves one comparative document 8 that is most dissimilar to 9, again using the retrieval model 0. This step is the framework’s explicit comparative mechanism: it contrasts a user’s own same-topic writing with a sharply different alternative.
Using these paired documents, ComPSum prompts an LLM to generate the structured analysis: 1 where 2 characterizes the user’s distinctive writing style and content focus. The paper emphasizes that this is not a generic profile summary. It is a constrained comparative description, and examples in the paper include phrases such as “Unlike other users…”. This suggests that the framework is designed to encode difference-aware user traits rather than merely recurring user topics.
The final personalized summary is then generated from the retrieved profile documents, the structured analysis, and the source document set: 3 The summarization prompt instructs the model to mimic the user’s writing style and content focus while ensuring that the summary contains only content supported by 4. This separation between user modeling and source grounding is a defining property of the framework.
3. Modeling choices and algorithmic structure
ComPSum is prompt-based and LLM-centered rather than a specialized trainable architecture. The retrieval model used in experiments is BM25. The generator LLMs evaluated for the framework are Llama3.1-8b-Instruct, Qwen2.5-14B-Instruct, and Llama3.3-70b-Instruct. For evaluation in the AuthorMap framework, the primary judge model is Llama3.3-70b-Instruct, with additional testing using Gemma-3-27b-it.
The schema of the structured analysis is deliberately minimal. It contains only two dimensions: Writing style and Content focus. The paper does not introduce a richer ontology, latent variable model, or aspect taxonomy for user preference. This suggests that the framework favors prompt interpretability and controllability over a more elaborate internal representation.
At inference time, the behavior of ComPSum is dynamic and instance-specific. For a given pair 5, it retrieves topic-relevant profile documents, retrieves comparative same-topic documents from other users, generates a comparative structured analysis, and only then generates the personalized summary. The system therefore does not rely on a fixed per-user embedding or parameter vector.
Several implementation details are specified. The framework retrieves 5 profile documents for personalization and comparison, truncates retrieved profile documents to 100 words, limits output personalized summaries to 100 words, uses default sampling parameters for the LLMs, and tunes prompts and hyperparameters on the validation set. An ablation over the number of retrieved profile documents reports that 6 performs best, that 7 is somewhat worse, and that 8 hurts performance. A plausible implication is that overly large retrieved contexts dilute user-specific signals rather than strengthening them.
4. Evaluation framework and dataset infrastructure
Because personalized MDS lacks gold personalized references, the paper introduces AuthorMap, a reference-free evaluation framework. AuthorMap evaluates personalization through authorship attribution between two personalized summaries generated by the same system for two different users on the same input document set. It measures personalization separately for writing style and content focus.
The evaluation protocol is defined as follows. Suppose 9 and 0 are two user profiles, and 1 and 2 are two summaries generated by the same system for the same document set 3. For each profile 4, AuthorMap retrieves the top 5 profile documents most similar to the concatenation of the two summaries: 6 where 7 denotes concatenation and 8 in experiments. The paper notes that concatenating both summaries avoids giving either summary an inherent retrieval advantage.
The judge then predicts which user more likely authored each retrieved profile, separately for style and content: 9
0
To mitigate position bias, predictions are made twice with the summary order swapped, yielding four predictions total. The final score is the percentage of samples where the judge correctly predicts the author of the retrieved profile in the majority of the four predictions.
The paper also introduces PerMSum, a personalized MDS dataset spanning news and review domains. In the news domain, the source is All The News; in the review domain, the source is the Amazon dataset, using the book category. The introduction summarizes PerMSum as containing about 45K document sets and 5.3K users. The split statistics are reported more precisely in the dataset table. For News, the counts are 828 / 293 / 296 users, 10730 / 1393 / 1463 document sets, 2085 validation evaluation samples, 2360 test evaluation samples, average profile size 39.72, document set size 3–10, and average document length 216.72 words. For Review, the counts are 2400 / 763 / 766 users, 27725 / 1878 / 1795 document sets, 2774 validation evaluation samples, 2757 test evaluation samples, average profile size 19.77, document set size 8, and average document length 86.40 words (Li et al., 25 Sep 2025).
The dataset construction is designed to avoid leakage. Only users with at least 10 documents are kept; users and document sets do not overlap across train, validation, and test; and when generating a personalized summary for a user 1 on a document set 2, all documents written by 3 that are inside 4 are removed from the input. Evaluation samples are restricted to user pairs who both wrote documents belonging to the same document set, which ensures that their profiles are likely informative for the topic at hand.
5. Empirical results and ablation findings
The main baselines are RAG, CICL, RAG+Summary, DPL, and Rehearsal. All baselines use 5 retrieved profile documents for fair comparison. Evaluation combines personalization and summary quality through AuthorMap style and content scores, FactScore for factuality, G-Eval for relevance, and an overall score defined as the arithmetic average of style, content, factuality, and relevance.
Across all three tested LLM backbones, ComPSum attains the strongest reported overall scores. With Llama3.1-8b-Instruct, the overall scores are 74.07 on News and 74.54 on Review for ComPSum, compared with second-best scores of 71.72 and 72.74 for RAG+Summary. With Qwen2.5-14B-Instruct, ComPSum reaches 74.78 on News and 76.08 on Review, again above all baselines. With Llama3.3-70b-Instruct, ComPSum achieves 69.74 on News and 72.95 on Review, remaining best in both domains. The paper states that differences between ComPSum and the second-best methods are statistically significant with paired bootstrap resampling at 5.
The paper emphasizes that ComPSum improves personalization without sacrificing factuality and relevance. This is especially visible in comparison with Rehearsal, which obtains extremely high personalization scores but very low factuality and relevance. The authors report that these outputs often resemble the user’s profile rather than faithful summaries of the source set. This is presented as evidence that personalization and source grounding must be balanced rather than optimized independently.
Ablation studies isolate the effect of the framework’s main components. The average overall performance is 74.87 for full ComPSum, 73.19 for w/o comp. doc., 73.77 for w/o structure, 74.00 for w/ sim. comp., and 72.02 for w/ multi. stage. These results support four specific claims: comparative documents help; structured analysis helps; most dissimilar comparative documents are more effective than similar ones; and generating one user-level analysis directly from multiple profile documents is better than a multi-stage document-wise aggregation scheme.
The paper also analyzes whether comparative documents make the structured analyses more user-distinctive. It measures embedding cosine similarity between analyses produced for two different users on the same input set. The average similarity is 80.89 for ComPSum and 82.52 for w/o comp. doc.. Lower similarity under ComPSum is interpreted as evidence that comparative personalization yields more differentiated user analyses.
6. Limitations, interpretation, and place in the literature
The paper identifies one especially important limitation: ComPSum depends on the availability of comparable same-topic documents written by different users. This dependency is natural in MDS, but it makes the framework harder to transfer to arbitrary personalized generation settings where such topic-controlled comparative documents may not exist. In such settings, differences between texts may reflect topic variation rather than user preference.
The construction of user profiles in the news domain is another constraint. In PerMSum, news profiles are defined by documents authored by the user. The paper notes that real-world news personalization might be better modeled by clicked or liked articles, but suitable public datasets with full text and time metadata were unavailable. This suggests that the current evaluation setting captures a specific, authorship-based notion of personalization rather than all practically relevant forms of personalization.
AuthorMap also has a narrower target than human preference modeling. It measures whether personalization is attributable through style and content-focus signatures, not whether end users would necessarily prefer the generated summaries. The paper validates AuthorMap in several ways: on human-written documents truncated to 100 words, it reports News accuracies of 76.65 for style and 71.64 for content, and Review accuracies of 89.00 for style and 82.69 for content; human evaluation reports Randolph’s kappa = 0.40 and agreement of AuthorMap with human choice of 80% style and 73% content on News, and 73% style and 80% content on Review. These results support the metric’s usefulness, but they do not make it equivalent to direct user satisfaction.
Within the broader summarization literature, ComPSum belongs to personalized multi-document summarization rather than generic MDS or source summarization. Its closest conceptual contribution is the idea of comparative personalization: personalization quality improves when user modeling explicitly identifies how one user differs from other users writing about the same topic. The paper’s evidence supports that claim empirically. A plausible interpretation is that ComPSum’s main novelty lies less in the choice of base LLM than in the structuring of personalization as retrieval, contrast, structured analysis, and grounded generation.
In that sense, ComPSum contributes a general design principle for personalized summarization: user history is most informative when transformed into contrastive, dimension-specific control signals. The framework’s strongest empirical result is that this comparative representation improves both style personalization and content-focus personalization while maintaining factuality and relevance on PerMSum (Li et al., 25 Sep 2025).