Papers
Topics
Authors
Recent
Search
2000 character limit reached

ComPSum: Comparative Personalized Summarization

Updated 12 July 2026
  • ComPSum is a personalized multi-document summarization framework that leverages comparative inference to highlight distinctive writing style and content focus.
  • It uses a two-stage prompting pipeline with retrieval of top 5 user documents and dissimilar comparative documents to build a structured analysis.
  • Empirical results show that ComPSum improves personalization while maintaining factuality and relevance, outperforming several baselines.

ComPSum is a personalized multi-document summarization framework that treats personalization as a comparative inference problem rather than a purely user-isolated one. In its formulation, the input consists of a document set DD containing multiple documents on the same topic and a user profile PuP_u containing documents previously authored by user uu; the output is a personalized summary sus_u that reflects that user’s writing style and content focus. Its central premise is that fine-grained personalization is easier to identify by comparing a target user’s preferences with other users’ preferences on the same topic, and it operationalizes that premise through a two-stage prompting pipeline with an intermediate structured analysis of user preferences (Li et al., 25 Sep 2025).

1. Conceptual basis and problem formulation

ComPSum is defined in the setting of personalized multi-document summarization (MDS). The framework assumes that a generic summary of a document set is often insufficient when different users prefer different emphases or different modes of expression. The paper therefore frames personalization along two explicit dimensions: writing style and content focus. Writing style concerns how the summary is written, whereas content focus concerns which aspects of the source set are emphasized.

The framework’s key methodological claim is that modeling a user only from their own profile can miss what is distinctive about that user. By contrast, comparing the user’s documents against comparable same-topic documents written by other users can make user-specific preferences more explicit. This comparative perspective is particularly natural in MDS because documents within the same document set are already about the same topic. This suggests that differences across authors in such sets are more plausibly attributable to preference differences than to topic mismatch.

The paper formalizes the task with the following objects. The document set DD is the source to be summarized, PuP_u is the profile of user uu, and sus_u is the personalized summary for that user. Internally, ComPSum also constructs a structured intermediate representation aua_u, described as a structured analysis of the user’s distinctive writing style and content focus. The paper does not define a learned objective or training loss for this framework, because ComPSum is used as a prompted inference-time system rather than a fine-tuned model.

2. Comparative personalization framework

ComPSum is a two-stage prompting pipeline. Its first stage builds a user representation by retrieval and comparison; its second stage uses that representation to guide personalized summary generation. The intermediate structure is central: the framework does not merely retrieve user documents and pass them directly to a generator, but instead converts comparative evidence into a short, explicit analysis organized into Style analysis and Content analysis.

The first retrieval step selects the top kk profile documents from PuP_u0 that are most relevant to the current summarization input PuP_u1. The paper denotes this retrieval with a retrieval model PuP_u2 applied to the concatenation of all documents in PuP_u3. In the reported experiments, retrieval is performed with BM25, and the number of retrieved profile documents is set to 5.

For each retrieved user profile document PuP_u4, ComPSum then identifies comparative evidence from other users. Let PuP_u5 denote the documents that belong to the same document set as PuP_u6 but are authored by users other than PuP_u7. From that set, the framework retrieves one comparative document PuP_u8 that is most dissimilar to PuP_u9, again using the retrieval model uu0. This step is the framework’s explicit comparative mechanism: it contrasts a user’s own same-topic writing with a sharply different alternative.

Using these paired documents, ComPSum prompts an LLM to generate the structured analysis: uu1 where uu2 characterizes the user’s distinctive writing style and content focus. The paper emphasizes that this is not a generic profile summary. It is a constrained comparative description, and examples in the paper include phrases such as “Unlike other users…”. This suggests that the framework is designed to encode difference-aware user traits rather than merely recurring user topics.

The final personalized summary is then generated from the retrieved profile documents, the structured analysis, and the source document set: uu3 The summarization prompt instructs the model to mimic the user’s writing style and content focus while ensuring that the summary contains only content supported by uu4. This separation between user modeling and source grounding is a defining property of the framework.

3. Modeling choices and algorithmic structure

ComPSum is prompt-based and LLM-centered rather than a specialized trainable architecture. The retrieval model used in experiments is BM25. The generator LLMs evaluated for the framework are Llama3.1-8b-Instruct, Qwen2.5-14B-Instruct, and Llama3.3-70b-Instruct. For evaluation in the AuthorMap framework, the primary judge model is Llama3.3-70b-Instruct, with additional testing using Gemma-3-27b-it.

The schema of the structured analysis is deliberately minimal. It contains only two dimensions: Writing style and Content focus. The paper does not introduce a richer ontology, latent variable model, or aspect taxonomy for user preference. This suggests that the framework favors prompt interpretability and controllability over a more elaborate internal representation.

At inference time, the behavior of ComPSum is dynamic and instance-specific. For a given pair uu5, it retrieves topic-relevant profile documents, retrieves comparative same-topic documents from other users, generates a comparative structured analysis, and only then generates the personalized summary. The system therefore does not rely on a fixed per-user embedding or parameter vector.

Several implementation details are specified. The framework retrieves 5 profile documents for personalization and comparison, truncates retrieved profile documents to 100 words, limits output personalized summaries to 100 words, uses default sampling parameters for the LLMs, and tunes prompts and hyperparameters on the validation set. An ablation over the number of retrieved profile documents reports that uu6 performs best, that uu7 is somewhat worse, and that uu8 hurts performance. A plausible implication is that overly large retrieved contexts dilute user-specific signals rather than strengthening them.

4. Evaluation framework and dataset infrastructure

Because personalized MDS lacks gold personalized references, the paper introduces AuthorMap, a reference-free evaluation framework. AuthorMap evaluates personalization through authorship attribution between two personalized summaries generated by the same system for two different users on the same input document set. It measures personalization separately for writing style and content focus.

The evaluation protocol is defined as follows. Suppose uu9 and sus_u0 are two user profiles, and sus_u1 and sus_u2 are two summaries generated by the same system for the same document set sus_u3. For each profile sus_u4, AuthorMap retrieves the top sus_u5 profile documents most similar to the concatenation of the two summaries: sus_u6 where sus_u7 denotes concatenation and sus_u8 in experiments. The paper notes that concatenating both summaries avoids giving either summary an inherent retrieval advantage.

The judge then predicts which user more likely authored each retrieved profile, separately for style and content: sus_u9

DD0

To mitigate position bias, predictions are made twice with the summary order swapped, yielding four predictions total. The final score is the percentage of samples where the judge correctly predicts the author of the retrieved profile in the majority of the four predictions.

The paper also introduces PerMSum, a personalized MDS dataset spanning news and review domains. In the news domain, the source is All The News; in the review domain, the source is the Amazon dataset, using the book category. The introduction summarizes PerMSum as containing about 45K document sets and 5.3K users. The split statistics are reported more precisely in the dataset table. For News, the counts are 828 / 293 / 296 users, 10730 / 1393 / 1463 document sets, 2085 validation evaluation samples, 2360 test evaluation samples, average profile size 39.72, document set size 3–10, and average document length 216.72 words. For Review, the counts are 2400 / 763 / 766 users, 27725 / 1878 / 1795 document sets, 2774 validation evaluation samples, 2757 test evaluation samples, average profile size 19.77, document set size 8, and average document length 86.40 words (Li et al., 25 Sep 2025).

The dataset construction is designed to avoid leakage. Only users with at least 10 documents are kept; users and document sets do not overlap across train, validation, and test; and when generating a personalized summary for a user DD1 on a document set DD2, all documents written by DD3 that are inside DD4 are removed from the input. Evaluation samples are restricted to user pairs who both wrote documents belonging to the same document set, which ensures that their profiles are likely informative for the topic at hand.

5. Empirical results and ablation findings

The main baselines are RAG, CICL, RAG+Summary, DPL, and Rehearsal. All baselines use 5 retrieved profile documents for fair comparison. Evaluation combines personalization and summary quality through AuthorMap style and content scores, FactScore for factuality, G-Eval for relevance, and an overall score defined as the arithmetic average of style, content, factuality, and relevance.

Across all three tested LLM backbones, ComPSum attains the strongest reported overall scores. With Llama3.1-8b-Instruct, the overall scores are 74.07 on News and 74.54 on Review for ComPSum, compared with second-best scores of 71.72 and 72.74 for RAG+Summary. With Qwen2.5-14B-Instruct, ComPSum reaches 74.78 on News and 76.08 on Review, again above all baselines. With Llama3.3-70b-Instruct, ComPSum achieves 69.74 on News and 72.95 on Review, remaining best in both domains. The paper states that differences between ComPSum and the second-best methods are statistically significant with paired bootstrap resampling at DD5.

The paper emphasizes that ComPSum improves personalization without sacrificing factuality and relevance. This is especially visible in comparison with Rehearsal, which obtains extremely high personalization scores but very low factuality and relevance. The authors report that these outputs often resemble the user’s profile rather than faithful summaries of the source set. This is presented as evidence that personalization and source grounding must be balanced rather than optimized independently.

Ablation studies isolate the effect of the framework’s main components. The average overall performance is 74.87 for full ComPSum, 73.19 for w/o comp. doc., 73.77 for w/o structure, 74.00 for w/ sim. comp., and 72.02 for w/ multi. stage. These results support four specific claims: comparative documents help; structured analysis helps; most dissimilar comparative documents are more effective than similar ones; and generating one user-level analysis directly from multiple profile documents is better than a multi-stage document-wise aggregation scheme.

The paper also analyzes whether comparative documents make the structured analyses more user-distinctive. It measures embedding cosine similarity between analyses produced for two different users on the same input set. The average similarity is 80.89 for ComPSum and 82.52 for w/o comp. doc.. Lower similarity under ComPSum is interpreted as evidence that comparative personalization yields more differentiated user analyses.

6. Limitations, interpretation, and place in the literature

The paper identifies one especially important limitation: ComPSum depends on the availability of comparable same-topic documents written by different users. This dependency is natural in MDS, but it makes the framework harder to transfer to arbitrary personalized generation settings where such topic-controlled comparative documents may not exist. In such settings, differences between texts may reflect topic variation rather than user preference.

The construction of user profiles in the news domain is another constraint. In PerMSum, news profiles are defined by documents authored by the user. The paper notes that real-world news personalization might be better modeled by clicked or liked articles, but suitable public datasets with full text and time metadata were unavailable. This suggests that the current evaluation setting captures a specific, authorship-based notion of personalization rather than all practically relevant forms of personalization.

AuthorMap also has a narrower target than human preference modeling. It measures whether personalization is attributable through style and content-focus signatures, not whether end users would necessarily prefer the generated summaries. The paper validates AuthorMap in several ways: on human-written documents truncated to 100 words, it reports News accuracies of 76.65 for style and 71.64 for content, and Review accuracies of 89.00 for style and 82.69 for content; human evaluation reports Randolph’s kappa = 0.40 and agreement of AuthorMap with human choice of 80% style and 73% content on News, and 73% style and 80% content on Review. These results support the metric’s usefulness, but they do not make it equivalent to direct user satisfaction.

Within the broader summarization literature, ComPSum belongs to personalized multi-document summarization rather than generic MDS or source summarization. Its closest conceptual contribution is the idea of comparative personalization: personalization quality improves when user modeling explicitly identifies how one user differs from other users writing about the same topic. The paper’s evidence supports that claim empirically. A plausible interpretation is that ComPSum’s main novelty lies less in the choice of base LLM than in the structuring of personalization as retrieval, contrast, structured analysis, and grounded generation.

In that sense, ComPSum contributes a general design principle for personalized summarization: user history is most informative when transformed into contrastive, dimension-specific control signals. The framework’s strongest empirical result is that this comparative representation improves both style personalization and content-focus personalization while maintaining factuality and relevance on PerMSum (Li et al., 25 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ComPSum.