---
title: 'ComPSum: Comparative Personalized Summarization'
url: https://www.emergentmind.com/topics/compsum
type: topic
---

# ComPSum: Comparative Personalized Summarization

ComPSum is a personalized multi-document summarization framework that treats personalization as a comparative inference problem rather than a purely user-isolated one. In its formulation, the input consists of a document set \(D\) containing multiple documents on the same topic and a user profile \(P_u\) containing documents previously authored by user \(u\); the output is a personalized summary \(s_u\) that reflects that user’s writing style and content focus. Its central premise is that fine-grained personalization is easier to identify by comparing a target user’s preferences with other users’ preferences on the same topic, and it operationalizes that premise through a two-stage prompting pipeline with an intermediate structured analysis of user preferences [2509.21562].

## 1. Conceptual basis and problem formulation

ComPSum is defined in the setting of personalized multi-document summarization (MDS). The framework assumes that a generic summary of a document set is often insufficient when different users prefer different emphases or different modes of expression. The paper therefore frames personalization along two explicit dimensions: **writing style** and **content focus**. Writing style concerns how the summary is written, whereas content focus concerns which aspects of the source set are emphasized.

The framework’s key methodological claim is that modeling a user only from their own profile can miss what is distinctive about that user. By contrast, comparing the user’s documents against comparable same-topic documents written by other users can make user-specific preferences more explicit. This comparative perspective is particularly natural in MDS because documents within the same document set are already about the same topic. This suggests that differences across authors in such sets are more plausibly attributable to preference differences than to topic mismatch.

The paper formalizes the task with the following objects. The document set \(D\) is the source to be summarized, \(P_u\) is the profile of user \(u\), and \(s_u\) is the personalized summary for that user. Internally, ComPSum also constructs a structured intermediate representation \(a_u\), described as a **structured analysis** of the user’s distinctive writing style and content focus. The paper does not define a learned objective or training loss for this framework, because ComPSum is used as a prompted inference-time system rather than a fine-tuned model.

## 2. Comparative personalization framework

ComPSum is a two-stage prompting pipeline. Its first stage builds a user representation by retrieval and comparison; its second stage uses that representation to guide personalized summary generation. The intermediate structure is central: the framework does not merely retrieve user documents and pass them directly to a generator, but instead converts comparative evidence into a short, explicit analysis organized into **Style analysis** and **Content analysis**.

The first retrieval step selects the top \(k\) profile documents from \(P_u\) that are most relevant to the current summarization input \(D\). The paper denotes this retrieval with a retrieval model \(\mathcal{R}\) applied to the concatenation of all documents in \(D\). In the reported experiments, retrieval is performed with **BM25**, and the number of retrieved profile documents is set to **5**.

For each retrieved user profile document \(p_u^i\), ComPSum then identifies comparative evidence from other users. Let \(C_{\neg u}^{p_i}\) denote the documents that belong to the same document set as \(p_u^i\) but are authored by users other than \(u\). From that set, the framework retrieves one comparative document \(p_{\neg u}^i\) that is **most dissimilar** to \(p_u^i\), again using the retrieval model \(\mathcal{R}\). This step is the framework’s explicit comparative mechanism: it contrasts a user’s own same-topic writing with a sharply different alternative.

Using these paired documents, ComPSum prompts an LLM to generate the structured analysis:
\[
a_u = LLM(p_u^1,p_{\neg u}^1,\ldots,p_u^k,p_{\neg u}^k)
\]
where \(a_u\) characterizes the user’s distinctive writing style and content focus. The paper emphasizes that this is not a generic profile summary. It is a constrained comparative description, and examples in the paper include phrases such as “Unlike other users…”. This suggests that the framework is designed to encode difference-aware user traits rather than merely recurring user topics.

The final personalized summary is then generated from the retrieved profile documents, the structured analysis, and the source document set:
\[
s_u = LLM(p_u^1,\ldots,p_u^k,a_u,D)
\]
The summarization prompt instructs the model to mimic the user’s writing style and content focus while ensuring that the summary contains only content supported by \(D\). This separation between user modeling and source grounding is a defining property of the framework.

## 3. Modeling choices and algorithmic structure

ComPSum is prompt-based and LLM-centered rather than a specialized trainable architecture. The retrieval model used in experiments is **BM25**. The generator LLMs evaluated for the framework are **Llama3.1-8b-Instruct**, **Qwen2.5-14B-Instruct**, and **Llama3.3-70b-Instruct**. For evaluation in the AuthorMap framework, the primary judge model is **Llama3.3-70b-Instruct**, with additional testing using **Gemma-3-27b-it**.

The schema of the structured analysis is deliberately minimal. It contains only two dimensions: **Writing style** and **Content focus**. The paper does not introduce a richer ontology, latent variable model, or aspect taxonomy for user preference. This suggests that the framework favors prompt interpretability and controllability over a more elaborate internal representation.

At inference time, the behavior of ComPSum is dynamic and instance-specific. For a given pair \((D,P_u)\), it retrieves topic-relevant profile documents, retrieves comparative same-topic documents from other users, generates a comparative structured analysis, and only then generates the personalized summary. The system therefore does not rely on a fixed per-user embedding or parameter vector.

Several implementation details are specified. The framework retrieves **5 profile documents** for personalization and comparison, truncates retrieved profile documents to **100 words**, limits output personalized summaries to **100 words**, uses default sampling parameters for the LLMs, and tunes prompts and hyperparameters on the validation set. An ablation over the number of retrieved profile documents reports that \(m=5\) performs best, that \(m=2\) is somewhat worse, and that \(m=10\) hurts performance. A plausible implication is that overly large retrieved contexts dilute user-specific signals rather than strengthening them.

## 4. Evaluation framework and dataset infrastructure

Because personalized MDS lacks gold personalized references, the paper introduces **AuthorMap**, a reference-free evaluation framework. AuthorMap evaluates personalization through authorship attribution between two personalized summaries generated by the same system for two different users on the same input document set. It measures personalization separately for **writing style** and **content focus**.

The evaluation protocol is defined as follows. Suppose \(P_{u_1}\) and \(P_{u_2}\) are two user profiles, and \(s_{u_1}\) and \(s_{u_2}\) are two summaries generated by the same system for the same document set \(D\). For each profile \(P_{u_*}\), AuthorMap retrieves the top \(n\) profile documents most similar to the concatenation of the two summaries:
\[
\mathcal{R}(s_{u_1} \circ s_{u_2}, P_{u_*}, n)
\]
where \(\circ\) denotes concatenation and \(n=5\) in experiments. The paper notes that concatenating both summaries avoids giving either summary an inherent retrieval advantage.

The judge then predicts which user more likely authored each retrieved profile, separately for style and content:
\[
\hat{u}_1 = LLM_{judge}(P_{u_1}, s_{u_1}, s_{u_2})
\]
\[
\hat{u}_2 = LLM_{judge}(P_{u_2}, s_{u_1}, s_{u_2})
\]
To mitigate position bias, predictions are made twice with the summary order swapped, yielding four predictions total. The final score is the percentage of samples where the judge correctly predicts the author of the retrieved profile in the majority of the four predictions.

The paper also introduces **PerMSum**, a personalized MDS dataset spanning **news** and **review** domains. In the news domain, the source is **All The News**; in the review domain, the source is the **Amazon** dataset, using the **book** category. The introduction summarizes PerMSum as containing about **45K document sets** and **5.3K users**. The split statistics are reported more precisely in the dataset table. For **News**, the counts are **828 / 293 / 296** users, **10730 / 1393 / 1463** document sets, **2085** validation evaluation samples, **2360** test evaluation samples, average profile size **39.72**, document set size **3–10**, and average document length **216.72** words. For **Review**, the counts are **2400 / 763 / 766** users, **27725 / 1878 / 1795** document sets, **2774** validation evaluation samples, **2757** test evaluation samples, average profile size **19.77**, document set size **8**, and average document length **86.40** words [2509.21562].

The dataset construction is designed to avoid leakage. Only users with at least **10 documents** are kept; users and document sets do not overlap across train, validation, and test; and when generating a personalized summary for a user \(u\) on a document set \(D\), all documents written by \(u\) that are inside \(D\) are removed from the input. Evaluation samples are restricted to user pairs who both wrote documents belonging to the same document set, which ensures that their profiles are likely informative for the topic at hand.

## 5. Empirical results and ablation findings

The main baselines are **RAG**, **CICL**, **RAG+Summary**, **DPL**, and **Rehearsal**. All baselines use **5 retrieved profile documents** for fair comparison. Evaluation combines personalization and summary quality through **AuthorMap** style and content scores, **FactScore** for factuality, **G-Eval** for relevance, and an **overall score** defined as the arithmetic average of style, content, factuality, and relevance.

Across all three tested LLM backbones, ComPSum attains the strongest reported overall scores. With **Llama3.1-8b-Instruct**, the overall scores are **74.07** on News and **74.54** on Review for ComPSum, compared with second-best scores of **71.72** and **72.74** for **RAG+Summary**. With **Qwen2.5-14B-Instruct**, ComPSum reaches **74.78** on News and **76.08** on Review, again above all baselines. With **Llama3.3-70b-Instruct**, ComPSum achieves **69.74** on News and **72.95** on Review, remaining best in both domains. The paper states that differences between ComPSum and the second-best methods are statistically significant with paired bootstrap resampling at \(p < 0.05\).

The paper emphasizes that ComPSum improves personalization without sacrificing factuality and relevance. This is especially visible in comparison with **Rehearsal**, which obtains extremely high personalization scores but very low factuality and relevance. The authors report that these outputs often resemble the user’s profile rather than faithful summaries of the source set. This is presented as evidence that personalization and source grounding must be balanced rather than optimized independently.

Ablation studies isolate the effect of the framework’s main components. The average overall performance is **74.87** for full ComPSum, **73.19** for **w/o comp. doc.**, **73.77** for **w/o structure**, **74.00** for **w/ sim. comp.**, and **72.02** for **w/ multi. stage**. These results support four specific claims: comparative documents help; structured analysis helps; **most dissimilar** comparative documents are more effective than similar ones; and generating one user-level analysis directly from multiple profile documents is better than a multi-stage document-wise aggregation scheme.

The paper also analyzes whether comparative documents make the structured analyses more user-distinctive. It measures embedding cosine similarity between analyses produced for two different users on the same input set. The average similarity is **80.89** for ComPSum and **82.52** for **w/o comp. doc.**. Lower similarity under ComPSum is interpreted as evidence that comparative personalization yields more differentiated user analyses.

## 6. Limitations, interpretation, and place in the literature

The paper identifies one especially important limitation: ComPSum depends on the availability of comparable same-topic documents written by different users. This dependency is natural in MDS, but it makes the framework harder to transfer to arbitrary personalized generation settings where such topic-controlled comparative documents may not exist. In such settings, differences between texts may reflect topic variation rather than user preference.

The construction of user profiles in the news domain is another constraint. In PerMSum, news profiles are defined by documents authored by the user. The paper notes that real-world news personalization might be better modeled by clicked or liked articles, but suitable public datasets with full text and time metadata were unavailable. This suggests that the current evaluation setting captures a specific, authorship-based notion of personalization rather than all practically relevant forms of personalization.

AuthorMap also has a narrower target than human preference modeling. It measures whether personalization is attributable through style and content-focus signatures, not whether end users would necessarily prefer the generated summaries. The paper validates AuthorMap in several ways: on human-written documents truncated to 100 words, it reports News accuracies of **76.65** for style and **71.64** for content, and Review accuracies of **89.00** for style and **82.69** for content; human evaluation reports **Randolph’s kappa = 0.40** and agreement of AuthorMap with human choice of **80%** style and **73%** content on News, and **73%** style and **80%** content on Review. These results support the metric’s usefulness, but they do not make it equivalent to direct user satisfaction.

Within the broader summarization literature, ComPSum belongs to personalized multi-document summarization rather than generic MDS or source summarization. Its closest conceptual contribution is the idea of **comparative personalization**: personalization quality improves when user modeling explicitly identifies how one user differs from other users writing about the same topic. The paper’s evidence supports that claim empirically. A plausible interpretation is that ComPSum’s main novelty lies less in the choice of base LLM than in the structuring of personalization as retrieval, contrast, structured analysis, and grounded generation.

In that sense, ComPSum contributes a general design principle for personalized summarization: user history is most informative when transformed into contrastive, dimension-specific control signals. The framework’s strongest empirical result is that this comparative representation improves both style personalization and content-focus personalization while maintaining factuality and relevance on PerMSum [2509.21562].

Source: https://www.emergentmind.com/topics/compsum