AuthorMap: Author-Centered Mapping Frameworks
- AuthorMap is a collection of distinct author-centered frameworks that visualize researcher profiles, resolve developer identities, and evaluate personalized summaries.
- Each usage employs specific techniques—TFIDF with PCA for scholarly visualization, logistic edge classifiers for identity resolution, and pairwise LLM evaluation for personalization.
- The design prioritizes precision over recall by preventing alias clumping, using methods like Gaussian mixture modeling and BM25 retrieval to ensure reliable author mapping.
In the cited literature, AuthorMap denotes more than one kind of author-centered mapping framework. One usage, derived from PeopleMap at Georgia Tech, is an interactive, open-source, web-based system that visually maps researchers from Google Scholar publication metadata by combining TFIDF embeddings, PCA, Gaussian mixture modeling, and cosine-similarity query matching (Saad-Falcon et al., 2020). A second usage, introduced by Audris Mockus for World of Code, is a curated global author-identity map that resolves raw author and committer strings into canonical developer identities across 5,866,595,698 commits and ships provenance-aware artifacts for downstream analytics (Mockus, 7 Jul 2026). A third usage, in personalized multi-document summarization, is a fine-grained, reference-free evaluation framework that measures personalization by asking whether a user’s profile can be correctly attributed from generated summaries along the dimensions of writing style and content focus (Li et al., 25 Sep 2025).
1. Terminological scope and conceptual commonality
The term is applied to distinct computational problems rather than to a single standardized system. In the researcher-discovery setting, AuthorMap is a visualization of institutional expertise. In software-repository mining, it is an identity-resolution layer that prevents alias inflation and over-merge in developer analytics. In personalized summarization, it is an evaluation instrument based on pairwise authorship attribution. These usages are technically different, but all treat the author as an entity that must be represented, aligned, disambiguated, or attributed.
| Usage | Primary input | Primary output |
|---|---|---|
| Researcher mapping | Google Scholar profiles and publication metadata | Interactive 2D map of researchers |
| Developer identity mapping | Raw author/committer strings over WoC commits | Canonical identities, classes, provenance tables |
| Personalized MDS evaluation | User profiles and two personalized summaries | Style/content attribution accuracy |
A related precursor is GMap, a graph-visualization method that draws relational data as geographic-like maps. The accompanying implementation-oriented summary explicitly frames GMap as a basis for building an AuthorMap for a co-authorship network, with authors as vertices and co-authorship relations as edges (0907.2585). This suggests that “AuthorMap” functions partly as a task label: a map whose semantic unit is the author, whether the underlying substrate is text, commits, or a collaboration graph.
2. Researcher mapping from Google Scholar profiles
In the PeopleMap line of work, AuthorMap is an institutional researcher-discovery system designed to complement rather than replace manually curated directories. The motivating problem is that university directories often become out of date and rarely provide visual summaries of a researcher’s work or an easy way to explore shared interests. The system requires only researchers’ Google Scholar profiles as input and collects profile URLs, Google Scholar keywords, citation counts, affiliations, profile photos, the top 50 most cited publications, and the top 50 most recent publications, including titles, abstracts, and years (Saad-Falcon et al., 2020).
The core representation is a per-author combined document produced by concatenating titles and abstracts of a researcher’s publications and, optionally, Google Scholar keywords. Text preprocessing removes words with non-English alphabet characters, eliminates words with fewer than two characters, lowercases all words, cleans HTML tags, removes stop-words, and stems words. TFIDF is then used to construct researcher embeddings, with each researcher represented as a column in a TFIDF matrix:
The platform includes a Keywords Emphasis control that increases or decreases the weight of self-identified keywords by concatenating multiples of them into the combined document. Dimensionality reduction is performed with PCA, reducing several-thousand-dimensional TFIDF vectors to two dimensions for visualization. The reported rationale is methodological rather than aesthetic: given small sample sizes relative to dimensionality, PCA is preferred as a linear model, whereas non-linear techniques such as UMAP or t-SNE may find structure within noise under these conditions (Saad-Falcon et al., 2020).
Topic alignment is computed by applying the same TFIDF preprocessing pipeline to a user-specified topic string and then measuring cosine similarity against each researcher embedding:
Higher scores indicate greater alignment with the query topic. In the interface, Map View renders each researcher as a dot in the PCA projection; cluster coloration is supplied by Gaussian mixture modeling on the reduced embeddings; Research Query shades dots by topical alignment; Researcher View reveals name, affiliation, Google Scholar keywords, citation count, profile link, and profile photo; and the Control Panel exposes Show Distributions, #Clusters, Show All Names, Keywords Emphasis, and Publication Set controls (Saad-Falcon et al., 2020).
Several technical points are important for interpretation. The system does not construct an explicit author graph with co-authorship edges; inter-author similarity is implied by proximity in the 2D PCA map. It also does not describe a separate name-disambiguation module within PeopleMap; Google Scholar profiles, including URLs, serve as canonical identifiers. Freshness is manual rather than automatic: users re-run the data collection and processing step when profiles change. Deployments reported in the papers include the Institute for Data Engineering and Science at Georgia Tech with 83 researchers, the Center for Machine Learning, and the Department of Chemistry and Biochemistry, and the generated maps are static web applications that can be hosted without a backend computation server (Saad-Falcon et al., 2020).
3. Global author-identity mapping for software repositories
In World of Code, AuthorMap is a curated identity-resolution layer for public version-control data. Its purpose is to solve the global author identity problem: the same developer commits under many name/email strings, while the same string is reused by many developers. The release for WoC version V2604 covers 5,866,595,698 commits and 106,826,059 distinct raw author/committer strings, which are folded into 62,670,110 canonical developer identities. The paper states that the design problem is clumping, not recall: naive transitive union over shared-attribute edges welds about three million unrelated people into one cluster, so preventing over-merge takes precedence over maximizing alias recovery (Mockus, 7 Jul 2026).
The system ships four co-versioned artifacts.
| Artifact | Schema | Role |
|---|---|---|
a2AFullSUG |
rawId;canonicalId |
Global alias map |
A2clsFull |
member;canonical;class |
Per-identity classification |
P2aAFull |
project;rawId;A;rule |
Within-project recovery |
c2AFull |
commit;A;prov |
Commit-to-identity join |
a2AFullSUG maps each raw author/committer string to a canonical developer identity and is described as mega-cluster free, with largest cluster size 6,910 and no clusters above 10k. A2clsFull assigns every raw id to one of five classes: good, bad-by-attribute, local, bot, or partial. The reported counts over 106,826,059 ids are: good 100,814,372 (94.37%), bad-by-attribute 2,652,369 (2.48%), local 2,562,118 (2.40%), bot 553,736 (0.52%), and partial 243,464 (0.23%). P2aAFull performs project-local recovery for ambiguous identities via anchor and solo rules, contributing 5,359,181 assignments. c2AFull tags every commit with a resolved identity and provenance code: g for global resolution, B for bots, p for within-project recovery, and s for self/unresolved; the reported distribution is g=91.53%, B=5.14%, p=1.59%, s=1.74% (Mockus, 7 Jul 2026).
The six-stage pipeline is ordered to prevent clumping before recovering recall: link generation, value gating, structural gating, edge classification, recall recovery, and representative selection with classification. Structural gating is especially central. On the exact union graph, sampled betweenness is used to cut high-betweenness bridge ids that weld otherwise-disjoint communities; about ~2,000 such bridges disconnect the mega-cluster, and 98% of them are reported as attribute-clean and therefore invisible to value-based rules. Residual homonym welds are then pruned by a logistic edge classifier trained on 2.6M free within-handle GitHub-id labels. Recall recovery composes pairwise shingle expansion v1 at \tau=0.9, map-level shingle expansion v2.1 at \tau=0.8, and ghid assertions for GitHub noreply same-account identities (Mockus, 7 Jul 2026).
The main reported outcomes are operationally significant. The map resolves 73.5% of all six billion commits into a multi-id identity. Human-id commit coverage is 96.49% with the global map alone and rises to 98.17% when within-project recovery is added. Against the ALFAA human-rated gold set of 469k pairs, the map scores precision 0.879 and recall 0.703, corresponding to splitting 0.297 and clumping 0.121. The comparison to the prior WoC V3 map is explicitly polemical: a headline precision of 0.949 collapses to 0.522 once the 3,006,318-id conflated region is counted, illustrating the claim that recall-only or clumping-blind evaluation can be badly misleading (Mockus, 7 Jul 2026).
For downstream usage, the paper recommends excluding bots, preferring prov ∈ {g, p} for human analyses, and treating s as unknown. It also advises joining external scholarly graphs on canonical A, not raw strings, and restricting to good-class ids when performing cross-corpus joins to resources such as OpenAlex or Semantic Scholar (Mockus, 7 Jul 2026).
4. AuthorMap as a reference-free evaluation framework for personalized summarization
In the personalized multi-document summarization literature, AuthorMap is neither a visualization system nor an identity map. It is a reference-free evaluation framework introduced alongside ComPSum to assess whether personalized summaries preserve fine-grained user preferences. The setup assumes a document set on a single topic, a set of users with profile documents , and a system that generates personalized summaries . Because no user-specific reference summaries exist, AuthorMap evaluates personalization indirectly by testing whether an LLM judge can attribute a user’s profile to the correct one of two candidate summaries, separately for writing style and content focus (Li et al., 25 Sep 2025).
The retrieval stage uses BM25. For two users and , AuthorMap queries each profile with the concatenation and retrieves the top profile documents, with . Each retrieved document is truncated to 100 words. The use of both summaries in the query is intended to avoid favoring either summary and to reflect the topic-conditional nature of user preferences. Judging is then performed twice per profile, once in each summary order, to mitigate positional bias:
0
1
with swapped-order counterparts for both profiles. Majority voting is applied per profile, and a sample is counted as correct for a dimension 2 only if both profiles are attributed correctly:
3
The dataset-level score is
4
where higher 5 indicates stronger personalization on that dimension (Li et al., 25 Sep 2025).
The framework explicitly separates style from content through dimension-specific prompts. The style prompt emphasizes linguistic features such as modal verbs, typos, and tone; the content prompt emphasizes aspects and topical focus. The main judge is Llama3.3-70b-Instruct, and Gemma-3-27b-it is reported to yield consistent system rankings. On human-written documents, the reported automatic evaluation scores are: News style 76.65% and content 71.64%; Reviews style 89.00% and content 82.69%. Retrieval improves accuracy relative to a no-retrieval baseline, and human evaluation reports Randolph’s Kappa of 0.40, with AuthorMap accuracy versus human labels of 80% and 73% for News style and content, and 73% and 80% for Reviews style and content (Li et al., 25 Sep 2025).
AuthorMap is also used to evaluate ComPSum against baselines such as RAG, CICL, RAG+Summary, DPL, and Rehearsal. The paper reports that ComPSum improves personalization while maintaining factuality and relevance, producing the best overall score across Llama3.1-8b-Instruct, Qwen2.5-14B-Instruct, and Llama3.3-70b-Instruct. A central interpretive point is that AuthorMap controls for system confounds by comparing two summaries produced by the same system for two users on the same document set, unlike comparisons that might conflate user-specific signals with model-specific stylistic idiosyncrasies (Li et al., 25 Sep 2025).
5. Cartographic metaphors and co-authorship visualization
A distinct but closely related tradition comes from graph drawing. GMap is not itself named AuthorMap, but the supplied implementation summary explicitly treats it as a practical basis for constructing an AuthorMap for co-authorship data. In this formulation, authors become vertices, co-authorship relationships become edges, and clusters of authors are rendered as contiguous geographic regions, or “countries,” within larger “continents” (0907.2585).
The pipeline has three major stages: 2D embedding, clustering, and map generation. Embedding choices include PCA, MDS, force-directed methods, LLE, and Isomap, although the paper most often couples a scalable force-directed layout with modularity-based clustering. For author networks, scalable force-directed layout is described as effective for revealing graph-theoretic proximity while distributing dense areas. Clusters are then used to merge multiple vertex cells into larger regions. The map itself is constructed from a Voronoi diagram over layout points augmented with artificial points around label bounding boxes to inflate area proportionally to label size and random outskirt points to create smoother coastlines, lakes, bays, and other natural-looking boundaries (0907.2585).
This geographic metaphor is not purely decorative. Label size can encode influence, such as publication count or citation count, and thus expand the area associated with an author. Countries may be non-contiguous if a cluster fragments in the embedding, and the coloring procedure operates on a derived “country graph” to maximize contrast along adjacent borders. The paper notes that physical adjacency on the map may produce false positives and that some graph adjacencies cannot be represented as shared borders, producing false negatives. These are treated as intrinsic trade-offs of map-based graph visualization rather than as implementation bugs (0907.2585).
The collaboration example in the paper uses the GD Symposium author graph with 509 vertices and 1517 edges. The resulting map is reported to reveal meaningful regions such as European clusters, Italian and Spanish groups, Australasian and North American clusters, and smaller islands for Japanese and Czech groups. In this sense, GMap supplies a cartographic vocabulary for author-centered maps that is structurally different from the TFIDF-PCA pipeline of PeopleMap but conceptually similar in its attempt to make author neighborhoods legible (0907.2585).
6. Limitations, misconceptions, and future directions
Across the cited uses, AuthorMap should not be treated as a single interchangeable methodology. In the PeopleMap lineage, a common misconception is to read the visualization as a collaboration network. The papers state the opposite: there is no explicit author graph with edges, and 2D position comes from PCA over TFIDF embeddings. Topic-query alignment is based on proportional term usage and is described as a reference point rather than a full assessment of contributions. Additional limitations include reliance on public Google Scholar profiles, absence of automatic updates, lack of a separate disambiguation module, and overplotting when the number of researchers grows. Planned improvements include richer embeddings from Transformer models such as BERT and SciBERT, topic labeling with methods such as NMF, and usability studies with administrators and industry stakeholders (Saad-Falcon et al., 2020).
In the World of Code usage, the central misconception is that high recall alone is desirable. Mockus argues that the binding constraint at global scale is clumping: a single careless merge source can create mega-clusters that corrupt population estimates, productivity measures, bus or truck-factor analyses, collaboration networks, and cross-corpus joins. The map is therefore deliberately conservative, leaving 1.83% of human commits unresolved and preserving ambiguous values as local, bad-by-attribute, or s where safe global merging is unavailable. Remaining limitations include residual splitting and clumping, heuristic sensitivity in name-spread and blacklist rules, incomplete benchmark coverage, and the need for better bot detection, stronger within-project disambiguation, and richer cross-corpus corroboration signals such as DOI anchoring and institutional email evidence (Mockus, 7 Jul 2026).
In personalized summarization, the key misconception is to treat AuthorMap as a generation model. It is an evaluation framework. Its limitations follow from that role: LLM judges may encode demographic or stylistic biases, positional bias must be mitigated by dual-order judging, complete disentanglement of style and content is difficult, and user profiles in public corpora may not fully represent real-world personalization signals. The framework also treats ties as non-matches, which can undercount partial signals in ambiguous cases. Even so, the reported robustness across judge models and the expected behavior under paraphrasing manipulations suggest that the method captures meaningful personalization differences rather than only superficial lexical overlap (Li et al., 25 Sep 2025).
Taken together, these lines of work show that AuthorMap is best understood as a family of author-centered mapping and attribution frameworks. One branch maps researchers by textual embeddings for discovery; another resolves developer aliases at global scale for trustworthy repository analytics; another evaluates personalized summaries through pairwise authorship attribution; and related cartographic graph-drawing methods provide a visual grammar for author neighborhoods and communities. The shared research challenge is not simply to display authors, but to make author identity, similarity, and differentiation computationally reliable.