---
title: Common-Set Prompting in Zero-Shot Species Recognition
url: https://www.emergentmind.com/topics/common-set-prompting
type: topic
---

# Common-Set Prompting in Zero-Shot Species Recognition

Searching arXiv for the target paper and nearby work on Common-Set Prompting.
arXiv search: "Common-Set Prompting zero-shot species recognition scientific names"
Common-Set Prompting is a zero-shot recognition strategy for fine-grained species classification in which the text prompt supplied to a vision–language model uses a species’ common English name rather than its scientific Latin or Greek name. In the setting studied with CLIP-like models, the method replaces prompts such as “a photo of *Lepus timidus*” with prompts such as “a photo of mountain hare,” on the premise that common names are far more likely than scientific names to occur in web-scale pretraining corpora such as LAION400M. The reported consequence is a substantial increase in top-1 accuracy on species-recognition benchmarks, with gains of approximately \(2\!\sim\!5\times\) on representative datasets [2310.09929].

## 1. Definition and conceptual basis

The method is defined over a set of species classes \(C\), where each class \(c\) has a scientific name \(s(c)\) and a common English name \(e(c)\). “Scientific-Name Prompting” uses the template
\[
p(c)=\text{“a photo of \{s(c)\}”},
\]
whereas “Common-Set Prompting” uses
\[
p(c)=\text{“a photo of \{e(c)\}”}.
\]
The central claim is that CLIP’s text encoder \(f_t\), having been trained on web text, has encountered common names such as “mountain hare” far more often than scientific names such as “Lepus timidus,” and therefore produces more effective text embeddings for prompts built from common names [2310.09929].

The motivating problem is specific but generalizable: zero-shot VLMs recognize many ordinary objects effectively, yet recognition degrades on highly specialized concepts whose canonical labels are under-represented in pretraining data. Species names are a salient case because scientific nomenclature is often absent from large-scale image–text corpora, even when the corresponding everyday name is common. This suggests that prompt performance depends not only on semantic correctness, but also on lexical alignment with the model’s pretraining distribution.

## 2. Formalization and inference procedure

The inference procedure follows the standard CLIP zero-shot pipeline. For each class \(c\), a text prompt is defined as
\[
p(c)=\text{“a photo of a 〈name(c)〉”}.
\]
Given an image \(I\), the image encoder produces
\[
v=f_v(I)\in\mathbb{R}^d,
\]
and the text encoder produces, for each class,
\[
t_c=f_t(p(c))\in\mathbb{R}^d.
\]
Classification is then performed by cosine similarity:
\[
s(c)=\cos\!\bigl(f_v(I),\,f_t(\text{``a photo of }c\text{''})\bigr),
\]
\[
\hat c=\arg\max_{c\in C} s(c).
\]

Within this formulation, the difference between prompting strategies lies entirely in the choice of \( \text{name}(c)\): the scientific name, the common name, or a frequency-based selection between them. The method therefore does not alter model parameters, training procedure, or embedding geometry; it changes only the textual surface form used to instantiate the class label. In practical terms, Common-Set Prompting is a prompt-construction intervention rather than a training-time adaptation.

## 3. Name translation and the “F-Name” variant

The name-translation pipeline begins by collecting the set of scientific names \(\{s(c): c\in C\}\). Off-the-shelf look-up sources, including Wikipedia infoboxes, the iNaturalist API, and museum catalogs, are then used to construct a dictionary
\[
D: s(c)\mapsto e(c).
\]
For species with no dictionary entry, an optional fallback queries a large-language model with the question: “What is the common English name for the species whose scientific name is \(\{s(c)\}\)?” If this also fails, the procedure falls back to the scientific name.

An optional “F-Name” variant augments this pipeline by precomputing frequency counts \(\mathrm{freqLAION}(w)\) for token sequences \(w\) in the LAION400M text corpus. For each class \(c\), the selected name is
\[
\text{name}(c)=\arg\max_{n\in\{s(c),e(c)\}} \mathrm{freqLAION}(n).
\]
According to the reported experiments, this frequency-based selection typically yields the best or near-best results across model sizes and datasets [2310.09929].

The significance of this variant is methodological. Common-Set Prompting, in its simplest form, assumes that the common name is the better lexical key to the pretrained model’s vocabulary. F-Name refines that assumption by explicitly consulting corpus frequency, allowing the method to retain scientific names in cases where they are not disadvantaged. This is particularly relevant for domains in which formal terminology already has substantial web presence.

## 4. Evaluation setting and empirical performance

The evaluation repurposes four fine-grained classification datasets for zero-shot testing using their held-out validation splits, with per-class averaged top-1 accuracy as the metric: semi-iNaturalist with 810 species spanning mammals, birds, fish, insects, fungi, plants, and related categories; semi-Aves with 200 bird species drawn from iNaturalist2018; Flowers102 with 102 flower categories; and CUB-200-2011 with 200 North American bird species [2310.09929].

| Dataset | Scientific-Name top-1 | Common-Name top-1 |
|---|---:|---:|
| iNat (810-way) | 6.84% / 9.21% | 13.51% / 20.17% |
| Aves (200-way) | 7.05% / 11.10% | 39.80% / 59.00% |
| CUB200 | \(\approx 6\!-\!7\%\) | \(\approx 56\%\) / \(\approx 76\%\) |
| Flowers102 | \(\approx 66\!-\!78\%\) | strong performance also reported |

The iNat and Aves results are reported for OpenCLIP ViT-B/32, with ViT-L/14 in parentheses. On iNat, the move from scientific-name prompts to common-name prompts corresponds to an approximately \(2\times\) gain. On Aves, the improvement is approximately \(5\times\). On CUB200, Scientific-Name Prompting remains poor at approximately \(6\!-\!7\%\), while Common-Set Prompting rises to approximately \(56\%\) with ViT-B/32 and approximately \(76\%\) with ViT-L/14. Flowers102 is the principal exception: all methods already perform strongly with scientific names, at approximately \(66\!-\!78\%\).

The paper also compares name substitution with description augmentation. Prior work had proposed using large-language models to generate species descriptions, such as color and shape attributes, and appending them to prompts. The reported effect is only marginal improvement or no benefit: on iNat, for example, Scientific-Name Prompting improves from \(6.84\%\) to \(8.17\%\) with GPT-4-generated descriptions, whereas Common-Name Prompting changes from \(13.51\%\) to \(14.42\%\); on Aves, description augmentation provides no benefit. A common misconception is therefore that richer textual descriptions necessarily solve specialized zero-shot recognition. The reported evidence indicates that lexical familiarity of the class name itself can matter more than added descriptive detail.

## 5. Why common names outperform scientific names

The proposed explanation is distributional coverage. Web-scale pretraining corpora such as LAION400M contain far more mentions of names like “red squirrel” or “bald eagle” than of “*Sciurus vulgaris*” or “*Haliaeetus leucocephalus*.” For iNat’s 810 species, the reported coverage counts are
\[
\mathrm{LAION}\cap\{s(c)\}=468,\qquad \mathrm{LAION}\cap\{e(c)\}=781.
\]
By using common names, prompts align more closely with the VLM’s learned vocabulary and co-occurrence statistics [2310.09929].

This account also explains the Flowers102 result. In that dataset, most flower Latin names occur in LAION as often as their common names, and all methods perform strongly with scientific-name prompts. The implication is not that common names are universally superior, but that the better prompt label is the one with stronger support in the model’s pretraining text. The “common set” in the method’s name is therefore not merely a linguistic preference for everyday wording; it denotes the subset of terms that the model has likely seen often enough for reliable text-side representation.

## 6. Limitations, best practices, and broader implications

The method requires constructing and curating an accurate scientific-to-common dictionary. Ambiguous or region-specific common names can introduce noise, and some species lack widely agreed common names or have multiple alternatives, such as “puma” versus “mountain lion,” which may confuse the model. The optional LLM fallback can also fail or hallucinate. Best practices are therefore reported as follows: begin with a high-quality external taxonomy or museum database; verify name coverage in the VLM’s pretraining corpus if possible; back off to scientific names only when no reliable common name exists; and optionally choose among variants with corpus frequency via F-Name [2310.09929].

The broader implication stated in the work is that the same principle may extend beyond taxonomy to specialized domains whose formal jargon is under-represented in pretraining data. Examples given include mapping chemical IUPAC names to common trade names, medical codes to layman terms, and technical product IDs to brand names. This suggests a more general prompt-engineering heuristic: when zero-shot recognition depends on label text, performance may improve if the label is mapped to the terminology that the pretrained model is more likely to have encountered. At the same time, the reported limitations indicate that such remapping is only as reliable as the underlying lexical resource and the stability of the target naming convention.

Source: https://www.emergentmind.com/topics/common-set-prompting