FinSim-2: Financial Hypernym Classification Task
- FinSim-2 is a shared task in financial-domain semantic classification that maps financial terms to predefined top-level hypernyms within an external ontology.
- The task employs hybrid methods combining rule-based lexical matching, domain-specific embeddings, and ontology traversal to address data scarcity and imbalance.
- Evaluation focuses on both top-1 accuracy and mean rank, highlighting the importance of precise classification and robust candidate ranking.
Searching arXiv for FinSim-2 papers and related task descriptions. FinSim-2 is a shared task in financial-domain semantic classification centered on hypernym detection: given a financial concept label or term, a system must identify the most relevant top-level concept, or hypernym, from a predefined label set in an external ontology. In the task formulations described in the literature, systems may return a ranked list of candidate labels, and performance is evaluated by both top-1 correctness and the position of the true label in that ranking. Published approaches situate FinSim-2 at the intersection of domain NLP, ontology-based reasoning, and knowledge-graph-enhanced classification, with representative systems drawing on the Financial Industry Business Ontology (FIBO), Word2vec, BERT, WordNet, Wikidata, and WebIsALOD (Perdih et al., 2021).
1. Task definition and semantic scope
FinSim-2 is described as a hypernym detection challenge within the financial services domain. The core problem is formulated as multi-class classification: for each input concept label, the system must assign the most relevant top-level hypernym from a fixed inventory of labels. One description states that the task is to classify concepts from the financial domain into the most relevant hypernym concept in an external ontology, specifically the Financial Industry Business Ontology (Perdih et al., 2021). Another describes the task as classifying financial terms into their most relevant hypernym, or top-level concept, from an external ontology (Keswani et al., 2020).
A more explicit FinSim-2 label inventory consists of ten predefined and mutually exclusive categories: Equity Index, Credit Index, Bonds, Swap, Option, Funds, Future, MMIs (Money Market Instruments), Stocks, and Forward (Portisch et al., 2021). The task permits ranked outputs rather than only single-label predictions, which is important because the official metrics include both accuracy and mean rank. This ranked-output requirement encourages systems that combine hard classification with candidate ordering, for example by merging ontology-derived predictions with probability distributions from supervised models (Perdih et al., 2021).
The task has a clear semantic orientation: it is not merely lexical matching, but top-level concept assignment under domain-specific constraints. At the same time, several system reports note that lexical surface structure remains highly informative. In particular, many terms contain their label as a substring, which creates a class of comparatively easy instances and motivates rule-based handling in some systems (Keswani et al., 2020). This dual character—semantic classification with strong lexical cues—shapes much of the methodological diversity in the FinSim-2 literature.
2. Data configurations and annotation regimes
Published descriptions of the task report two distinct data configurations. One report describes a training set of 100 terms, each paired with its correct hypernym label across 8 possible classes, together with 150 financial prospectus PDFs in English that serve as unlabeled corpus data for embedding training (Keswani et al., 2020). A later FinSim-2 system paper describes a training set of 614 hyponym-hypernym pairs labeled with one of ten classes, plus a test set of 212 entries (Portisch et al., 2021).
This discrepancy indicates that the literature spans closely related task settings rather than a single invariant benchmark. A plausible implication is that FinSim-2 should be understood as a family of shared-task formulations around financial hypernym assignment, with continuity at the level of problem structure even when dataset size and label inventory differ.
Several recurring data characteristics are reported. The labeled sets are small, which constrains the use of data-hungry supervised methods (Keswani et al., 2020). The class distribution can be imbalanced; for the 614-example training set, Equity Index appears 286 times whereas Forward appears only 9 times (Portisch et al., 2021). The corpora used for domain adaptation may also be noisy: PDF conversion introduces noise and token or sentence boundary detection issues during conversion to text (Keswani et al., 2020). Another reported coverage problem is that some terms in the labeled data do not appear in the provided corpus, which has direct consequences for embedding quality when systems rely on corpus-trained vector representations (Keswani et al., 2020).
These properties are not incidental. They explain why many successful systems favor hybrid designs: rule-based components for overt lexical regularities, ontology traversal for explainable structure, and relatively simple classifiers or shallow neural models to exploit whatever supervised signal is available.
3. Ontologies, label spaces, and external knowledge sources
A defining feature of FinSim-2 is its reliance on external knowledge resources. In the ontology-centered formulation, the target space is grounded in FIBO, and each financial label is assumed to correspond to a node in the ontology graph (Perdih et al., 2021). The ontology is transformed into a NetworkX directed, unweighted graph, technically a NetworkX MultiGraph, after which classification becomes a problem of mapping an input concept into that graph and traversing upward to locate an admissible shared-task label.
The JSI system operationalizes ontology mapping in several stages. If a concept exists in FIBO as a node with the same name, direct mapping is used. For multi-word concepts or concepts absent from FIBO, the system splits them into component words, beginning with the last word, assumed to be the noun or head, and attempts singular or plural matching against ontology nodes. If that fails, it searches for synonyms using WordNet. Two labels, Equity Index and Credit Index, receive manual treatment because they lacked direct FIBO representations: Equity Index was manually added as the parent of Index in FIBO, and unmappable concepts default to Credit Index (Perdih et al., 2021).
A different knowledge-centric line of work, represented by FinMatcher, uses three publicly available knowledge graphs—WordNet, Wikidata, and WebIsALOD—to derive both explicit and latent features for each label-class pair (Portisch et al., 2021). Here the external resources are not used primarily for symbolic ontology reasoning, but for feature construction. Explicit features include word overlap, Wikidata hypernym lookup, WordNet hypernym lookup, and WebIsALOD hypernym lookup. Latent features are generated via pretrained RDF2vec embeddings for WebIsALOD nodes, with cosine similarity used to measure semantic relatedness between linked label and class entities (Portisch et al., 2021).
The contrast between these approaches is instructive. In the ontology-graph approach, background knowledge drives prediction directly through graph search. In the knowledge-graph feature approach, background knowledge is converted into numerical signals for a classifier. Both paradigms treat external structure as indispensable, but they place it at different points in the inference pipeline.
4. Representation learning and classification strategies
The FinSim-2 literature exhibits three broad modeling families: embedding-based term-label similarity, ontology-augmented symbolic classification, and feature-based supervised learning with knowledge-graph signals.
The IITK system combines context-independent and context-dependent embeddings. It trains Word2vec embeddings from scratch on the financial prospectus corpus and explores vector sizes from 50 to 500. Each term is represented as the average of its constituent word vectors, and label embeddings are constructed similarly. The same work also uses BERT Base Uncased, with 12 layers, 12 attention heads, and 110 million parameters, fine-tuned on the financial PDF corpus; for each word in a term, the embedding from the last hidden layer is extracted and then averaged to obtain the term representation (Keswani et al., 2020).
Its preprocessing pipeline consists of removal of punctuation, stop words, and special characters; lowercasing; lemmatization and tokenization using NLTK; embedding lookup for tokens; and construction of the term vector as the average of token vectors (Keswani et al., 2020). Classification is then split by a domain rule. Terms containing exactly one label are placed in Group 1 and handled by rule-based direct mapping. Group 2 contains terms where the label appeared twice or the label was missing; these are passed to either unsupervised distance-based methods or simple supervised classifiers. The unsupervised variants compute cosine similarity, minimum distance, or absolute vector distance between a term embedding and each label embedding, assigning the label with maximum cosine similarity. The supervised variants use Bernoulli Naive Bayes and one-vs-rest 8-class logistic regression on average term embeddings (Keswani et al., 2020).
The JSI system takes a different route. It performs ontology-based classification by mapping a concept into FIBO and then generalizing upward through parent nodes until a valid shared-task label is found; this is described as essentially a breadth-first search up the hierarchy with stop conditions. To supplement this symbolic prediction, the system trains a Random Forest classifier on Word2vec vector representations of concepts. For each concept, the classifier predicts a probability distribution over all possible labels, and the ontology-based label is then moved to the top of the ranked list from Random Forest (Perdih et al., 2021).
FinMatcher constructs per-class signal vectors from explicit and latent knowledge-graph features and concatenates them into a single feature vector,
where each is a length-10 signal vector (Portisch et al., 2021). The classifier is a simple feedforward artificial neural network with a single fully connected layer and output size 10, trained with Mean Squared Error for 100 epochs and batch size 25. To address imbalance, it uses SMOTE to upsample minority classes to at least one third of the majority-class count in the training split (Portisch et al., 2021).
Across these systems, a consistent theme is restraint in model complexity relative to data scale. This is explicit in the IITK paper, which states that the small dataset precludes use of complex models (Keswani et al., 2020). Even when neural models are used, as in FinMatcher, the architecture remains shallow and heavily dependent on manually designed knowledge signals rather than large end-to-end fine-tuning (Portisch et al., 2021).
5. Evaluation protocol and reported performance
Two evaluation metrics recur across the reported systems: accuracy and mean rank. Accuracy is the fraction of terms correctly assigned to their true hypernym, or, equivalently, correct top-1 predictions (Keswani et al., 2020). Mean rank, also described as average label rank, measures the average position of the correct label in the returned ranked list; lower values are better (Perdih et al., 2021).
The IITK submission reports that its system ranks 1st based on both metrics, namely mean rank and accuracy (Keswani et al., 2020). The paper attributes this performance to a hybrid strategy combining rule-based mapping, custom-trained Word2vec, fine-tuned BERT, and simple classifiers.
The JSI paper reports several internal and official results. On training data, the ontology-based method achieves 0.87 accuracy, corresponding to 535 out of 614 correctly classified instances, outperforming the provided distance-based classifier at 0.62 and slightly exceeding logistic regression at 0.86. Its Random Forest classifier attains 0.85 accuracy on a 10% held-out set. The merged method yields a slight boost in average label rank, improving from 1.37 for Random Forest alone to 1.19 when the ontology-based prediction is placed first. On the official test set, the system reaches accuracy 0.811, ranking 15th out of 18, and average label rank 1.316, ranking 12th of 18 (Perdih et al., 2021).
FinMatcher reports 81.1% accuracy and mean rank 1.415 on test data. In five-fold cross-validation on the training data, the submission entry reaches 86.7% accuracy and mean rank 1.43 (Portisch et al., 2021). Its ablation study shows that word overlap is particularly important: removing it drops accuracy from 86.69 to 60.51 and worsens mean rank from 1.432 to 2.007. Removing Wikidata, WordNet, WebIsALOD hypernym, or RDF2vec features causes only modest degradation, suggesting partial redundancy among those knowledge signals (Portisch et al., 2021).
These results collectively indicate that FinSim-2 rewards systems that optimize not only class assignment but also ranking quality. They also show that no single modeling family dominates all settings: first-place performance is reported for a hybrid embedding-and-rule system in one configuration, while ontology-centric and knowledge-graph-centric systems achieve competitive but different trade-offs in other settings.
6. Methodological themes, bottlenecks, and interpretive issues
One major methodological theme is hybridization. The IITK system explicitly divides the test data into rule-based and classifier-based subsets (Keswani et al., 2020). The JSI system combines ontology-based prediction with a Random Forest ranked list (Perdih et al., 2021). FinMatcher integrates lexical overlap, symbolic hypernym lookup, and latent graph embeddings in a single classifier (Portisch et al., 2021). This convergence suggests that FinSim-2 is not well served by purely distributional or purely symbolic methods in isolation.
A second theme is the importance of lexical structure. Word overlap is reported as critical in FinMatcher, and the IITK report notes that many terms contain their label as a substring (Portisch et al., 2021). A plausible implication is that a substantial fraction of FinSim-2 instances can be approached through label-aware lexical heuristics before deeper semantic machinery is applied. This does not reduce the task to string matching, but it does mean that robust systems exploit lexical transparency when present.
A third theme is mapping as bottleneck. In the ontology-driven JSI system, errors are typically due to failure in mapping, including ambiguous multi-word expressions and lack of domain synonyms (Perdih et al., 2021). The paper explicitly suggests that better word-to-ontology matching and improved head noun detection could further improve performance. This is a distinct challenge from downstream classification: even a well-structured ontology cannot help if the concept cannot be aligned to the graph.
A fourth issue is data scarcity and imbalance. Small labeled sets limit model complexity (Keswani et al., 2020), while class skew motivates explicit imbalance handling such as SMOTE (Portisch et al., 2021). Combined with corpus noise from PDF conversion and coverage gaps between labeled terms and unlabeled corpora, these constraints explain why many systems emphasize feature engineering, ontology use, and controlled model capacity rather than large supervised architectures.
7. Position within financial NLP and related research directions
Within financial NLP, FinSim-2 occupies a specialized niche: semantic categorization of domain concepts against a structured hypernym inventory. Its significance lies in the fact that the target labels are not arbitrary classes but top-level financial concepts, often grounded in FIBO or analogous ontology structures (Perdih et al., 2021). This gives the task both practical and methodological relevance for semantic indexing, ontology population, and explainable concept classification.
The published systems illustrate three related research directions. One is domain-adaptive embedding learning from financial prospectus text, including both Word2vec trained from scratch and BERT fine-tuned on the provided corpus (Keswani et al., 2020). Another is explainable ontology reasoning through graph transformation and upward traversal in FIBO (Perdih et al., 2021). A third is the combination of explicit and latent knowledge-graph features, including hypernym reachability and RDF2vec similarity, within a classifier (Portisch et al., 2021).
The literature also surfaces a recurring misconception: that ontologies eliminate the need for empirical modeling altogether. The JSI paper notes that ontology-based classification does not require training data for that part of the system, but it also shows that mapping quality is a practical limitation and supplements ontology reasoning with a learned Random Forest ranker (Perdih et al., 2021). Conversely, the strong ablation effect of word overlap in FinMatcher cautions against assuming that sophisticated graph features will necessarily dominate simpler lexical signals (Portisch et al., 2021).
Taken together, the available work presents FinSim-2 as a benchmark family where structured finance knowledge, lexical regularities, and compact statistical models interact under severe data constraints. This suggests that future progress is likely to depend less on scaling model size than on better concept-to-ontology alignment, more robust handling of multi-word financial expressions, and more systematic integration of ranking-oriented inference with domain-specific background knowledge.