---
title: 'DIT: Dissimilarity in Trades'
url: https://www.emergentmind.com/topics/dissimilarity-in-trades-dit
type: topic
---

# DIT: Dissimilarity in Trades

Searching arXiv for the cited works and recent context on “Dissimilarity in Trades”.
Dissimilarity in Trades (DIT) is a non-unified term used in several research literatures to denote formally different measures of non-equivalence across traded objects, trading relations, or trading agents. Across the cited works, DIT refers to partition dissimilarity in multilayer international trade networks, structural and attribute-level dissimilarity in industrial trade-credit graphs, the Index of Dissimilarity applied to trades or occupations, quality-gap measures in intra-industry trade, co-occurrence-based distinctions in high-frequency equity trading, a stress-based rarity metric and embedding for NFT markets, and cosine-angle dissimilarity for pairing fantasy-football teams in player-trade optimization. This suggests that DIT is best understood not as a single canonical statistic but as a family of domain-specific constructions whose common purpose is to quantify how different two trade-related entities or structures are.

## 1. Taxonomy of usages

The term is attached to at least seven distinct mathematical objects.

| Domain | DIT construction | Primary object compared |
|---|---|---|
| International-trade multi-network | $\mathrm{DIT}(A,B)=1-\mathrm{NMI}(A,B)$ | Community partitions |
| Industrial trade-credit network | $\mathrm{DIT}_{\mathrm{degree}}=-r$; $\mathrm{DIT}_{\mathrm{attr}}=\frac{1}{|E|}\sum_{(i,j)\in E}|x_i-x_j|$ | Neighbor connectivity and node attributes |
| Labor-market segregation | DIT as the Index of Dissimilarity (ID), and standardized ID (SID) | Group distributions across trades/occupations |
| Intra-industry trade | $DIT_i=\left|\ln VUX_i-\ln VUM_i\right|=\left|\ln r_i\right|$ | Export/import unit values |
| Equity-market microstructure | Composition- and signal-based DIT derived from trade co-occurrence and COI | Stock-day trade-flow patterns |
| NFT markets | Stress-style $F$ and a one-dimensional NM-wMDS embedding | Pairwise trade-derived dissimilarities |
| Fantasy-football trade optimization | $\theta(x,y)=\arccos\!\left(\frac{x\cdot y}{\|x\|\|y\|}\right)$ | Team feature vectors |

The shared intuition is comparison under a trade-induced geometry: partitions, degrees, attributes, prices, exposure patterns, or roster states are embedded into a metric or quasi-metric space, and DIT quantifies separation within that space. The technical content, however, is domain-dependent and not interchangeable across applications [1009.1731] [1409.8588] [2503.02763] [2307.10660] [2209.10334] [2508.12671] [2111.02859].

## 2. Community-partition dissimilarity in international trade networks

In the international-trade multi-network literature, DIT is defined on community partitions of directed weighted trade layers. The underlying data comprise a balanced panel of $N=162$ countries, $C=97$ commodities at HS1996 2-digit resolution, and years $1992$–$2003$. Each commodity layer is a weighted directed network with weight matrix $X_t^c=\{x_{ij,t}^c\}$, where $x_{ij,t}^c$ is the value of exports of commodity $c$ from $i$ to $j$ in year $t$. The aggregate International Trade Network is
\[
x_{ij,t}=\sum_{c=1}^{C} x_{ij,t}^c.
\]
Communities are uncovered by modularity maximization using a tabu search heuristic on weighted directed networks, and similarity between partitions is measured by the confusion-matrix-based normalized mutual information index (NMI), bounded in $[0,1]$ with $1$ for identical partitions and $0$ for independent partitions. DIT is then the complement of NMI,
\[
\mathrm{DIT}(A,B)=1-\mathrm{NMI}(A,B),
\]
so that $0$ denotes identical community structure and $1$ maximal dissimilarity.

This formulation is used for several comparisons: commodity versus aggregate partitions, commodity versus geography-induced or RTA-induced partitions, and inter-commodity comparisons. Geography is encoded by an inverse-distance network $s_{ij}=d_{ij}^{-1}$, whereas RTAs form a weighted undirected network whose entries count agreements in force. The main empirical result is that commodity-specific community structures are highly heterogeneous and substantially more fragmented than the aggregate ITN. Aggregate density rises from $0.2260$ in 1992 to about $0.5400$ in 2003, whereas commodity-layer densities are typically only $12$–$65\%$ of the aggregate; the aggregate network has fewer communities than most commodity layers, with $2$ communities in 1992 and $4$ in 2003, while arms rises from $7$ to $9$. The Herfindahl index of community-size concentration for the aggregate falls from $0.50$ to $0.31$, while commodity patterns differ sharply. Commodity-versus-aggregate NMI generally increases over time, especially in chemical-related sectors such as mineral fuels and plastics, implying declining DIT for those sectors. Aggregate and commodity partitions are also more similar to geography-induced communities than to RTA-induced communities, and the minimum spanning tree built from $1-\mathrm{NMI}(i,j)$ places science- and technology-based industries together while arms is the most dissimilar commodity layer. The broader implication is that aggregation masks sectoral fragmentation: the relatively coherent aggregate ITN can arise from the superposition of structurally diverse commodity layers [1009.1731].

## 3. Structural and behavioral dissimilarity in industrial trade-credit networks

In the industrial trade-credit literature, the central distinction is between structural dissimilarity in connectivity and behavioral similarity in node attributes. The network is built from invoice discounting at a large Italian bank in 2007, with firms as nodes and directed edges from buyer to supplier, the direction of payments. The trade-credit dataset contains $1{,}578{,}812$ firms connected by $7{,}290{,}072$ links; the intersection with balance-sheet data contains $345{,}403$ firms and $2{,}874{,}830$ links. The graph has a “dandelion-like” topology, a giant component of $101{,}186$ nodes, diameter $20$, many single-link clusters, a power-law supplier in-degree distribution over six orders of magnitude, and a roughly log-normal buyer out-degree distribution.

The paper’s headline characterization is “dissortative from the outside, assortative from the inside.” Structurally, high-degree suppliers tend to connect to low-degree buyers, and the evidence for degree dissortativity is the scaling of average neighbor degree with supplier in-degree,
\[
k_{nn}(K)\approx K^{\alpha}, \qquad \alpha=-1.246.
\]
A precise DIT construction consistent with this framework sets
\[
\mathrm{DIT}_{\mathrm{degree}}=-r,
\]
where $r$ is Newman’s assortativity coefficient or its edge-level correlation analog; since the sign is negative by the paper’s evidence, $\mathrm{DIT}_{\mathrm{degree}}$ is positive. By contrast, trading partners are highly similar in credit rating. Ratings lie on a $1$–$9$ scale and are grouped as $A=1$–$3$, $B=4$–$6$, and $C=7$–$9$. A cross-tabulation over $2{,}802{,}976$ rating pairs yields $\chi^2=2803$ with $56$ degrees of freedom and $p\approx 0$, with strong same-rating tiles in the mosaic; after removing intra-industry pairs at the two-digit NACE level, homophily persists with $\chi^2=1456.1$, $df=49$, $p\approx 0$. A corresponding attribute-level DIT is
\[
\mathrm{DIT}_{\mathrm{attr}}=\frac{1}{|E|}\sum_{(i,j)\in E}|x_i-x_j|,
\]
with normalized form obtained by dividing by $x_{\max}-x_{\min}=8$ for ratings.

A further construct links observed dissimilarity to missingness through information exposure. For supplier $s$,
\[
R_s=\sum_{\{c:(c,s)\in CS\}}R_{cs}, \qquad a_s=\frac{R_s}{S_s},
\]
where $S_s$ is net sales. The relation between average exposure and rating is U-shaped: mid-rated firms minimize exposure, while ratings $1$ and $9$ maximize it. This exposure variable is used to quantify how much of a supplier’s transactional neighborhood is visible to the bank and therefore how much unobserved trade volume remains. The substantive interpretation is dual. Degree dissortativity tends to dampen hub-to-hub distress propagation, but attribute assortativity concentrates vulnerability within rating-homogeneous neighborhoods; low exposure among mid-rated firms further obscures potential contagion channels because missingness is not at random [1409.8588].

## 4. DIT as segregation and quality differentiation

A separate literature uses DIT in the classical segregation sense: the Index of Dissimilarity applied to trades, occupations, or sectors. Let $F_i$ and $M_i$ denote the counts of women and men in occupation $i$, with totals $F=\sum_i F_i$ and $M=\sum_i M_i$. After collapsing occupations into “female” and “male” categories relative to the workforce female share, crude DIT is
\[
\mathrm{ID}=\frac{F_f}{F}-\frac{M_f}{M}
=\frac{1}{2}\sum_i\left|\frac{F_i}{F}-\frac{M_i}{M}\right|.
\]
Its interpretation is the proportion of one group’s workers who would need to change occupational category to equalize distributions. The core critique is that crude ID is inconsistent for cross-country and time-series comparison because it is sensitive to row and column margins: changes in female labor-force participation and in the size of female versus male occupations can alter ID even when the underlying association is unchanged. To remove these structural effects, the Basic Segregation Table is standardized by iterative proportional fitting (IPF) to common target marginals, and the standardized DIT is then computed on the resulting table:
\[
\mathrm{SID}=\frac{\tilde F_f}{\tilde F}-\frac{\tilde M_f}{\tilde M}.
\]
In the paper’s worked example, the difference between two countries’ crude IDs is decomposed into $63\%$ “true segregation” and $37\%$ marginal composition. Empirically, for occupational segregation across $39$ countries and sectoral segregation across $84$ countries, standardization weakens the estimated relationship between log GDP per capita and segregation relative to crude ID, implying that crude DIT overstates the explanatory power of development by conflating segregation with structural composition [2503.02763].

In intra-industry trade, by contrast, DIT measures quality dissimilarity between exports and imports within an industry. Standard IIT intensity is captured by Grubel–Lloyd-type overlap measures, but horizontal versus vertical differentiation requires a quality-gap statistic. A natural, threshold-free DIT metric consistent with the unit-value logic is
\[
DIT_i=\left|\ln VUX_i-\ln VUM_i\right|=\left|\ln r_i\right|,
\]
where $VUX_i=X_i/x_i$, $VUM_i=M_i/m_i$, and $r_i=VUX_i/VUM_i$. This converts price or unit-value asymmetry into a scale-free distance from parity: zero denotes similarity and larger values denote stronger vertical differentiation. Threshold-based methods such as Greenaway–Hine–Milner and Fontagné–Freudenberg classify industries as horizontal or vertical using fixed $\alpha$ bands around $r_i=1$, but the paper emphasizes three limitations of that approach: dependence on arbitrary $\alpha$, aggregation bias from coarse product classifications, and imperfect correspondence between unit values and quality. To address those problems, the paper proposes endogenous decomposition of overlapped trade,
\[
VIIT_i^{endo}=IIT_i\cdot \phi\!\left(\left|\ln r_i\right|\right), \qquad
HIIT_i^{endo}=IIT_i\cdot \left[1-\phi\!\left(\left|\ln r_i\right|\right)\right],
\]
with $\phi(\cdot)$ estimated from the data rather than imposed ex ante. In this usage, DIT is not a segregation index but a continuous measure of quality distance embedded within the decomposition of intra-industry trade into horizontal and vertical components [2307.10660].

## 5. DIT in market microstructure and digital-asset markets

In high-frequency equity trading, the paper does not define DIT explicitly, but a rigorously grounded operationalization follows from its trade co-occurrence framework. For each trade $x_a$ at time $t_a$, all trades arriving within $(t_a-\delta,t_a+\delta)$ are in its $\delta$-neighborhood. This induces five trade types: isolated (iso), non-isolated (nis), non-self-isolated (nis-s), non-cross-isolated (nis-c), and non-both-isolated (nis-b). A Poisson null model yields analytic type probabilities, and the empirical analysis chooses $\delta=1$ ms by maximizing the distance between observed and null type frequencies. At that scale, co-occurrence is abundant: the empirical fractions are iso $28.55\%$, nis $71.45\%$, nis-s $29.75\%$, nis-c $17.27\%$, and nis-b $24.43\%$, versus null values of $81.35\%$, $18.65\%$, $0.09\%$, $18.53\%$, and $0.03\%$. Conditional order imbalance (COI) for each type is
\[
COI_{i,t}^{type}=
\frac{N_{i,t}^{type,buy}-N_{i,t}^{type,sell}}
{N_{i,t}^{type,buy}+N_{i,t}^{type,sell}}.
\]
A plausible DIT construction based on type composition is
\[
DIT^{comp}_{i,t}=p_{i,t}^{iso}-\bigl(p_{i,t}^{nis-c}+p_{i,t}^{nis-b}\bigr),
\]
and a signal-weighted version combines type-specific COIs with the paper’s predictive sign pattern. The empirical findings show strong positive contemporaneous associations between returns and COIs, positive predictive effects for isolated trades, and negative predictive effects for non-isolated cross-stock types. Portfolio strategies built from these decompositions are economically material; the best double-sort, iso/nis-c, attains an annualized return of $34.87\%$ and a Sharpe ratio of $1.79$ [2209.10334].

In NFT markets, DIT is both an explicit performance measure and an explicit rarity meter. The premise is that the only objective signal of market value is trading history, so rarity should be a one-dimensional embedding that preserves pairwise trade-derived dissimilarities. For item pair $(i,j)$, the framework defines a non-negative weight $w_{ij}$ from a time kernel and a trade-derived dissimilarity $\delta_{ij}$ based on weighted absolute log price ratios. A rarity meter $R$ induces distances $d_{ij}(R)=|R_i-R_j|$, and the weighted NM-wMDS stress is
\[
S(\vec R)=\sum_{i<j} w_{ij}\bigl(d_{ij}(R)-\delta_{ij}\bigr)^2.
\]
The associated DIT evaluation metric is the scale-normalized stress
\[
F(R;\mathbb X_N)=
\sqrt{
\frac{\sum_{i,j=1}^{N} w_{ij}\bigl(d_{ij}(R)-\delta_{ij}\bigr)^2}
{\sum_{i,j=1}^{N} w_{ij}d_{ij}(R)^2}
}.
\]
Lower $F$ means that a rarity meter better preserves the trade-implied geometry. On the updated ROAR benchmark of $100$ Ethereum NFT collections and about $4.1$ million trades, using a $70\%$ earliest-trades training split and $30\%$ later-trades test split, DIT achieves the best performance in over $65\%$ of collections. The construction is explicitly non-interpretable: unlike trait-frequency or entropy-based rarity meters, its scores arise from one-dimensional embedding under NM-wMDS rather than from trait-level semantics [2508.12671].

## 6. Complementarity, optimization, and recurring methodological issues

In large-scale player-trade recommendation for fantasy football, DIT is operationalized as cosine-angle dissimilarity between team feature vectors. Each team vector concatenates position-specific importance features and strength features, and candidate team pairs are ranked by
\[
\theta(x,y)=\arccos\!\left(\frac{x\cdot y}{\|x\|\|y\|}\right).
\]
Pairs are sorted in descending order from $90^\circ$, the most dissimilar. This dissimilarity is not merely a matching heuristic: once a pair is selected, opponent-aware “cross valuation” and “positional decay” alter player values before a $0$–$1$ knapsack determines outgoing players subject to a risk-scaled cost constraint. Postprocessing then evaluates parity, impact, pain, and acceptance-oriented signals. In production evaluation over the 2020 and 2021 NFL seasons, high-quality trades rose from $76.9\%$ to $97.3\%$, and the system reported $100\%$ trade uniqueness across its quantum, classical, and rules-based compute modes [2111.02859].

Across these literatures, several recurrent methodological tensions appear. First, DIT is often null-model dependent or target-dependent: community DIT depends on NMI after modularity maximization, standardized segregation DIT depends on chosen IPF marginals, co-occurrence DIT depends on $\delta$ and the market reference set, and NFT DIT depends on kernel weighting and sparse pairwise connectivity [1009.1731] [2503.02763] [2209.10334] [2508.12671]. Second, several papers stress that naive aggregation obscures the object of interest: aggregate ITN communities mask commodity fragmentation, crude ID confounds segregation with margins, and unit-value thresholding can mix horizontal and vertical trade by construction [1009.1731] [2503.02763] [2307.10660]. Third, data incompleteness and observability are central: bank-mediated trade-credit data are missing not at random and require exposure-based reasoning, while NFT DIT can be distorted by wash trading and illiquidity, and fantasy-football cosine dissimilarity can miss magnitude-sensitive roster effects because angle-based measures emphasize direction over scale [1409.8588] [2508.12671] [2111.02859]. Finally, interpretability is uneven. The segregation and community versions are relatively transparent, whereas the NFT meter is explicitly non-interpretable and the high-frequency equity constructions are derived rather than natively defined by the source paper. The common lesson is that any use of “DIT” is only meaningful after the compared objects, the embedding or null model, and the relevant invariances have been made explicit.

Source: https://www.emergentmind.com/topics/dissimilarity-in-trades-dit