---
title: Instance Jigsaw Puzzle Problem
url: https://www.emergentmind.com/topics/instance-jigsaw-puzzle-problem
type: topic
---

# Instance Jigsaw Puzzle Problem

Searching arXiv for recent and foundational papers on the instance jigsaw puzzle problem.
The instance jigsaw puzzle problem concerns reconstructing an original whole from an unordered collection of fragments. In the classical square-tile formulation, the input is the unordered multiset of all \(n\cdot m\) non-overlapping \(K\times K\) image tiles extracted from an unknown original image, with no information about global placement, orientation, or puzzle dimensions, and the objective is to assign each tile to a grid position and orientation so that all abutting edges match [1711.08762]. Closely related formulations distinguish Type-1 puzzles, in which piece orientations are known and fixed, from Type-2 puzzles, in which both locations and orientations are unknown [2203.06488]; model the unknowns as a permutation \(\sigma\in S_n\) together with rotations \(R_i\in Z_4\) for each patch [1811.03188]; extend the pieces from equal squares to convex polygons generated by straight cuts through a global polygonal shape [2008.07644]; or reinterpret the “pieces” as shuffled instances in a bag of whole-slide image features and seek a permutation-restoring operator \(\widehat{\mathbf P}\) [2507.08178].

## 1. Formalizations and problem variants

A standard square-piece instance consists of image tiles arranged originally on a rectangular grid. Each piece has four edges, and a valid reconstruction requires a globally consistent assignment of pieces to cells such that neighboring edges are compatible [1711.08762]. In the terminology used by later reconstruction work, Type-1 refers to unknown locations with known orientations, whereas Type-2 additionally allows each piece to rotate by \(0^\circ,90^\circ,180^\circ,270^\circ\) [2203.06488].

Several important specializations alter the observational model rather than the combinatorial core. In the graph-connection formulation, the observation is a shuffled and arbitrarily rotated set of patches \(Q=\{(R_i(P_{\sigma(i)}),(R_i\circ f)|_{P_{\sigma(i)}})\}_{i=1}^n\), and reconstruction amounts to recovering the image \(f\), or equivalently the permutation \(\sigma\) and rotations \(\{R_i\}\) [1811.03188]. In Deepzzle, the task is image reassembly with wide space between fragments, missing fragments, outsider fragments, and a distinguished central fragment, which changes the scoring problem from direct border continuity to fragment-to-position prediction [2005.12548].

The notion of an instance can also depart from square pieces entirely. Polygonal jigsaw puzzles generated by the lazy caterer model begin with a convex polygonal canvas \(S\subset\mathbb R^2\) and a set of straight cuts \(\{c_1,\dots,c_a\}\), producing convex pieces \(P=\{p_1,\dots,p_n\}\) whose vertices, edges, and internal angles become the primitive data for reconstruction [2008.07644]. A different abstraction appears in the line-only puzzle setting, where each piece \(i\) is a bounded planar region \(\Omega_i\) carrying a finite set of oriented line segments \(L_i\), and compatibility ignores color and conventional shape cues in favor of Wertheimer’s law of good continuation [2410.16857].

A recent domain-specific reinterpretation arises in whole-slide image analysis. There, one starts from a bag of instance embeddings \(\mathbf X=\{\mathbf x_1,\dots,\mathbf x_n\}\in\mathbb R^{n\times d}\), applies a random shuffling operator \(\mathcal S\) to obtain \(\mathbf X_\sigma=\mathbf P_\sigma \mathbf X\), and trains an encoder \(f_\theta\) so that \(f_\theta(\mathbf X_\sigma)\approx \mathbf P_\sigma f_\theta(\mathbf X)\) [2507.08178]. The combinatorial object is again a hidden permutation, but the semantic content is spatial correlation among instances rather than physical edge matching.

## 2. Search space, structure, and hardness

The basic computational obstacle is combinatorial explosion. For square puzzles, exhaustive search over all permutations is infeasible, with complexity \(O((nm)!)\) [1711.08762]. A key structural observation is that full reconstruction need not identify every correct adjacency immediately: the assembly can be viewed as selecting a spanning tree of correct adjacent-edge pairs, and only \(n\cdot m-1\) true adjacencies are needed to connect all pieces [1711.08762]. This is why high-precision compatibility estimates are often more valuable than uniformly high recall; the DNN-Buddies analysis notes that, in the extreme, perfect precision with recall \(<1/8\) can still yield a correct spanning selection of edges [1711.08762].

Complexity-theoretic results show that this difficulty persists even in highly restricted layouts. For rotating and placing \(n\) square tiles into a \(1\times n\) array, exact signed edge-matching (jigsaw) is NP-hard [1701.00146]. The same work proves stronger inapproximability results: it is NP-hard to approximate the max-placement version within factor \(67038719/67038720\approx 0.9999999851\), and the max-matched variant inherits a corresponding hardness result on the number of matched edges [1701.00146]. On the positive side, the same paper gives an easy \(1/2\)-approximation and a \(2/3\)-approximation for the max-tiles version [1701.00146].

Random models illuminate a different aspect of solvability: identifiability from local clues. In shotgun assembly of random jigsaw puzzles on an \(n\times n\) grid whose edge colors are drawn uniformly from \(\{1,\dots,q\}\), earlier work had shown unique assembly for \(q=\omega(n^2)\) and failure of unique edge assembly for \(q=o(n^{2/3})\); the improved result proves that for every fixed \(\varepsilon>0\), if \(q\ge n^{1+\varepsilon}\), then with high probability the puzzle has unique vertex assembly, and a deterministic algorithm with running time \(n^{\Theta(1/\varepsilon)}\) reconstructs the planted assembly [1605.03086]. This line of work is not about natural images, but it clarifies when local matching information is information-theoretically sufficient.

These hardness and identifiability results delimit what compatibility learning and global optimization can plausibly achieve. A plausible implication is that practical solvers are best understood as exploiting strong structure in natural-image statistics, restricted motion groups, or specialized acquisition models rather than circumventing the worst-case combinatorics.

## 3. Compatibility measures and adjacency estimation

Most modern solvers factor the problem into a local compatibility stage and a global assembly stage. Classical square-piece methods use hand-crafted dissimilarity functions on candidate abutting edges. One such boundary score compares the right edge of \(x_i\) to the left edge of \(x_j\) by
\[
D(x_i,x_j,\text{right})=\sqrt{\sum_{k=1}^K\sum_{c=1}^3 [x_i(k,K,c)-x_j(k,1,c)]^2},
\]
which serves as a simple dissimilarity baseline in DNN-Buddies [1711.08762]. The graph-connection pipeline instead initializes pairwise relations with the Mahalanobis-Gradient Compatibility metric across all four relative orientations and directions [1811.03188].

DNN-Buddies introduced the first deep neural network-based estimation metric for the jigsaw puzzle problem [1711.08762]. Its feed-forward fully connected network takes 336 scalar inputs formed from the two abutting pixel columns and their immediate neighboring columns on each piece, using YUV channels, and outputs a two-way softmax indicating whether the two edges should be adjacent [1711.08762]. Training on 970,224 labeled edge-pairs from 2,755 IAPR TC-12 images with informed undersampling yields 95.04% training accuracy and 94.62% held-out test accuracy; when used to select the top-1 most compatible edge on the 20-image, 432-piece benchmark of Cho et al., the reported precision is 94.83% [1711.08762].

Later neural compatibility models concentrated on eroded-boundary settings, where the outermost boundary pixels are missing or unreliable. TEN represents a piece by an embedding produced by twin CNN encoders \(f_l\) and \(f_r\), compares the embeddings with Euclidean distance, and trains them with a triplet loss of margin \(\gamma=1\) [2203.06488]. In the reported experiments, TEN-Large improves top-1 accuracy over the best classical method MGC on all three datasets, and for reconstruction with the GA-based solver raises Type-2 neighbor accuracy from 40.2% to 55.5% on DIV2K, from 47.5% to 59.4% on PIRM, and from 32.3% to 49.1% on MIT [2203.06488]. The same study reports that computing all \(16N^2\) compatibility measures is 450–2770× faster than a conventional end-to-end CNN across puzzle sizes from \(N=100\) to \(N=3200\) [2203.06488].

Edge2Vec refines the embedding approach by modifying the architecture and replacing standard triplet training with hard batch triplet loss [2211.07771]. It defines compatibility as the negative Euclidean distance between learned edge embeddings, uses grouped fully connected projections to reduce parameter count, and adds an \(\ell_2\) regularizer on batch embeddings [2211.07771]. On the reported benchmarks, Edge2Vec improves Type-2 top-1 edge-matching accuracy from 64.1% to 71.4% over the DNN-E2E baseline, and with the GA solver improves Type-2 reconstruction neighbor accuracy from 73.3% to 82.3% [2211.07771].

Not all compatibility measures are color-based. The good-continuation framework computes a purely line-based score \(R_{ij\gamma}\in[0,1]\) by matching oriented line segments across pieces with a linear assignment problem, penalizing unmatched lines and normalizing by a threshold \(\tau\) [2410.16857]. In polygonal puzzles, compatibility can also be defined geometrically through relaxed edge-length and angle predicates under bounded noise \(\varepsilon\), e.g., \(|\,|\hat y|-|\hat y'|\,|\le 4\varepsilon\) together with angle-sum constraints [2008.07644]. This diversity of compatibility models suggests that the local signal depends strongly on the degradation model and on whether the puzzle is pictorial, apictorial, or geometric.

## 4. Global reconstruction and optimization frameworks

A compatibility measure alone does not solve the puzzle; it must be integrated into a global optimizer. In DNN-Buddies, the base solver is a genetic algorithm whose candidate assemblies are represented by relative edge-to-edge assignments [1711.08762]. Its crossover operator builds one child from two parents in five ordered phases: assign all relative relations common to both parents; assign all “DNN-buddy” relations present in either parent; assign all best-buddy relations; assign the most-compatible relations by simple dissimilarity; and assign random relations until \(n\cdot m-1\) relations have been selected [1711.08762]. Elevating DNN-buddy pairs to second priority heavily biases the search toward high-precision learned adjacencies.

The graph connection Laplacian approach separates orientation recovery from shuffle recovery. It builds a weighted connection graph \(G=(V,E,W,R)\), forms the connection adjacency \(S\), the degree matrix \(D\), the graph connection weight matrix \(C=D^{-1}S\), and the normalized graph connection Laplacian \(L=I-C\) [1811.03188]. Rotations are estimated by taking the top two eigenvectors of \(C\), extracting the \(2\times 2\) block for each piece, and projecting it onto the nearest element of \(Z_4\) by Frobenius norm [1811.03188]. Once the rotations are fixed, any standard Type-1 solver can be used to recover the permutation. The authors further describe an iterative update cycle based on a neighbor-averaged MGC score, with up to five iterations, retaining the solution with minimum global Err score [1811.03188].

Deepzzle changes the optimization target from local edge matching to global position assignment. A Siamese CNN predicts \(P_r(x_i=j)\) for a fragment \(i\) occupying position \(j\), and these scores define a directed acyclic graph whose shortest path represents the maximum-probability assignment under a simplified product model [2005.12548]. The graph is pruned by a branch-cut threshold \(\theta\): with \(\theta=0\), inference takes approximately 20,000 s per puzzle; with \(\theta=0.01\), approximately 2,000 s with no accuracy loss; with \(\theta=0.05\), approximately 20 s with an approximately 1% accuracy drop [2005.12548]. This is specifically suited to settings with large gaps, outsider fragments, and missing pieces.

GANzzle reframes the problem as retrieval from a generated “mental image” rather than pairwise edge comparison [2207.05634]. A permutation-invariant encoder pools piece features into a global latent code \(z\), a multi-scale GAN reconstructs an image \(G(z)\), RoIAlign extracts per-slot features from a generator layer matching the puzzle grid, and a shared embedding space yields similarities \(C_{ij}=u_i^\top v_j\) between piece and slot embeddings [2207.05634]. Global assignment is then enforced with a Sinkhorn relaxation followed by Hungarian attention, making the matching process end-to-end [2207.05634].

Alphazzle applies single-player Monte Carlo Tree Search to the full permutation space. A state consists of the next patch to place, the current partial reassembly, and the occupancy dictionary of the canvas; legal actions are the empty slots [2302.00384]. Selection uses a PUCT rule with policy prior \(\pi_\theta(a\mid s)\), leaf evaluation uses a learned value network \(V(s)\) instead of rollouts, and self-play fine-tuning on MCTS-visited states provides further improvement [2302.00384]. In contrast, polygonal puzzles are solved through a multi-body spring-mass system combined with hierarchical loop constraints and layered reconstruction [2008.07644], while the good-continuation formulation encodes all placements as a polymatrix game and finds a Nash equilibrium with multi-population replicator dynamics [2410.16857]. The field therefore spans evolutionary search, spectral synchronization, shortest-path optimization, end-to-end assignment, tree search, physical simulation, and game-theoretic equilibrium computation.

## 5. Evaluation protocols and reported performance

Reported metrics are highly problem-dependent. Classical square-piece benchmarks frequently use neighbor comparison, defined as the fraction of correctly reconstructed abutting-edge pairs, together with the number of perfectly solved puzzles [1711.08762]. The graph-connection literature additionally reports direct accuracy, largest component, and perfect reconstruction counts [1811.03188]. Deepzzle introduces an “almost-perfect” reassembly criterion: for each slot \(j\), if the pixel-wise norm between the ground-truth fragment and the predicted fragment is at most \(T\), the swap is deemed visually acceptable, with \(T=20\) reported as best matching human tolerance [2005.12548]. GANzzle evaluates piece-to-slot retrieval by \(R@1\), and the whole-slide image reinterpretation uses Acc, F1, AUC, Avg.3, and C-index depending on whether the task is classification or survival prediction [2207.05634] [2507.08178].

Representative results illustrate how strongly performance depends on the task definition and solver.

| Method | Setting | Reported result |
|---|---|---|
| DNN-Buddies + GA | 540-piece benchmark | 96.37% neighbor accuracy; 11 perfect puzzles [1711.08762] |
| Deepzzle | Full \(9!\) permutation space | 39.2% image-level reassembly accuracy [2005.12548] |
| Alphazzle | Full \(3\times3\) permutation problem | 51.5% full-puzzle, 75.1% patch-wise, 77.5% neighbor-wise, in \(\sim 16\) s/puzzle [2302.00384] |
| TEN-Large + GA | Type-2 eroded puzzles | 55.5%, 59.4%, 49.1% neighbor accuracy on DIV2K, PIRM, MIT [2203.06488] |
| Edge2Vec + GA | Type-2 eroded puzzles | 82.3% neighbor accuracy [2211.07771] |

On the standard benchmarks used in DNN-Buddies, adding the learned metric to the GA baseline raises neighbor accuracy from 94.88% to 95.65% on 432-piece puzzles, from 94.08% to 96.37% on 540-piece puzzles, and from 94.12% to 95.86% on 805-piece puzzles; the number of perfectly reconstructed puzzles increases by \(+1,+3,+2\), respectively [1711.08762]. The 96.37% result on the 540-piece set exceeds the previous best of 95.4% reported by Paikin and Tal (2015) [1711.08762]. The same paper also notes that no formal statistical-significance tests or ablations beyond GA versus GA+DNN are provided [1711.08762].

Performance claims in difficult settings are more mixed. Deepzzle reports 24.7% almost-perfect puzzles and 64.6% well-placed fragments with no outsiders and no missing fragments, and under known versus unknown central fragment reports 44.4% versus 39.2% puzzle accuracy and 89.9% versus 71.1% fragment accuracy [2005.12548]. GANzzle is explicitly puzzle-size agnostic and, on PuzzleCelebA, reports direct accuracies of 72.18% on \(6\times6\), 53.26% on \(8\times8\), 32.84% on \(10\times10\), and 12.94% on \(12\times12\) [2207.05634]. This suggests that evaluation should be read together with the observation model: wide gaps, erosion, outsiders, unknown central fragments, variable puzzle sizes, and arbitrary shapes are not commensurate regimes.

## 6. Generalizations, cross-domain uses, and technical limits

A major trend is the move from ideal square tiles toward degraded or structurally richer inputs. Eroded-boundary puzzles motivate learned embeddings that use the entire piece rather than only the outermost pixels, because the boundary is unreliable by construction [2203.06488] [2211.07771]. Archaeological-style reconstruction with large gaps and outsider fragments motivates position-prediction models and branch-cut graph search rather than strict edge continuity [2005.12548]. Polygonal puzzles introduce bounded geometric noise, relaxed mate predicates, hierarchical loop consensus, and physical settling via a spring-mass dynamical system [2008.07644]. The line-only problem goes further by deliberately ignoring conventional color and shape features and relying solely on geometric line continuation [2410.16857].

Another trend is conceptual transfer beyond image reassembly. Jigsaw percolation defines two graphs on the same vertex set—a social graph \(G_{\mathrm{soc}}\) and a puzzle graph \(G_{\mathrm{puz}}\)—and iteratively merges clusters when there is both a social edge and a puzzle edge between them [1207.1927]. For an Erdős–Rényi social network and the ring puzzle, the critical probability satisfies \(p_c(n)=\Theta(1/\log n)\); for any connected puzzle graph of bounded maximum degree, \(p_c(n)=O(1/\log n)\) and \(p_c(n)=\omega(1/n^b)\) for every \(b>0\); and power-law social networks with \(\alpha>2\) cannot solve bounded-degree puzzles with high probability [1207.1927]. This is not a reconstruction algorithm, but it recasts “solving a puzzle” as a phase transition in joint graph connectivity.

In whole-slide image analysis, the instance jigsaw puzzle becomes a regularizer against permutation-invariant aggregation. The proposed shuffling-equivariance loss is
\[
\mathcal L_{\rm Equiv}=\frac{1}{2n}\left\|f_\theta(\mathbf X_\sigma)-\mathbf P_\sigma f_\theta(\mathbf X)\right\|_2^2,
\]
which encourages the network to undo the random permutation in feature space [2507.08178]. On CAMELYON16, the Transformer variant reports 96.64 Acc, 96.39 F1, 98.00 AUC, and 97.01 Avg.3, compared with 96.31 Avg.3 for the best MIL baseline TransMIL; on TCGA-NSCLC, it reports 96.13 Avg.3 versus 95.59 for the best MIL baseline [2507.08178]. For survival prediction, the same model reaches C-index 0.732 on TCGA-LUAD and 0.715 on TCGA-BRCA, compared with 0.714 and 0.701 for the best MIL baselines [2507.08178]. Here the “pieces” are not physical fragments, yet the hidden permutation and the recovery of spatial correlations preserve the central logic of the jigsaw formulation.

The literature also records clear limitations. DNN-Buddies addresses class imbalance by informed undersampling, but does not explore more advanced remedies and does not detail the computational cost of classifying all \(O(n^2m^2)\) edge pairs [1711.08762]. The graph-connection method is dominated by the \(O(n^2)\) MGC precomputation and by the Type-1 solver [1811.03188]. Deepzzle explicitly notes that the independence approximation \(P(x_i\mid x_{i-1})\approx P(x_i)\) ignores spatial context beyond center-fragment pairs and that exponential graph size limits massively large puzzles beyond approximately 20 fragments despite pruning [2005.12548]. These limitations help explain why the field remains methodologically plural: no single formulation currently covers natural-image square puzzles, eroded fragments, polygonal pieces, line-only evidence, and instance-order recovery in bag-level learning with equal effectiveness.

Source: https://www.emergentmind.com/topics/instance-jigsaw-puzzle-problem