---
title: 'PuzzlePlex: Benchmark for Foundation Models'
url: https://www.emergentmind.com/topics/puzzleplex
type: topic
---

# PuzzlePlex: Benchmark for Foundation Models

PuzzlePlex is a benchmark for evaluating the reasoning, planning, and generalization abilities of foundation models through a curated suite of interactive puzzles. Introduced in “PUZZLEPLEX: Benchmarking Foundation Models on Reasoning and Planning with Puzzles” [2510.06475], it comprises 15 rule-based puzzle types spanning single-player and two-player settings, deterministic and stochastic dynamics, text-only and text-image modalities, multiple difficulty levels, and two evaluation modes—instruction-based interaction and code-based execution. Rather than functioning as a static question-answer dataset, it is structured as an extensible environment for studying long-horizon problem solving, dynamic adaptation, strategic interaction, and executable program synthesis.

## 1. Rationale and benchmark scope

PuzzlePlex was introduced to address a gap in existing evaluations of foundation models: many prior puzzle benchmarks emphasize single-player tasks, short-horizon interactions, mostly deterministic settings, little or no competition, limited multimodality, no code-execution-based evaluation, and puzzle families that may already be common in model pretraining data [2510.06475]. Its stated aim is therefore broader than measuring isolated logical correctness. It evaluates whether a model can play, adapt, plan, and execute in structured environments.

The benchmark explicitly targets reasoning under several axes at once. It includes deterministic and stochastic tasks, so repeated action in the same state may or may not yield identical outcomes. It includes both single-player and two-player settings, thereby probing not only optimization and deduction but also adversarial response and first-mover effects. It also separates text-only from text-image interaction, with multimodal tasks limited to two-player deterministic puzzles in the current release [2510.06475].

A central design claim is extensibility. Puzzle instances are generated from templates rather than fixed example sets, and difficulty is scaled through puzzle size and randomized generation. This makes PuzzlePlex an adaptive testbed rather than a frozen leaderboard. The paper explicitly presents it as a framework that can generate more challenging instances as frontier models evolve [2510.06475].

## 2. Puzzle families and environment architecture

PuzzlePlex contains 15 puzzle types organized across scenario and dynamics categories. Text-based puzzles cover all four scenario/dynamics combinations, while text-image puzzles are restricted to two-player deterministic settings [2510.06475].

| Category | Count | Puzzles |
|---|---:|---|
| Single-player deterministic | 3 | TidyTower, OptimalTouring, CountMaximalCocktails |
| Two-player deterministic | 5 | SudoKill, CardNim, MaxMaximalCocktails, ExclusivityParticles, Superply |
| Single-player stochastic | 3 | ExclusivityProbes, RubyRisks, Max Target |
| Two-player stochastic | 2 | BeatOrBombSto., LargerTarget |
| Text-image two-player deterministic | 2 | SudoKill M., Superply M. |

The benchmark also annotates the main reasoning mode each puzzle is intended to probe. Examples include logical and spatial reasoning in SudoKill, numerical reasoning in CardNim and OptimalTouring, deduction under uncertainty in ExclusivityProbes and Max Target, strategic competition in CardNim, SudoKill, Superply, BeatOrBombSto., and LargerTarget, and multimodal strategic reasoning in SudoKill M. and Superply M. [2510.06475].

Each game is embedded in a common environment loop with five components: Puzzle Generator, Solver, Transition Checker, State Transition, and Evaluator. The generator maps a puzzle template to an instance, and the generated instance is the initial state:
$$
\text{instance}(p)=S_0.
$$
After a move $M$, the transition module maps state $S_n$ to the next state and feedback,
$$
M : S_n > (S_{n+1}, F_n),
$$
and at termination the evaluator computes a raw score over the trajectory,
$$
rsp = E_p(S_0, S_1, \ldots, S_n).
$$
The implementation also includes a web UI called Simulator for visualizing game states, move histories, model reasoning steps, and state transitions [2510.06475].

## 3. Evaluation protocols and metrics

PuzzlePlex uses two distinct evaluation settings. In the instruction-based setting, the model acts as an interactive agent through natural-language prompts and turn-by-turn state updates. In the code-based setting, the model is given puzzle rules and I/O templates once and must generate executable code that then interacts autonomously with the environment [2510.06475].

The instruction-based protocol covers deterministic puzzles only. For deterministic single-player puzzles, each model is evaluated on 10 randomly generated instances with fixed seeds 1 to 10 at both easy and normal difficulty. For deterministic two-player puzzles, each model pair competes on 5 instances with seeds 1 to 5, each match is repeated twice, and first-player role is alternated to reduce first-mover bias. Stochastic puzzles are excluded from instruction-based evaluation because variance is high and enough runs for robust conclusions would be expensive [2510.06475].

The code-based protocol is more exhaustive. For each puzzle, each model is sampled 32 times to generate code. Deterministic puzzles use the same evaluation structure as in instruction-based testing. For single-player stochastic puzzles, each generated program is evaluated over 100 runs with seeds 1 to 100 at both difficulty levels. For two-player stochastic puzzles, each program plays 50 runs with seeds 1 to 50 and alternating roles [2510.06475].

The scoring system combines raw score, normalization, Elo-style comparison, and termination-status analysis. For two-player puzzles, raw score is win \(=1\), loss \(=0\), tie \(=0.5\). For single-player puzzles, raw scores may be binary or continuous. Single-player results are normalized into \([0,1]\): if higher is better, the top-performing model gets 1 and others receive \(\text{score}/\text{max}\); if lower is better, normalization uses \(\text{min}/\text{score}\). To compare models across both single-player and two-player tasks, PuzzlePlex uses an Elo system with initial rating
$$
R=1000,
$$
update rule
$$
R'_A = R_A + K \cdot (S_A - E_A),
$$
with
$$
K=32,
$$
and expected score
$$
E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}}.
$$
For single-player games, pairwise Elo comparisons are induced by normalized scores [2510.06475].

The benchmark also reports detailed termination-status categories: LEGAL, RULE VIOLATION, NOT FOLLOWING INSTRUCTION, TIMEOUT, SYNTAX ERROR, and RUNTIME ERROR. These are diagnostically important because they distinguish strategic weakness from failures of parsing, code generation, or execution reliability. In both evaluation settings, malformed output is handled by allowing up to five attempts per turn to produce correctly formatted output using a regex checker; if all fail, the model loses the game [2510.06475].

## 4. Empirical findings

The central empirical result is that reasoning models outperform non-reasoning models in the instruction-based setting. On deterministic puzzles, the top average normalized scores are DeepSeek-R1 at 0.62, o4-mini at 0.59, Gemini-2.5-pro at 0.58, QwQ-32B at 0.58, and grok-3-mini at 0.49, with DeepSeek-V3 and GPT-4.1 around 0.42–0.43. The paper highlights that all top-5 models in this regime use reasoning strategies, and that QwQ-32B outperforms larger non-reasoning systems such as GPT-4.1 and DeepSeek-V3 [2510.06475].

PuzzlePlex also reports that open-source models are highly competitive. DeepSeek-R1 achieves the highest average normalized score, 0.62, surpassing the best proprietary model listed, Gemini-2.5-pro at 0.58. At the same time, algorithmic baselines remain stronger in aggregate: the custom strategy averages 0.70, above the best models overall, although the paper notes that strong LLMs can match or exceed the custom strategy in some two-player deterministic puzzles [2510.06475].

Code-based evaluation is substantially harder. DeepSeek-R1, for example, drops on single-player deterministic puzzles from 0.64 to 0.33 on easy and from 0.48 to 0.25 on normal when moving from instruction-based to code-based evaluation. Yet the ranking also changes: GPT-4.1, which is not among the strongest instruction-based models, enters the top tier in code generation. The main code-based average normalized scores reported are o4-mini at 0.53, DeepSeek-R1 at 0.52, GPT-4.1 at 0.51, and Gemini-2.5-pro at 0.50, while best-of-32 sampling raises some of these substantially, with best scores of 0.73, 0.66, 0.72, and 0.74 respectively [2510.06475].

Error analysis clarifies why the two settings differ. In instruction-based mode, legal completion rates are often around 0.79 for DeepSeek-R1 and o4-mini, 0.78 for QwQ-32B, but only 0.37 for Gemma-3-27B and 0.42 for Phi-4-multimodal. In code-based mode, legal rates fall to 0.61 for o4-mini, 0.57 for GPT-4.1, 0.54 for DeepSeek-R1, 0.26 for DeepSeek-V3, and 0.00 for Phi-4-multimodal. Syntax errors, runtime errors, timeouts, and failures to produce compliant code account for a large share of the performance drop [2510.06475].

Puzzle-level and prompting analyses reveal additional structure. TidyTower is particularly difficult in instruction mode, with many models scoring 0.00 in the base setting. Legality-aware prompting consistently helps, especially for o4-mini, suggesting that many failures arise from weak internal legality filtering rather than from entirely absent strategic reasoning. On TidyTower and SudoKill, 1-shot prompting yields little improvement, Tree-of-Thought helps TidyTower but not SudoKill, and removing history dramatically improves TidyTower, which the authors interpret as evidence that models can be misled by their own prior reasoning traces [2510.06475].

PuzzlePlex also studies token-based scaling. DeepSeek-R1’s normalized score tends to improve as token count increases, whereas DeepSeek-V3 shows a flat or negative trend. DeepSeek-R1 also allocates more tokens to normal than easy puzzles, indicating difficulty-aware test-time computation. In the multimodal setting, stronger systems usually benefit from image input; the paper gives GPT-4.1’s improvement of \(+0.38\) on SUPERPLYM (Normal) as a representative example [2510.06475].

## 5. Relation to adjacent puzzle-solving research

PuzzlePlex is primarily a benchmark rather than a solver paper, but its design sits in a larger methodological landscape. In Wordle, “Constraint Satisfaction Approaches to Wordle: Novel Heuristics and Cross-Lexicon Validation” formalizes the puzzle as a CSP
$$
\mathcal{W}=\langle X,D,C\rangle
$$
with monotone domain reduction, cardinality-aware duplicate-letter handling, and propagation-aware action scoring. Its CSP-Aware Entropy solver reaches 3.54 average guesses with 99.9% success on 2,315 English words, and the paper explicitly frames this layered reasoning architecture as relevant to PuzzlePlex-style systems that compare solvers across structured puzzle domains [2510.02855].

Visual planning papers define another nearby regime. “Alphazzle: Jigsaw Puzzle Solver with Deep Monte-Carlo Tree Search” casts square-tile image reassembly as a deterministic single-player planning problem in which neural policy and value networks guide PUCT-based MCTS; with fine-tuning, \(Q(a \mid s_t)\) action choice, and 10 attempts, it reports 75.12% patch-wise, 77.54% neighbor-wise, and 51.49% puzzle-wise accuracy on \(3\times 3\) puzzles [2302.00384]. Polygonal reassembly work broadens this space further. “Pictorial and apictorial polygonal jigsaw puzzles: The lazy caterer model, properties, and solvers” combines noisy geometric predicates, hierarchical loop constraints, and a spring-mass dynamical system for convex polygonal crossing-cuts puzzles [2008.07644], while “Solving Convex Partition Visual Jigsaw Puzzles” uses geometric filtering, pictorial compatibility, greedy cycle-based aggregation, and Box2D-based spring-mass pose recovery on a 75-puzzle benchmark [2511.04450]. These papers indicate a family of structured assembly tasks that PuzzlePlex does not presently enumerate but that are technically close to its reasoning-and-planning agenda.

The benchmark’s scope can also be interpreted against reconfiguration and communication-oriented puzzle research. “Higher-dimensional cubical sliding puzzles” gives exact solvability thresholds \(S(d,k)\) and a phase structure ranging from isolated to connected mobile regimes on hypercubic state spaces [2307.14143]. “AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation” isolates communication under information asymmetry, with search space \((N!)^2\), turn budget \(2N\), and 100% success across all feedback modes on 5-piece puzzles for GPT-5 and Claude-4.0 [2512.03466]. Together, such work suggests a broader design space in which PuzzlePlex can be read as one node among several benchmark paradigms: formal CSP reasoning, visual assembly planning, structured reconfiguration, and cooperative inference under partial observability.

## 6. Limitations, extensibility, and computational context

PuzzlePlex’s current release has several explicit limitations. It contains only 15 puzzles, and the authors note that results may depend on puzzle mix and random seeds. Rapid model evolution can date the comparison set, and the benchmark does not study fine-tuned or task-specialized systems. Stochastic instruction-based evaluation is omitted because variance is high and enough runs for robust conclusions would be expensive. Some normalized scoring depends on best or worst achievable outcomes under fixed initialization, and output-format compliance materially influences outcomes because parsing and code-generation failures can terminate runs independently of strategic quality [2510.06475].

Its extensibility claim is nonetheless substantial. Because puzzles are template-generated and difficulty-scaled, the framework is designed to admit harder instances as models improve. A plausible implication is that PuzzlePlex’s long-term value lies less in a fixed ranking than in its capacity to expose changing failure modes across interaction styles, especially the contrast between turn-by-turn reasoning and reusable code synthesis [2510.06475].

The broader puzzle literature situates that claim in a well-developed computational context. “Tetravex is NP-complete” establishes NP-completeness for exact edge-matching tilings [0903.1147]. “Gourds: a sliding-block puzzle with turning” separates universally solvable reconfiguration on proper hole-free boards from NP-complete colored placement with unbounded colors [2011.00968]. “Solving Tantrix via Integer Programming” shows that integer programming with placement variables, edge-color constraints, additional short-subloop exclusions, and an artificial compactness objective solves Tantrix challenge numbers up to 50 [1202.6438]. “Computational Complexity of Games and Puzzles” provides the reduction-based toolkit—decision formulations, gadget constructions, and NP/PSPACE-style classifications—used to analyze such puzzle families formally [1807.04724]. In that setting, PuzzlePlex functions not as a replacement for solver theory or complexity analysis, but as an empirical benchmark that tests how contemporary foundation models behave across puzzle classes that the broader literature already understands through CSPs, search, optimization, and hardness reductions.

Source: https://www.emergentmind.com/topics/puzzleplex