Papers
Topics
Authors
Recent
Search
2000 character limit reached

PuzzlePlex: Benchmark for Foundation Models

Updated 14 July 2026
  • PuzzlePlex is a benchmark that uses a suite of 15 interactive puzzles to assess reasoning, planning, and generalization in foundation models.
  • It features diverse puzzle types across deterministic/stochastic and single/two-player settings, including multimodal challenges.
  • The framework employs both instruction-based and code-based evaluation protocols to measure strategic adaptation and error handling in complex scenarios.

PuzzlePlex is a benchmark for evaluating the reasoning, planning, and generalization abilities of foundation models through a curated suite of interactive puzzles. Introduced in “PUZZLEPLEX: Benchmarking Foundation Models on Reasoning and Planning with Puzzles” (Long et al., 7 Oct 2025), it comprises 15 rule-based puzzle types spanning single-player and two-player settings, deterministic and stochastic dynamics, text-only and text-image modalities, multiple difficulty levels, and two evaluation modes—instruction-based interaction and code-based execution. Rather than functioning as a static question-answer dataset, it is structured as an extensible environment for studying long-horizon problem solving, dynamic adaptation, strategic interaction, and executable program synthesis.

1. Rationale and benchmark scope

PuzzlePlex was introduced to address a gap in existing evaluations of foundation models: many prior puzzle benchmarks emphasize single-player tasks, short-horizon interactions, mostly deterministic settings, little or no competition, limited multimodality, no code-execution-based evaluation, and puzzle families that may already be common in model pretraining data (Long et al., 7 Oct 2025). Its stated aim is therefore broader than measuring isolated logical correctness. It evaluates whether a model can play, adapt, plan, and execute in structured environments.

The benchmark explicitly targets reasoning under several axes at once. It includes deterministic and stochastic tasks, so repeated action in the same state may or may not yield identical outcomes. It includes both single-player and two-player settings, thereby probing not only optimization and deduction but also adversarial response and first-mover effects. It also separates text-only from text-image interaction, with multimodal tasks limited to two-player deterministic puzzles in the current release (Long et al., 7 Oct 2025).

A central design claim is extensibility. Puzzle instances are generated from templates rather than fixed example sets, and difficulty is scaled through puzzle size and randomized generation. This makes PuzzlePlex an adaptive testbed rather than a frozen leaderboard. The paper explicitly presents it as a framework that can generate more challenging instances as frontier models evolve (Long et al., 7 Oct 2025).

2. Puzzle families and environment architecture

PuzzlePlex contains 15 puzzle types organized across scenario and dynamics categories. Text-based puzzles cover all four scenario/dynamics combinations, while text-image puzzles are restricted to two-player deterministic settings (Long et al., 7 Oct 2025).

Category Count Puzzles
Single-player deterministic 3 TidyTower, OptimalTouring, CountMaximalCocktails
Two-player deterministic 5 SudoKill, CardNim, MaxMaximalCocktails, ExclusivityParticles, Superply
Single-player stochastic 3 ExclusivityProbes, RubyRisks, Max Target
Two-player stochastic 2 BeatOrBombSto., LargerTarget
Text-image two-player deterministic 2 SudoKill M., Superply M.

The benchmark also annotates the main reasoning mode each puzzle is intended to probe. Examples include logical and spatial reasoning in SudoKill, numerical reasoning in CardNim and OptimalTouring, deduction under uncertainty in ExclusivityProbes and Max Target, strategic competition in CardNim, SudoKill, Superply, BeatOrBombSto., and LargerTarget, and multimodal strategic reasoning in SudoKill M. and Superply M. (Long et al., 7 Oct 2025).

Each game is embedded in a common environment loop with five components: Puzzle Generator, Solver, Transition Checker, State Transition, and Evaluator. The generator maps a puzzle template to an instance, and the generated instance is the initial state:

instance(p)=S0.\text{instance}(p)=S_0.

After a move MM, the transition module maps state SnS_n to the next state and feedback,

M:Sn>(Sn+1,Fn),M : S_n > (S_{n+1}, F_n),

and at termination the evaluator computes a raw score over the trajectory,

rsp=Ep(S0,S1,…,Sn).rsp = E_p(S_0, S_1, \ldots, S_n).

The implementation also includes a web UI called Simulator for visualizing game states, move histories, model reasoning steps, and state transitions (Long et al., 7 Oct 2025).

3. Evaluation protocols and metrics

PuzzlePlex uses two distinct evaluation settings. In the instruction-based setting, the model acts as an interactive agent through natural-language prompts and turn-by-turn state updates. In the code-based setting, the model is given puzzle rules and I/O templates once and must generate executable code that then interacts autonomously with the environment (Long et al., 7 Oct 2025).

The instruction-based protocol covers deterministic puzzles only. For deterministic single-player puzzles, each model is evaluated on 10 randomly generated instances with fixed seeds 1 to 10 at both easy and normal difficulty. For deterministic two-player puzzles, each model pair competes on 5 instances with seeds 1 to 5, each match is repeated twice, and first-player role is alternated to reduce first-mover bias. Stochastic puzzles are excluded from instruction-based evaluation because variance is high and enough runs for robust conclusions would be expensive (Long et al., 7 Oct 2025).

The code-based protocol is more exhaustive. For each puzzle, each model is sampled 32 times to generate code. Deterministic puzzles use the same evaluation structure as in instruction-based testing. For single-player stochastic puzzles, each generated program is evaluated over 100 runs with seeds 1 to 100 at both difficulty levels. For two-player stochastic puzzles, each program plays 50 runs with seeds 1 to 50 and alternating roles (Long et al., 7 Oct 2025).

The scoring system combines raw score, normalization, Elo-style comparison, and termination-status analysis. For two-player puzzles, raw score is win =1=1, loss =0=0, tie =0.5=0.5. For single-player puzzles, raw scores may be binary or continuous. Single-player results are normalized into [0,1][0,1]: if higher is better, the top-performing model gets 1 and others receive score/max\text{score}/\text{max}; if lower is better, normalization uses MM0. To compare models across both single-player and two-player tasks, PuzzlePlex uses an Elo system with initial rating

MM1

update rule

MM2

with

MM3

and expected score

MM4

For single-player games, pairwise Elo comparisons are induced by normalized scores (Long et al., 7 Oct 2025).

The benchmark also reports detailed termination-status categories: LEGAL, RULE VIOLATION, NOT FOLLOWING INSTRUCTION, TIMEOUT, SYNTAX ERROR, and RUNTIME ERROR. These are diagnostically important because they distinguish strategic weakness from failures of parsing, code generation, or execution reliability. In both evaluation settings, malformed output is handled by allowing up to five attempts per turn to produce correctly formatted output using a regex checker; if all fail, the model loses the game (Long et al., 7 Oct 2025).

4. Empirical findings

The central empirical result is that reasoning models outperform non-reasoning models in the instruction-based setting. On deterministic puzzles, the top average normalized scores are DeepSeek-R1 at 0.62, o4-mini at 0.59, Gemini-2.5-pro at 0.58, QwQ-32B at 0.58, and grok-3-mini at 0.49, with DeepSeek-V3 and GPT-4.1 around 0.42–0.43. The paper highlights that all top-5 models in this regime use reasoning strategies, and that QwQ-32B outperforms larger non-reasoning systems such as GPT-4.1 and DeepSeek-V3 (Long et al., 7 Oct 2025).

PuzzlePlex also reports that open-source models are highly competitive. DeepSeek-R1 achieves the highest average normalized score, 0.62, surpassing the best proprietary model listed, Gemini-2.5-pro at 0.58. At the same time, algorithmic baselines remain stronger in aggregate: the custom strategy averages 0.70, above the best models overall, although the paper notes that strong LLMs can match or exceed the custom strategy in some two-player deterministic puzzles (Long et al., 7 Oct 2025).

Code-based evaluation is substantially harder. DeepSeek-R1, for example, drops on single-player deterministic puzzles from 0.64 to 0.33 on easy and from 0.48 to 0.25 on normal when moving from instruction-based to code-based evaluation. Yet the ranking also changes: GPT-4.1, which is not among the strongest instruction-based models, enters the top tier in code generation. The main code-based average normalized scores reported are o4-mini at 0.53, DeepSeek-R1 at 0.52, GPT-4.1 at 0.51, and Gemini-2.5-pro at 0.50, while best-of-32 sampling raises some of these substantially, with best scores of 0.73, 0.66, 0.72, and 0.74 respectively (Long et al., 7 Oct 2025).

Error analysis clarifies why the two settings differ. In instruction-based mode, legal completion rates are often around 0.79 for DeepSeek-R1 and o4-mini, 0.78 for QwQ-32B, but only 0.37 for Gemma-3-27B and 0.42 for Phi-4-multimodal. In code-based mode, legal rates fall to 0.61 for o4-mini, 0.57 for GPT-4.1, 0.54 for DeepSeek-R1, 0.26 for DeepSeek-V3, and 0.00 for Phi-4-multimodal. Syntax errors, runtime errors, timeouts, and failures to produce compliant code account for a large share of the performance drop (Long et al., 7 Oct 2025).

Puzzle-level and prompting analyses reveal additional structure. TidyTower is particularly difficult in instruction mode, with many models scoring 0.00 in the base setting. Legality-aware prompting consistently helps, especially for o4-mini, suggesting that many failures arise from weak internal legality filtering rather than from entirely absent strategic reasoning. On TidyTower and SudoKill, 1-shot prompting yields little improvement, Tree-of-Thought helps TidyTower but not SudoKill, and removing history dramatically improves TidyTower, which the authors interpret as evidence that models can be misled by their own prior reasoning traces (Long et al., 7 Oct 2025).

PuzzlePlex also studies token-based scaling. DeepSeek-R1’s normalized score tends to improve as token count increases, whereas DeepSeek-V3 shows a flat or negative trend. DeepSeek-R1 also allocates more tokens to normal than easy puzzles, indicating difficulty-aware test-time computation. In the multimodal setting, stronger systems usually benefit from image input; the paper gives GPT-4.1’s improvement of MM5 on SUPERPLYM (Normal) as a representative example (Long et al., 7 Oct 2025).

5. Relation to adjacent puzzle-solving research

PuzzlePlex is primarily a benchmark rather than a solver paper, but its design sits in a larger methodological landscape. In Wordle, “Constraint Satisfaction Approaches to Wordle: Novel Heuristics and Cross-Lexicon Validation” formalizes the puzzle as a CSP

MM6

with monotone domain reduction, cardinality-aware duplicate-letter handling, and propagation-aware action scoring. Its CSP-Aware Entropy solver reaches 3.54 average guesses with 99.9% success on 2,315 English words, and the paper explicitly frames this layered reasoning architecture as relevant to PuzzlePlex-style systems that compare solvers across structured puzzle domains (Arafat et al., 3 Oct 2025).

Visual planning papers define another nearby regime. “Alphazzle: Jigsaw Puzzle Solver with Deep Monte-Carlo Tree Search” casts square-tile image reassembly as a deterministic single-player planning problem in which neural policy and value networks guide PUCT-based MCTS; with fine-tuning, MM7 action choice, and 10 attempts, it reports 75.12% patch-wise, 77.54% neighbor-wise, and 51.49% puzzle-wise accuracy on MM8 puzzles (Paumard et al., 2023). Polygonal reassembly work broadens this space further. “Pictorial and apictorial polygonal jigsaw puzzles: The lazy caterer model, properties, and solvers” combines noisy geometric predicates, hierarchical loop constraints, and a spring-mass dynamical system for convex polygonal crossing-cuts puzzles (Harel et al., 2020), while “Solving Convex Partition Visual Jigsaw Puzzles” uses geometric filtering, pictorial compatibility, greedy cycle-based aggregation, and Box2D-based spring-mass pose recovery on a 75-puzzle benchmark (Ohayon et al., 6 Nov 2025). These papers indicate a family of structured assembly tasks that PuzzlePlex does not presently enumerate but that are technically close to its reasoning-and-planning agenda.

The benchmark’s scope can also be interpreted against reconfiguration and communication-oriented puzzle research. “Higher-dimensional cubical sliding puzzles” gives exact solvability thresholds MM9 and a phase structure ranging from isolated to connected mobile regimes on hypercubic state spaces (Beyer et al., 2023). “AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation” isolates communication under information asymmetry, with search space SnS_n0, turn budget SnS_n1, and 100% success across all feedback modes on 5-piece puzzles for GPT-5 and Claude-4.0 (Cadet et al., 3 Dec 2025). Together, such work suggests a broader design space in which PuzzlePlex can be read as one node among several benchmark paradigms: formal CSP reasoning, visual assembly planning, structured reconfiguration, and cooperative inference under partial observability.

6. Limitations, extensibility, and computational context

PuzzlePlex’s current release has several explicit limitations. It contains only 15 puzzles, and the authors note that results may depend on puzzle mix and random seeds. Rapid model evolution can date the comparison set, and the benchmark does not study fine-tuned or task-specialized systems. Stochastic instruction-based evaluation is omitted because variance is high and enough runs for robust conclusions would be expensive. Some normalized scoring depends on best or worst achievable outcomes under fixed initialization, and output-format compliance materially influences outcomes because parsing and code-generation failures can terminate runs independently of strategic quality (Long et al., 7 Oct 2025).

Its extensibility claim is nonetheless substantial. Because puzzles are template-generated and difficulty-scaled, the framework is designed to admit harder instances as models improve. A plausible implication is that PuzzlePlex’s long-term value lies less in a fixed ranking than in its capacity to expose changing failure modes across interaction styles, especially the contrast between turn-by-turn reasoning and reusable code synthesis (Long et al., 7 Oct 2025).

The broader puzzle literature situates that claim in a well-developed computational context. “Tetravex is NP-complete” establishes NP-completeness for exact edge-matching tilings (0903.1147). “Gourds: a sliding-block puzzle with turning” separates universally solvable reconfiguration on proper hole-free boards from NP-complete colored placement with unbounded colors (Hamersma et al., 2020). “Solving Tantrix via Integer Programming” shows that integer programming with placement variables, edge-color constraints, additional short-subloop exclusions, and an artificial compactness objective solves Tantrix challenge numbers up to 50 (Kino et al., 2012). “Computational Complexity of Games and Puzzles” provides the reduction-based toolkit—decision formulations, gadget constructions, and NP/PSPACE-style classifications—used to analyze such puzzle families formally (Costa, 2018). In that setting, PuzzlePlex functions not as a replacement for solver theory or complexity analysis, but as an empirical benchmark that tests how contemporary foundation models behave across puzzle classes that the broader literature already understands through CSPs, search, optimization, and hardness reductions.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PuzzlePlex.