MechBench: Mechanical Benchmark Suites
- MechBench is a collection of benchmark datasets covering CAD assembly motion prediction, visual mechanical reasoning, and FEM-based crashworthiness optimization.
- The CAD benchmark uses segmented point clouds, procedural generation, and ground-truth SE(3) transforms to evaluate dependency-aware motion in synthetic gear assemblies.
- The VLM benchmark features over 550 image–question pairs for zero-shot evaluation, while the optimization suite uses expensive FEM crash simulations for black-box optimization tasks.
MechBench denotes multiple recent benchmark resources centered on mechanical structure, mechanical reasoning, and mechanically grounded optimization rather than a single unified artifact. In current arXiv usage, the name refers to a benchmark dataset of 693 synthetic gear assemblies with part-wise ground-truth motion trajectories introduced alongside DYNAMO for articulated assembly motion prediction (Patel et al., 15 Sep 2025), an image-based benchmark of 155 cognitive experiments for testing mechanical reasoning in vision–LLMs (Sun et al., 2024), and MECHBench, an open-source suite of black-box optimization problems rooted in structural-mechanics crashworthiness simulations (Rodríguez et al., 13 Nov 2025). The shared label masks substantial differences in modality, supervision, and evaluation target.
1. Nomenclature and domain of use
The term appears in three distinct technical settings.
| Resource | Primary data or problem unit | Target task |
|---|---|---|
| MechBench in DYNAMO | 693 synthetic gear assemblies | Per-part SE(3) motion trajectory prediction from segmented CAD point clouds |
| MechBench in CogDevelop2K | 155 cognitive experiments, yielding 550+ image–question pairs | Zero-shot evaluation of mechanical reasoning in 26 VLMs |
| MECHBench | Three crashworthiness benchmark cases | Black-box optimization on expensive structural-mechanics simulations |
The first MechBench is a 3D perception and design-automation benchmark in which motion arises from geometric coupling through meshing teeth or aligned axes rather than predefined joints. The second is an image-based assay of pure mechanical reasoning spanning system stability, pulleys, gears, levers, inertia, and fluids. The third is a benchmarking suite for derivative-free optimization built around explicit-FEM crash simulations and exposed through a Python/IOH-style interface (Patel et al., 15 Sep 2025, Sun et al., 2024, Rodríguez et al., 13 Nov 2025).
A common source of confusion is that the same name is attached to resources with incompatible data models. One uses segmented point clouds, rigid transforms, twists, and coupling matrices; one uses schematic or lightly rendered images paired with multiple-choice or binary questions; one uses parameterized design variables, mesh generation, solver calls, and post-processing. The capitalization MECHBench is also associated specifically with the structural-mechanics optimization suite.
2. MechBench as a benchmark for coupled mechanical motion in CAD assemblies
In "DYNAMO: Dependency-Aware Deep Learning Framework for Articulated Assembly Motion Prediction" (Patel et al., 15 Sep 2025), MechBench is defined as a benchmark dataset of 693 diverse synthetic gear assemblies with part-wise ground-truth motion trajectories. Each assembly contains between 2 and 7 parts, for a total of 2,445 movable parts. The dataset covers 13 gear categories, including spur gears, helical gears, double-helical gears, bevel gears, worm gears, rack-and-pinion, planetary systems, and compound configurations such as multi-stage gear trains. The representation includes a STEP file capturing full CAD hierarchy and exact solid geometry, together with part-segmented point clouds in which each moving part is uniformly sampled with up to surface points and stored as separate point clouds .
The assemblies are procedurally generated in FreeCAD by uniformly sampling key design parameters: gear type, number of teeth, pitch diameter, pressure angle, and tooth profile. The number of parts is uniformly sampled between 2 and 7, with checks to ensure valid meshing, including no interpenetration, correct tooth alignment, and compatible modules. Procedural rules further enforce that no two gears share the same axis unless explicitly intended, as in planetary carriers, and that contacts are enforced by distance thresholds on tooth surfaces.
Motion trajectories are generated by selecting one part as the driver and assigning it a constant angular increment per frame. The rotations of coupled parts follow simple gear-ratio laws: if gear meshes with gear ,
where and are the numbers of teeth on the two gears. The sign encodes direction, clockwise versus counterclockwise. Each sequence is recorded over frames, described as one full cycle, although the same motion can be repeated arbitrarily.
The benchmark is positioned as a structured setting for studying coupled motion in which part dynamics are induced by contact and transmission rather than predefined joints. This differs from prior work on everyday articulated objects such as doors and laptops that typically assumes simplified kinematic structures or relies on joint annotations. In that formulation, MechBench is not merely a geometry corpus; it is an evaluation environment for relational motion induced by meshing and transmission.
3. Ground-truth structure, file organization, and evaluation protocol in the CAD benchmark
The MechBench dataset accompanying DYNAMO stores per-frame rigid transforms for part 0 at frame 1 as 2. Each transform is built from a pure rotation about the part’s center 3, and the dataset also stores a 6D Lie-algebra twist vector 4,
5
with 6 recovered by the matrix exponential 7 (Patel et al., 15 Sep 2025).
Mechanical coupling is encoded by a binary matrix 8 with
9
Here 0 and 1 are thresholds used to detect mechanical coupling. For evaluation only, the dataset also records mobility units
2
where 3, 4 is the axis, 5 the center of motion, and 6 the active degrees of freedom.
The assembly-size distribution is explicitly reported: 187 assemblies with 7, 214 with 8, 127 with 9, 98 with 0, 38 with 1, and 29 with 2. The train/test split is 90%/10%, approximately 624 training assemblies and 69 test assemblies, with validation performed by holding out approximately 10% of train for hyperparameter tuning. Diversity is described as covering 13 gear categories, a broad range of tooth counts such as 12–80 teeth, diameters of 10 mm–200 mm, straight versus helical profiles, and mixed assemblies such as spur + worm + bevel.
The on-disk organization is defined at the assembly level. Each sample contains meshes/ with one .step file per assembly, clouds/ with segmented point clouds saved as NumPy arrays of shape 3, transforms/ with arrays of shape 4, twists/ with arrays of shape 5, coupling.npy storing the binary matrix 6, and mobility_units.npy of shape 7 storing the tuple encoded as integer type, axis, center, and degrees of freedom.
Evaluation is reported on the test set, with errors computed per part per frame and then averaged over parts and frames. Rotation error uses the best-fit rotation 8 obtained via Kabsch from predicted point clouds, with
9
and frame-wise error
0
Translation error is the mean per-point Euclidean distance between predicted and ground-truth transformed point sets. Temporal consistency is reported as standard deviation across frames for rotation and translation errors, with lower standard deviation indicating a more stable model. Dependency-aware performance is analyzed by breaking down error versus the number of coupled neighbors via 1, and the paper reports consistent DYNAMO performance even for 2 up to 7.
The training objectives are also explicit:
3
4
5
Within that formulation, MechBench serves simultaneously as a data source, an annotation standard, and a motion-prediction evaluation protocol.
4. MechBench as an image-based benchmark for mechanical reasoning in VLMs
In "Probing Mechanical Reasoning in Large Vision LLMs" (Sun et al., 2024), MechBench is described as part of the CogDevelop2K suite and comprises 155 distinct cognitive experiments designed to probe six core facets of mechanical intuition. The six categories are System Stability with 30 experiments, Pulley Systems with 30, Gear Systems with 25, Seesaw-Like Levers with 25, Inertia & Motion with 25, and Fluid Mechanics with 20. Each experiment presents one or more schematic or lightly rendered photographic stimuli accompanied by a multiple-choice or binary question. For each cognitive experiment, 3–5 variants were generated by randomly perturbing object sizes, positions, surface textures, and related factors, yielding a total of 550+ image–question pairs.
Ground-truth labels were assigned by consensus of four mechanical-engineering and cognitive-science graduate students, and ambiguous items with inter-annotator agreement below 75% were discarded. The stimuli were drawn or adapted from classical developmental-psychology and physics-education experiments and then re-rendered to a uniform schematic style described as monochrome with minimal texture. The benchmark is therefore explicitly designed as an assay of reasoning from visualized mechanics rather than from textual description or from simulation traces.
All 26 VLMs were evaluated in a zero-shot, open-ended generation setting. Each prompt appended the instruction, “Please answer by selecting the option letter (A, B, C, D) directly.” Responses were string-matched to the correct choice. Accuracy was computed as
6
and 95% confidence intervals were reported with the normal approximation
7
using 8. Pairwise differences, including model-versus-human comparisons, were tested using two-proportion 9-tests.
Human baseline performance was measured using 20 undergraduates with freshman-level physics training, each answering a randomized subset of approximately 100 items under a strict 2-minute-per-item time limit. Aggregate human accuracy was 94.7% ± 1.2% overall, with per-category accuracies ranging from 91% on Inertia & Motion to 98% on Gear Systems. Representative model scores reported in the paper include GPT-4v at 68.2% ± 1.4 overall, Gemini-Ultra at 65.7% ± 1.5, Claude-3 at 64.3% ± 1.6, LLaVA-13B at 55.8% ± 1.7, and Qwen-VL-6B at 51.2% ± 1.8. Models uniformly beat chance, approximately 16% for four-way multiple choice, but none approach human benchmarks.
The diagnostic analysis identifies Pulley Systems as the hardest category, with average model accuracy under 30%, and Gear Systems as the easiest, with models at approximately 70% versus humans at approximately 98%. Reported failures include misidentification of movable versus fixed pulleys, systematic inversion of predicted motion, friction confusions, pendulum misperceptions, and missed fluid-efflux details. Parameter-count scaling across the full suite is described as weak, with Pearson 0 and 1, indicating that larger transformer-based VLMs do not reliably gain mechanical-reasoning capability simply by scaling up. Sun et al. attribute these deficits to absent or weak mental simulation of causal chains, analog processing of continuous dynamics, and rich spatial-working memory.
5. MECHBench as a black-box optimization suite from structural mechanics
In "MECHBench: A Set of Black-Box Optimization Benchmarks originated from Structural Mechanics" (Rodríguez et al., 13 Nov 2025), MECHBench is an open-source suite of black-box optimization problems drawn directly from structural-mechanics crashworthiness simulations. Its motivation is to bridge the gap between abstract benchmark suites such as COCO/BBOB and CEC and the non-smooth, highly nonlinear, high-cost models encountered in engineering applications. The suite encapsulates FEM model setup, mesh generation, solver call through OpenRadioss, and post-processing behind a simple Python/IOH-style interface, while remaining compatible with optimizers written in any language through an ask–evaluate–tell loop.
The current implementation contains three single-objective crashworthiness problems, all using shell-element FEM models with the Belytschko-Lin-Tsay formulation, surface-to-surface contact, and Cowper–Symonds strain-rate plasticity for steel or aluminum.
| Case | Optimization formulation | Evaluation cost |
|---|---|---|
| Star-shaped crash box | Maximize 2 subject to 3 mm; reformulated as a piecewise unconstrained minimization | ~300–800 s |
| Three-point bending of layered beam | Minimize 4 subject to 5 mm; reformulated as a piecewise unconstrained minimization | ~70–230 s |
| Long crash tube | Minimize 6 | ~500–1900 s |
The star-shaped crash box is a prismatic, star-shaped shell structure clamped at its base and struck from above by a rigid impactor. For 7, a small number of geometric parameters control the outer profile and, for 8, one uniform thickness variable. For 9, the first four variables fix overall box dimensions and the remaining variables define a piecewise-linear thickness profile. The original constrained formulation is
0
with unconstrained reformulation
1
The three-point bending problem models a five-rib aluminum beam simply supported and loaded by a central impactor. The task is to minimize structural mass while limiting maximum deflection to 50 mm. For low-dimensional settings, rib thicknesses or grouped rib thicknesses serve as variables; for 2, each rib’s thickness profile along its length is controlled by interpolated control points. Its reformulated objective is
3
The long crash tube is a rectangular steel shell tube with ten triggers on its four faces, struck axially by an impactor. Ten triplets 4 describe trigger position, protrusion, and height, with mirroring in lower-dimensional parameterizations and full independence at 5. The objective is unconstrained:
6
MECHBench normalizes all decision variables internally to 7, mapping them to physical bounds. Integration follows a fixed workflow: instantiate Problem(case, d, objectives, path_to_solver), evaluate a proposal by writing the mesh, calling OpenRadioss, and parsing output files, then return the objective to the optimizer. The suite provides Python API components under src/sob/problems.py, fem.py, and solver.py, mesh scripts in lib/, and logging compatible with IOHExperimenter and COCO-like practices. The paper emphasizes that high per-evaluation cost restricts purely global, population-based methods to a few dozen calls unless surrogate or multi-fidelity strategies are used, and it notes planned extensions such as trivial landscape transformations and low/high-fidelity surrogates.
6. Comparative significance and recurrent sources of confusion
Across these three usages, MechBench is consistently associated with mechanical constraints and benchmarked algorithmic behavior, but the underlying scientific questions are different. The DYNAMO benchmark asks whether a model can infer coupled rigid-body motion from static CAD geometry with explicit ground-truth trajectories and coupling structure. The CogDevelop2K benchmark asks whether VLMs can answer visual questions derived from cognitive experiments in mechanics. The structural-mechanics suite asks how derivative-free optimizers behave on explicit-FEM crashworthiness problems with expensive, non-smooth evaluations (Patel et al., 15 Sep 2025, Sun et al., 2024, Rodríguez et al., 13 Nov 2025).
The overlapping nomenclature can obscure differences in supervision. In the CAD benchmark, supervision is dense and kinematic: transforms, twists, and coupling matrices are provided. In the VLM benchmark, supervision is categorical: correct choice labels attached to image-question prompts. In MECHBench for optimization, supervision is replaced by objective evaluation through simulation; the benchmark returns a scalar objective or penalty-defined value rather than labels or trajectories. A plausible implication is that results are not comparable across the three benchmarks except at the level of broad methodological theme.
There is also a difference in how “mechanical reasoning” is operationalized. In the DYNAMO setting, reasoning is embodied in dependency-aware motion prediction over articulated assemblies, where coupled motion is induced by contact and transmission rather than predefined joints. In the VLM setting, reasoning is operationalized as answer accuracy on schematic tasks such as pulley direction, lever balance, or fluid efflux. In MECHBench for optimization, the relevant capability is effective search under nonlinear, constrained, and computationally expensive physics. This suggests that the shared name should be interpreted as a label for benchmarked interaction with mechanical systems, not as a single canonical benchmark with a unified task definition.
For researchers, the practical consequence is terminological precision. “MechBench” in discussions of SE(3) trajectories, segmented CAD point clouds, and DYNAMO refers to the articulated-assembly benchmark. “MechBench” in discussions of zero-shot VLM accuracy, 155 cognitive experiments, and CogDevelop2K refers to the visual reasoning benchmark. “MECHBench” in discussions of OpenRadioss, Gmsh, crash boxes, and ask–evaluate–tell optimization refers to the structural-mechanics benchmark suite.