All figures arose with the same evaluation standard, kept unchanged across the entire campaign. Green means: the dimension's threshold was met. Seven models across eleven runs over 26 cases – the top systems claude-opus-5 and gpt-5.6-sol with three runs each as a variance indicator.
In brief: Seven AI models, eleven runs, 26 drawings. The best system hits the body shape almost every time — yet even its runs still often fail to meet all requirements at once. The tables show both. Unfamiliar terms are explained in the glossary.
Open the case-by-case comparison — every model in 3D next to the original
Sorted descending by shape fidelity; one row per agent-model system — the top systems measured three times enter with the mean of their runs, and the individual runs appear under repeatability and in the case-by-case comparison. A click on a dimension header sorts by that figure. The delivered column shows how many of the 26 corpus cases yielded a scorable model. Passing is stated in two rates. The headline figure is “passed (reachable)”: each dimension above the threshold or at its case-specific maximum. The technical comparison figure “formal (raw)” requires all raw values above the thresholds — which 16 of the 26 cases cannot achieve by construction because of their dimensional-accuracy ceiling; the explanation with a worked example is below the table.
Figures at a glance: The table shows the four dimensions measured per case (dimensional accuracy, feature fidelity, shape fidelity, drawing fidelity). Thread fidelity applies only to thread cases and is therefore absent here; the delivery rate is not a dimension but counts whether a scorable model arrived. Details: Methodology — the five dimensions.
One bar per model for each dimension; the vertical line marks the pass threshold.
Seven of the cases are anchoring cases: the agents' rules of thumb were sharpened on them, so their result is optimistically biased. The clean generalization test is the remaining independent cases. Both subsets are stated separately – a mean over everything would blur the two. That the anchoring cases nonetheless score worse on average is no contradiction: they were chosen for anchoring precisely because early runs failed on them — they are the harder parts. The bias warning concerns the closeness to training (training = test), not the absolute score level.
Agent runs are stochastic. For the two top systems measured three times, this table shows how strongly the shape fidelity of a case scatters between runs – the noise band without which a model comparison at case level cannot be interpreted. Notably: the aggregates are stable across three runs (claude-opus-5 0.918–0.954; gpt-5.6-sol 0.668–0.691), while individual cases still jump by several tenths.
Who wins: By the leading dimension shape fidelity, Claude Code with claude-opus-5 is clearly ahead (shape fidelity 0.935 as the mean of three runs, individual runs 0.918–0.954; 78/78 delivered); gpt-5.6-sol reaches the best dimensional accuracy and drawing fidelity in individual runs, but falls off on shape (≈0.68).
What this means in practice: The best system reconstructs the bodies of these tasks almost consistently congruent and passes 20 of 26 cases on average under the reachable dimensions (best run: 21) — with visible scatter between runs of the same model.
What it does not mean: Engineering work is not thereby replaced. Tolerances, material, simulation and manufacturing are not part of the measurement; review and sign-off remain with a qualified engineer.
What exactly is compared: Each row is an agent-model system — a command-line agent (Claude Code or codex) together with its language model. In this setup the influence of the model cannot be separated from the influence of the agent.
The case-by-case comparison shows, for each of the 26 cases, the original source video next to the interactive 3D views of all eleven runs — including the fully shown bearing block demo case of our own.
To the case-by-case comparison