Results

The measurement matrix

All figures arose with the same evaluation standard, kept unchanged across the entire campaign. Green means: the dimension's threshold was met. Seven models across eleven runs over 26 cases – the top systems claude-opus-5 and gpt-5.6-sol with three runs each as a variance indicator.

In brief: Seven AI models, eleven runs, 26 drawings. The best system hits the body shape almost every time — yet even its runs still often fail to meet all requirements at once. The tables show both. Unfamiliar terms are explained in the glossary.

Open the case-by-case comparison — every model in 3D next to the original

Aggregate per model

Sorted descending by shape fidelity; one row per agent-model system — the top systems measured three times enter with the mean of their runs, and the individual runs appear under repeatability and in the case-by-case comparison. A click on a dimension header sorts by that figure. The delivered column shows how many of the 26 corpus cases yielded a scorable model. Passing is stated in two rates. The headline figure is “passed (reachable)”: each dimension above the threshold or at its case-specific maximum. The technical comparison figure “formal (raw)” requires all raw values above the thresholds — which 16 of the 26 cases cannot achieve by construction because of their dimensional-accuracy ceiling; the explanation with a worked example is below the table.

Figures at a glance: The table shows the four dimensions measured per case (dimensional accuracy, feature fidelity, shape fidelity, drawing fidelity). Thread fidelity applies only to thread cases and is therefore absent here; the delivery rate is not a dimension but counts whether a scorable model arrived. Details: Methodology — the five dimensions.

Loading aggregates …

Dimensions compared

One bar per model for each dimension; the vertical line marks the pass threshold.

Loading charts …

Anchored and independent cases

Seven of the cases are anchoring cases: the agents' rules of thumb were sharpened on them, so their result is optimistically biased. The clean generalization test is the remaining independent cases. Both subsets are stated separately – a mean over everything would blur the two. That the anchoring cases nonetheless score worse on average is no contradiction: they were chosen for anchoring precisely because early runs failed on them — they are the harder parts. The bias warning concerns the closeness to training (training = test), not the absolute score level.

Loading subsets …

Repeatability (n = 3)

Agent runs are stochastic. For the two top systems measured three times, this table shows how strongly the shape fidelity of a case scatters between runs – the noise band without which a model comparison at case level cannot be interpreted. Notably: the aggregates are stable across three runs (claude-opus-5 0.918–0.954; gpt-5.6-sol 0.668–0.691), while individual cases still jump by several tenths.

Loading repeatability …

What follows from the matrix

Who wins: By the leading dimension shape fidelity, Claude Code with claude-opus-5 is clearly ahead (shape fidelity 0.935 as the mean of three runs, individual runs 0.918–0.954; 78/78 delivered); gpt-5.6-sol reaches the best dimensional accuracy and drawing fidelity in individual runs, but falls off on shape (≈0.68).

What this means in practice: The best system reconstructs the bodies of these tasks almost consistently congruent and passes 20 of 26 cases on average under the reachable dimensions (best run: 21) — with visible scatter between runs of the same model.

What it does not mean: Engineering work is not thereby replaced. Tolerances, material, simulation and manufacturing are not part of the measurement; review and sign-off remain with a qualified engineer.

What exactly is compared: Each row is an agent-model system — a command-line agent (Claude Code or codex) together with its language model. In this setup the influence of the model cannot be separated from the influence of the agent.

Check every case yourself

The case-by-case comparison shows, for each of the 26 cases, the original source video next to the interactive 3D views of all eleven runs — including the fully shown bearing block demo case of our own.

To the case-by-case comparison