The lab does not only measure models – it improves the environment in which they work. The following story from v1 to v2 shows both: what a targeted intervention delivers and what price a carelessly phrased one demands for it.
Same 25 cases, same model, same isolation, same order. Only the environment was changed — as one package of two coupled components: the extended rulebook in an evidence-free edition and the projection check as an obligation – after the build, produce a technical drawing of your own, measure target against actual, and only then deliver. A movement in the numbers is therefore attributable primarily to the package, not to a shifted setup — a residue of single-run chance remains; which of the two components contributes how much is something this comparison deliberately does not separate.
Two confounders remain and are reported: the single-run variance (every run is stochastic) and training = test on the seven anchoring cases on which the rules were sharpened. That is why the attribution below does not rest on the number alone, but on the agents' self-documentation – wherever an agent traces its own decision back to a rule.
In brief: We improved the agents' rules of thumb once in a targeted way and re-measured everything: on the practice cases it got markedly better, on new cases slightly worse. Both stand here, including the cause of the regression. Unfamiliar terms are explained by the glossary.
All 25 self-documentations in v2 contained a real section on the projection check with measured target-actual lines. The check grounds every delivery in re-measured values; its time cost per run was not recorded separately.
In the aggregate across the 25 cases, shape fidelity rose from 0.689 to 0.727 and drawing fidelity from 0.947 to 0.962; dimensional accuracy stayed practically unchanged (0.843 → 0.834). The formal pass rate – all thresholds at once – climbed from 1 to 4 of 25, and the number of clear shape errors (shape fidelity < 0.40) fell from 7 to 4.
What matters is the look at the subsets: on the 7 anchoring cases shape fidelity jumped from 0.241 to 0.516 (+0.28) – the intervention works strongly there. On the 18 independent cases of the corpus at the time it fell slightly from 0.863 to 0.809 (−0.05). Four of the seven target cases were lifted over the 0.4 mark, two of them to nearly perfect values; on the independent cases shape fidelity fell by 0.05 at the same time. Both figures are part of the result.
The most expensive regression hit a single case – and is documented verbatim in its self-documentation. A manufacturing heuristic along the lines of “when in doubt, air rather than material” was applied to a zone for which the source visibly shows material. The heuristic overruled a correct standard interpretation.
One case thereby fell in shape fidelity from 0.917 (v1, correctly built full) to 0.119 (v2). The classic error of an unconditional blanket rule: phrased for undimensioned gaps, applied to a spot where the projection shows material edges.
The lesson is not a new rule, but a conditioning of the existing one: a heuristic must never overrule an interpretation that the projection of the source supports – evidence from the projection beats the heuristic. The same run showed two further limits: the projection check captures spans, contour types and directions, but it does not count repeated features (one case got the right direction but too few of the same kind of element), and wall thicknesses of rotational parts need to be explicitly cross-checked as a Ø difference.
“When in doubt, air” applies only to undimensioned gaps and only where the source shows no closed material contours. Ranking made explicit: evidence before heuristic.
Count repeated features per view (claws, ribs, windows, holes) and check against your own projection – not only dimensions, but counts.
Every environment change runs as a paired comparison. Every case with a shape-fidelity drop > 0.1 is attributed via its self-documentation: a rule reference = a problem in the rulebook, no reference = a variance candidate.
At least one repeat run per condition, to know the noise band of the single-case values. Without it, no ±0.2 movement at case level is interpretable.
The answer to the opening question “Can the rulebook be improved?” is therefore: Yes, with conditions. Targeted rules repair the error classes they were built for – but unconditional blanket rules create side effects on cases that were previously correct. The next step is therefore not a new rule, but the conditioning of the existing one – and the regression check as a standing process.
To the current measurement matrix