Methodology

How the measuring works

Every reconstruction is scored against a reference model in five dimensions. Metric and environment are described here precisely enough to be traceable – without publishing the scoring code.

In brief: Every AI reconstruction is compared with the correct model — by dimensions, design features, body shape and its own check drawing. For each figure there is a pass threshold, and per case a stated ceiling on what is measurable at all. Unfamiliar terms are explained in the glossary.

The five dimensions

Each dimension measures a different facet of “correct”. The threshold is the value from which a dimension counts as passed; a case is formally passed only when it meets all applicable thresholds at once.

How the figures relate: Four dimensions are measured for every case — dimensional accuracy, feature fidelity, shape fidelity and drawing fidelity. Thread fidelity is the fifth dimension, but applicable only to thread cases; that is why it is absent from the main results table and scored case by case. The delivery rate is not an evaluation dimension but a quality figure in its own right: it counts whether a scorable model arrived at all. Missing deliveries do not lower the means — so always read both figures side by side.

The five evaluation dimensions arranged around the comparison of reconstruction and reference
Five perspectives on the same comparison: reconstruction against reference.

Dimensional accuracy

Threshold ≥ 0.95

How many of the drawing's dimensions reappear in the model? The share of the target dimensions specified in the drawing that can be found again on the finished model by measurement. The dimensions are re-measured on the finished body, not taken over from the model parameters: a spectrum of measurable quantities is extracted from the geometry (edge lengths, distances between parallel faces, diameters, radii, angles), and every drawing value must appear within it in tolerance. Not every drawing dimension is measurable this way: dimensions on axes or single points — such as the edge distance of a hole axis — have no measurable pair of faces in the solid model. The maximum reachable per case is therefore stated as the reference coverage on each case, as a concrete fraction (e.g. 14 of 17 dimensions).

Feature fidelity

Threshold ≥ 0.90

Do the dimensions sit on the right design feature? Agreement of the design structure: the right features in the right places (holes, chamfers, grooves, add-ons, breakthroughs) – not just the right outer dimensions. A part can be correct in its dimensions and still structurally wrong; this dimension separates the two.

Shape fidelity

Threshold ≥ 0.90

How congruent is the body with the reference (volume overlap)? The ratio of intersection to union volume (“Intersection over Union”) between the candidate and reference body after best-possible alignment. The strictest overall shape test – it penalizes both too much and too little material and is safeguarded against the silent pitfalls of boolean volume measurement (cross-check plus independent point sampling).

Thread fidelity

Threshold = 1.0

Are threads correct in size and position? Applicable only to thread cases: is the thread modeled to standard by type and pitch and equal in volume? Here it is “all or nothing” – hence the threshold of exactly 1.0. Cases without a thread carry no value here and are not scored.

Drawing fidelity

Threshold ≥ 0.90

Does the model pass the check against a technical drawing of its own? Self-consistency of the technical drawing generated from the reconstruction against the target dimensions: does the candidate produce a drawing whose measured dimensions match those required? This dimension tests the documenting capability – the ability to represent one's own model to standard and true to dimension.

Important — self-consistency, not a proof of shape: because the reconstruction is checked here against its own drawing, even a wrongly built shape can achieve perfect drawing fidelity (measured: a run with shape fidelity 0.423 achieved drawing fidelity 1.000). Distinction from dimensional accuracy: dimensional accuracy measures the finished body against the dimensions of the source drawing; drawing fidelity measures the projection generated from the reconstruction. Whether the shape is correct is stated by shape fidelity alone.

For those recomputing — the internal names: dimensional accuracy = dimensionRecall · feature fidelity = semanticMatch · shape fidelity = shapeIoU · thread fidelity = threadMatch · drawing fidelity = drawingMatch · reference coverage = referenceCoverage · drawing self-test = drawingSelfTest · random-hit rate = randomHitRate. Under these names the values appear in the machine-readable report data of the results page.

Per-case ceilings

Not every case even permits a perfect value. Each case therefore carries three calibrated ceilings, which are tracked on the Cases page and taken into account in the results.

Ceiling for dimensional accuracy

Reference coverage

The share of the drawing's dimensions that can be found again by measurement on the reference model at all. If it lies below 1, even a perfect reconstruction cannot reach a dimensional accuracy of 1 – the value caps the reachable score. Example from our own bearing block case: 14 of 17 dimensions are measurable (0.82), because three dimensions refer to hole axes or a single point — a geometrically exact reconstruction reaches exactly 0.82 there, and the case cards state this maximum as well.

Consequence for passing — worked through: “Formally passed” compares the raw value with the threshold of 0.95. A case with 16 drawing dimensions, of which 13 are measurable, has the ceiling 13/16 = 0.813 — even a perfect reconstruction cannot formally pass it; this affects 16 of the 26 cases. That is why the results state two rates — the headline figure is passed (reachable): each dimension above the threshold or at its case-specific maximum (in the example: 13 of 13 measurable dimensions hit). Formal (raw) remains beside it as a technical comparison figure. If the reconstruction misses one of the 13 measurable dimensions, it stays a reconstruction error in both rates.

Ceiling for drawing fidelity

Drawing self-test

The drawing fidelity of the reference drawing against itself. It quantifies how far the drawing check can get in the best case, and separates model errors from limits of the projection/measurement chain.

Floor (null model)

Random-hit rate

How often a purely random value would hit a dimension – separately for lengths and angles. This rate makes visible how much dimensional accuracy would arise from chance alone, and prevents a high value from being taken for skill where it is statistics.

Isolation protocol

The agent works with open aids – which makes it all the more important that the solution itself remains out of reach. Isolation is therefore not a side step but a hard entry condition for every measurement.

Contrast: main environment with references and version history, isolated working environment of the agent with only the drawing, tools and rulebook
Before every run the working environment is cleared of every trace of the solution – and the empty finding is proven.

Working copy without version history

Every run receives a fresh working copy fully detached from the project's version history. The agent cannot reach reference data through earlier states.

Truth removal in the file system

All reference artifacts – drawing, reference model (FCStd/STEP) and every analysis or result file derived from them with a case reference – are physically removed from the working tree before the run. The deletion list is extended with every new kind of artifact.

grep proof

After removal, a full-text scan of the working tree proves that no trace of the truth remains. A non-empty hit aborts the run – the empty grep is the entry ticket.

Evidence-free edition of the rulebook

The rulebook given to the agent contains the rules but not the solution-revealing evidence. A rule such as “when in doubt, air not material” stays – the concrete dimension from which it was once distilled is taken out, because it would be solution knowledge.

Cleared agent memory

Every agent starts with a minimal profile: no project-bound memory, no working notes loaded in from the development environment. Case-free base settings are recorded as a documented exception in the isolation proof.

Reproducibility: the metric fingerprint

For a comparison to hold across weeks and models, it must be established that all figures arose with the same evaluation standard. To that end every evaluation carries a metric fingerprint – a content-stable hash over the scoring logic and its parameters. If anything about the standard changes, the fingerprint changes; if it stays the same, the figures are comparable bit for bit. The fingerprint seals the standard the way a lead seal secures a measuring instrument: not hidden, but protected against unnoticed change.

The currently valid stamp is d32c33768f4d3e52. It stays untouched across the campaign; thresholds are never lowered. That way a movement in the figures is always a model or environment effect – never a shifted standard.

A parallel checksum safeguards the drawing scoring chain (drawing fingerprint). Both stamps appear in the results on every evaluation.

Who measures here? The author and person professionally responsible is Dipl.-Ing. Sören Gebbert, managing director of the Institut für holistische Technologieforschung GmbH (operator of this website). Measurement environment, reference models, thresholds and evaluation all come from one hand — which is why this lab is a documented in-house study, not an independently audited benchmark.

Limits of what the figures can say

Four points determine how far the figures reach.

  • Open aids. The task is open: drawing plus documented environment. The test does not measure rote knowledge but the implementation of visible specifications – with all the freedoms a real drawing leaves.
  • Single-run variance. Every agent run is stochastic. The shape fidelity of a case scatters between runs – for individual cases by several tenths – without anything having changed about the model. That is why we measure the top systems several times and state the ranges; a single case figure without context cannot be interpreted.
  • Rule-shaped cases. Seven of the cases are anchoring cases – the agents' rules of thumb were sharpened on them. Their result is optimistically biased and is therefore stated separately from the remaining independent cases. The clean generalization test is the independent cases. Chosen for anchoring were precisely the cases on which early runs failed — that these harder parts lie below the independent ones on average does not contradict the bias warning.
  • Only one stage. What is measured is drawing → model → self-documentation. FEM and manufacturing are later stages and not included here.
To the cases