standard tier · 72 briefs · 1,368 runs · results db2e24e

CityFormGen

Generative and optimisation methods for middle housing, compared under matched compute

RESEARCH GENERATIVE-DESIGN BENCHMARK - NOT CONSTRUCTION DOCUMENTATION OR CODE CERTIFICATION.

Seventy-two lots, seven methods, one checker.

Every method gets the same seeded middle-housing briefs, the same 17 benchmark constraints and the same six proxy objectives, under a matched compute budget. What they find is compared as a trade-off surface, never as a single score.

Plan drawing of a adu pair on a 16.5 by 19.9 metre lot, 2 storeys, 2 dwellings, 13 rooms. Dwelling U0 (2BR) has 6 rooms totalling 101 square metres. Dwelling U1 (1BR) has 5 rooms totalling 101 square metres. Building footprint 116 square metres; open space 218 square metres. This design satisfies every benchmark hard constraint. Objective scores, higher is better: density 0.97, daylight 0.65, privacy 0.93, circulation 0.56, carbon 0.80, cost 0.67.
B0000 · knee of the pooled frontRandom sketch + repair✓ all 17 constraints met

The brief with the most methods finding a feasible design (seven). Inspect this plan · see the whole brief.

62/72
briefs where at least one method found a feasible design
57%
highest solve rate: Heuristic-seeded NSGA-II
0.084
highest mean hypervolume: Heuristic
1,368
runs: 72 briefs × 3 seeds × 7 methods, 1,500 evaluations each

Pre-registered hypotheses

  1. H0Optimisation beats random search SupportedNSGA-II adds 0.061 hypervolume over random search and simulated annealing 0.040, paired over briefs.
  2. H1Constraint programming is the most feasible Not supportedCP-SAT solves 45% of runs; Heuristic-seeded NSGA-II solves 57%. CP-SAT timed out on 97 of 216 runs.
  3. H2Evolutionary search finds broader fronts SupportedNSGA-II beats simulated annealing on hypervolume (+0.022) and spread (+0.017).
  4. H3LLM proposals: better adjacency, worse geometry Not runNot run: no LLM credentials were configured for this release.
  5. H4Repair recovers LLM feasibility at a cost Not runNot run: no LLM credentials were configured for this release.
  6. H5No method dominates SupportedThe best method differs across the six objectives and runtime (4 different winners).

Statistics, intervals and tests

No method leads everywhere

Solve rate by typology: the share of runs that found at least one design meeting every constraint. The method that is best on duplexes is not the method that is best on courtyards.

Mean over briefs of the per-brief solve rate across seeds.
methodadu paircourtyardduplexfourplexrowhousetriplex
Heuristic0.330.671.000.500.420.25
NSGA-II0.780.190.970.030.690.53
Random sketch + repair0.690.170.890.060.500.36
CP-SAT0.920.001.000.170.060.58
Heuristic-seeded NSGA-II0.690.750.890.220.330.56
Simulated annealing0.720.111.000.000.280.42
Random search0.330.000.720.000.000.00

Read the benchmark