standard tier · 72 briefs · 1,368 runs · results db2e24e

CityFormGen

Generative and optimisation methods for middle housing, compared under matched compute

RESEARCH GENERATIVE-DESIGN BENCHMARK - NOT CONSTRUCTION DOCUMENTATION OR CODE CERTIFICATION.

Algorithms under matched compute

Each run gets 1,500 evaluations through one metered evaluator; CP-SAT, which does not produce evaluations internally, gets 6.4 s — the median time NSGA-II needs for that budget. Means are over briefs after averaging seeds; intervals are 95% bootstrap intervals over briefs.

Heuristic-seeded NSGA-II57% ±
NSGA-II53% ±
Heuristic53% ±
CP-SAT45% ±
Random sketch + repair44% ±
Simulated annealing42% ±
Random search18% ±
Solve rate: share of runs with at least one feasible design.
Heuristic0.084 ±
NSGA-II0.078 ±
Random sketch + repair0.067 ±
CP-SAT0.063 ±
Heuristic-seeded NSGA-II0.061 ±
Simulated annealing0.056 ±
Random search0.017 ±
Hypervolume of the feasible 6-objective front (origin reference; 0 when nothing feasible).
CP-SAT100.0% ±
Heuristic21.5% ±
Heuristic-seeded NSGA-II11.8% ±
NSGA-II11.4% ±
Random sketch + repair7.4% ±
Simulated annealing6.0% ±
Random search0.1% ±
Share of evaluated designs that were feasible, over runs that evaluated anything (a CP-SAT timeout evaluates nothing and counts only as unsolved).
CP-SAT0.086 ±
NSGA-II0.076 ±
Simulated annealing0.059 ±
Heuristic-seeded NSGA-II0.052 ±
Random sketch + repair0.047 ±
Random search0.032 ±
Heuristic0.023 ±
Spread: mean pairwise distance on the front, in objective space.
NSGA-II6.1 s ±
Heuristic-seeded NSGA-II6.1 s ±
Random sketch + repair5.7 s ±
CP-SAT5.2 s ±
Simulated annealing5.1 s ±
Random search5.0 s ±
Heuristic2.0 s ±
Wall-clock per run, single process.
NSGA-II8.5
Random sketch + repair6.3
Heuristic-seeded NSGA-II4.9
Simulated annealing4.5
CP-SAT3.9
Heuristic0.9
Random search0.8
Distinct room-adjacency partis on a run's front: a method that returns one plan many times scores low.

Hypotheses

Frozen before the main run; paired over briefs; p-values from a sign-flip permutation test, Holm-corrected within each hypothesis.

TestMean difference95% CIp (Holm)briefs
H0 optimisation beats random search · Supported
hypervolume: NSGA-II − random search0.0610.047–0.0760.00072
hypervolume: Simulated annealing − random search0.0400.030–0.0500.00072
H1 CP-SAT has the highest solve rate · Not supported
solve rate: CP-SAT − Heuristic-seeded NSGA-II-0.120-0.241–0.0000.06472
H2 NSGA-II finds broader fronts than simulated annealing · Supported
hypervolume: NSGA-II − simulated annealing0.0220.013–0.0310.00072
spread: NSGA-II − simulated annealing0.0170.009–0.0250.00072
H3, H4 LLM proposals and repair · Not run no LLM credentials for this release
H5 no method dominates · Supported winners: density → Heuristic · daylight → Random search · privacy → NSGA-II · circulation → CP-SAT · carbon → Heuristic · cost → Heuristic · runtime → Heuristic

By difficulty band

Solve rate by band (tightness of programme against buildable volume).
methodeasymediumhard
Heuristic67%58%33%
Random search29%11%13%
Simulated annealing51%40%35%
NSGA-II58%60%42%
CP-SAT53%50%33%
Random sketch + repair57%40%36%
Heuristic-seeded NSGA-II69%58%44%

By typology

Solve rate by typology.
methodadu paircourtyardduplexfourplexrowhousetriplex
Heuristic33%67%100%50%42%25%
Random search33%0%72%0%0%0%
Simulated annealing72%11%100%0%28%42%
NSGA-II78%19%97%3%69%53%
CP-SAT92%0%100%17%6%58%
Random sketch + repair69%17%89%6%50%36%
Heuristic-seeded NSGA-II69%75%89%22%33%56%

Convergence

Archive hypervolume against evaluations, mean over all runs. CP-SAT reports its solutions at once, so its curve is flat.

Heuristic
00.081500 ev
Random search
00.081500 ev
Simulated annealing
00.081500 ev
NSGA-II
00.081500 ev
CP-SAT
00.081500 ev
Random sketch + repair
00.081500 ev
Heuristic-seeded NSGA-II
00.081500 ev

Does more compute help?

Separate sweep on one balanced cycle of 18 briefs, two seeds.

Method250 ev500 ev1000 ev2000 ev4000 ev
Simulated annealing0.010
11%
0.016
22%
0.040
33%
0.095
56%
0.138
78%
NSGA-II0.020
22%
0.033
28%
0.068
50%
0.100
61%
0.134
69%
Random search0.010
11%
0.016
17%
0.020
19%
0.029
28%
0.037
28%

Cell: mean hypervolume, and solve rate beneath.

CP-SAT with more time

All 72 briefs, seed 0.

TimeRoleSolveHV
3.2 spre-registered robustness (half)42%0.054
6.4 smatched budget49%0.065
12.8 sexploratory53%0.070
25.6 sexploratory61%0.077

Priority-regime sensitivity

For each named weighting, the best weighted score each method's feasible archive offers, mean over briefs. Top method per regime: balanced → Heuristic; density → Heuristic; daylight → Random search; carbon → Heuristic; privacy → Random sketch + repair; cost → Heuristic. The top method changes with the weighting.

Methodbalanceddensitydaylightcarbonprivacycost
Heuristic0.733 #10.805 #10.690 #30.764 #10.753 #50.699 #1
Random search0.668 #70.705 #60.694 #10.708 #70.747 #60.573 #7
Simulated annealing0.701 #30.756 #40.685 #50.740 #30.768 #20.633 #3
NSGA-II0.698 #50.748 #50.686 #40.737 #40.767 #30.632 #4
CP-SAT0.700 #40.771 #30.693 #20.726 #50.753 #40.600 #6
Random sketch + repair0.711 #20.779 #20.681 #60.748 #20.775 #10.659 #2
Heuristic-seeded NSGA-II0.671 #60.705 #70.634 #70.712 #60.721 #70.616 #5

Robustness checks

Hypervolume ranking vs reference point 0.1
Kendall τ 1.00
Hypervolume ranking vs reference point 0.2
Kendall τ 1.00
Best-carbon method unchanged under ±50% coefficients
100% of 14 perturbations

CP-SAT outcomes by typology

Typologymodel infeasiblesolvedtimeout
adu pair3330
courtyard0036
duplex0360
fourplex6624
rowhouse0234
triplex12213

“model infeasible” means the CP encoding, which is stricter in places than the checker, has no solution; it is not a proof the brief is infeasible.