Algorithms under matched compute
Each run gets 1,500 evaluations through one metered evaluator; CP-SAT, which does not produce evaluations internally, gets 6.4 s — the median time NSGA-II needs for that budget. Means are over briefs after averaging seeds; intervals are 95% bootstrap intervals over briefs.
Hypotheses
Frozen before the main run; paired over briefs; p-values from a sign-flip permutation test, Holm-corrected within each hypothesis.
| Test | Mean difference | 95% CI | p (Holm) | briefs |
|---|---|---|---|---|
| H0 optimisation beats random search · Supported | ||||
| hypervolume: NSGA-II − random search | 0.061 | 0.047–0.076 | 0.000 | 72 |
| hypervolume: Simulated annealing − random search | 0.040 | 0.030–0.050 | 0.000 | 72 |
| H1 CP-SAT has the highest solve rate · Not supported | ||||
| solve rate: CP-SAT − Heuristic-seeded NSGA-II | -0.120 | -0.241–0.000 | 0.064 | 72 |
| H2 NSGA-II finds broader fronts than simulated annealing · Supported | ||||
| hypervolume: NSGA-II − simulated annealing | 0.022 | 0.013–0.031 | 0.000 | 72 |
| spread: NSGA-II − simulated annealing | 0.017 | 0.009–0.025 | 0.000 | 72 |
| H3, H4 LLM proposals and repair · Not run no LLM credentials for this release | ||||
| H5 no method dominates · Supported winners: density → Heuristic · daylight → Random search · privacy → NSGA-II · circulation → CP-SAT · carbon → Heuristic · cost → Heuristic · runtime → Heuristic | ||||
By difficulty band
| method | easy | medium | hard |
|---|---|---|---|
| Heuristic | 67% | 58% | 33% |
| Random search | 29% | 11% | 13% |
| Simulated annealing | 51% | 40% | 35% |
| NSGA-II | 58% | 60% | 42% |
| CP-SAT | 53% | 50% | 33% |
| Random sketch + repair | 57% | 40% | 36% |
| Heuristic-seeded NSGA-II | 69% | 58% | 44% |
By typology
| method | adu pair | courtyard | duplex | fourplex | rowhouse | triplex |
|---|---|---|---|---|---|---|
| Heuristic | 33% | 67% | 100% | 50% | 42% | 25% |
| Random search | 33% | 0% | 72% | 0% | 0% | 0% |
| Simulated annealing | 72% | 11% | 100% | 0% | 28% | 42% |
| NSGA-II | 78% | 19% | 97% | 3% | 69% | 53% |
| CP-SAT | 92% | 0% | 100% | 17% | 6% | 58% |
| Random sketch + repair | 69% | 17% | 89% | 6% | 50% | 36% |
| Heuristic-seeded NSGA-II | 69% | 75% | 89% | 22% | 33% | 56% |
Convergence
Archive hypervolume against evaluations, mean over all runs. CP-SAT reports its solutions at once, so its curve is flat.
Does more compute help?
Separate sweep on one balanced cycle of 18 briefs, two seeds.
| Method | 250 ev | 500 ev | 1000 ev | 2000 ev | 4000 ev |
|---|---|---|---|---|---|
| Simulated annealing | 0.010 11% | 0.016 22% | 0.040 33% | 0.095 56% | 0.138 78% |
| NSGA-II | 0.020 22% | 0.033 28% | 0.068 50% | 0.100 61% | 0.134 69% |
| Random search | 0.010 11% | 0.016 17% | 0.020 19% | 0.029 28% | 0.037 28% |
Cell: mean hypervolume, and solve rate beneath.
CP-SAT with more time
All 72 briefs, seed 0.
| Time | Role | Solve | HV |
|---|---|---|---|
| 3.2 s | pre-registered robustness (half) | 42% | 0.054 |
| 6.4 s | matched budget | 49% | 0.065 |
| 12.8 s | exploratory | 53% | 0.070 |
| 25.6 s | exploratory | 61% | 0.077 |
Priority-regime sensitivity
For each named weighting, the best weighted score each method's feasible archive offers, mean over briefs. Top method per regime: balanced → Heuristic; density → Heuristic; daylight → Random search; carbon → Heuristic; privacy → Random sketch + repair; cost → Heuristic. The top method changes with the weighting.
| Method | balanced | density | daylight | carbon | privacy | cost |
|---|---|---|---|---|---|---|
| Heuristic | 0.733 #1 | 0.805 #1 | 0.690 #3 | 0.764 #1 | 0.753 #5 | 0.699 #1 |
| Random search | 0.668 #7 | 0.705 #6 | 0.694 #1 | 0.708 #7 | 0.747 #6 | 0.573 #7 |
| Simulated annealing | 0.701 #3 | 0.756 #4 | 0.685 #5 | 0.740 #3 | 0.768 #2 | 0.633 #3 |
| NSGA-II | 0.698 #5 | 0.748 #5 | 0.686 #4 | 0.737 #4 | 0.767 #3 | 0.632 #4 |
| CP-SAT | 0.700 #4 | 0.771 #3 | 0.693 #2 | 0.726 #5 | 0.753 #4 | 0.600 #6 |
| Random sketch + repair | 0.711 #2 | 0.779 #2 | 0.681 #6 | 0.748 #2 | 0.775 #1 | 0.659 #2 |
| Heuristic-seeded NSGA-II | 0.671 #6 | 0.705 #7 | 0.634 #7 | 0.712 #6 | 0.721 #7 | 0.616 #5 |
Robustness checks
- Hypervolume ranking vs reference point 0.1
- Kendall τ 1.00
- Hypervolume ranking vs reference point 0.2
- Kendall τ 1.00
- Best-carbon method unchanged under ±50% coefficients
- 100% of 14 perturbations
CP-SAT outcomes by typology
| Typology | model infeasible | solved | timeout |
|---|---|---|---|
| adu pair | 3 | 33 | 0 |
| courtyard | 0 | 0 | 36 |
| duplex | 0 | 36 | 0 |
| fourplex | 6 | 6 | 24 |
| rowhouse | 0 | 2 | 34 |
| triplex | 12 | 21 | 3 |
“model infeasible” means the CP encoding, which is stricter in places than the checker, has no solution; it is not a proof the brief is infeasible.