World
Global identity, urban semantics, shared visual language.
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually synthesized independently, lacking the correspondence required for a coherent urban world. We present HoloWorld, a unified indoor-outdoor urban world generation framework built on a continuously updated cross-scale world context.
Initializing from a user description, HoloWorld progressively represents and updates the diverse world information, from city-scale planning to individual buildings, allowing generated interiors to maintain explicit correspondence with their associated exterior buildings. Conditioned on the evolving context and previously generated neighboring blocks, HoloWorld autoregressively generates urban exteriors with consistent spatial organization and visual identity across blocks. The generated exterior representations are further grounded in 3D building instances and footprints, enabling building-specific indoor generation with geometry-constrained layouts and inherited appearance characteristics.
To our knowledge, HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world. Extensive experiments demonstrate that HoloWorld achieves superior urban exterior generation performance, improving the average AQS score over the SOTA by 7.68% and obtaining the highest average RDR score, while maintaining strong building-level indoor-outdoor correspondence and cross-block continuity within a unified 3D urban world.

Global identity, urban semantics, shared visual language.
Local organization conditioned on surrounding city regions.
A grounded instance with function, appearance, and footprint.
A corresponding space that belongs to its exterior building.



Table 1: Quantitative comparison of urban exterior generation. All methods receive identical urban descriptions as inputs. GPT-5.5-based and human evaluations are reported separately. The best results are in bold.
| Method | AQS | RDR | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SVC↑ | SRC↑ | MTF↑ | LA↑ | SVC↑ | SRC↑ | MTF↑ | LA↑ | |||||||||
| GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | |
| CityCraft | 7.11 | 7.17 | 7.22 | 6.00 | 4.89 | 5.00 | 5.89 | 5.00 | 23.65 | 19.21 | 20.76 | 19.61 | 19.90 | 14.99 | 17.35 | 15.97 |
| SynCity | 7.70 | 7.17 | 8.20 | 7.83 | 4.90 | 6.33 | 5.30 | 6.50 | 21.95 | 18.27 | 25.69 | 21.72 | 16.97 | 15.12 | 14.91 | 12.00 |
| MajutsuCity | 8.00 | 7.83 | 8.60 | 8.00 | 7.00 | 7.50 | 7.40 | 7.83 | 22.62 | 20.99 | 24.99 | 19.91 | 23.89 | 20.52 | 25.84 | 22.70 |
| Ours | 8.75 | 8.67 | 9.00 | 8.50 | 7.63 | 8.00 | 8.00 | 8.67 | 27.89 | 23.83 | 27.64 | 24.65 | 29.43 | 24.11 | 29.57 | 25.89 |
Table 2: Quantitative evaluation of building-level indoor–outdoor correspondence using GPT-5.5 and human evaluation. The best results are in bold. Dashes denote metrics unavailable for TRELLIS because its independently generated interior is not grounded in the exterior building footprint.
| Method | AQS | RDR | Shape IoU↑ | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Functional↑ | Visual↑ | Spatial↑ | Average↑ | Functional↑ | Visual↑ | ||||||||
| GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | ||
| TRELLIS | 6.23 | 6.30 | 5.88 | 4.85 | — | — | — | — | 19.17 | 20.03 | 1.83 | 10.90 | — |
| Ours | 7.47 | 7.45 | 8.21 | 7.60 | 8.77 | 8.65 | 8.15 | 7.90 | 22.35 | 20.36 | 24.29 | 23.51 | 0.997 |
Table 3: Effect of dynamic world-context localization on building-level indoor–outdoor correspondence under GPT-5.5 and human evaluation. The best results are in bold.
| Configuration | AQS | RDR | Shape IoU↑ | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Functional↑ | Visual↑ | Spatial↑ | Average↑ | Functional↑ | Visual↑ | Spatial↑ | |||||||||
| GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | GPT | Hum. | ||
| Static World Context | 5.75 | 5.88 | 4.58 | 6.00 | 3.25 | 4.75 | 4.53 | 5.54 | 16.37 | 11.75 | 1.40 | 0.33 | 1.40 | 11.09 | 0.670 |
| Full Model | 7.67 | 7.88 | 8.17 | 8.25 | 7.75 | 8.50 | 7.86 | 8.21 | 21.82 | 18.86 | 22.85 | 18.82 | 22.85 | 17.50 | 0.994 |
Table 4: Effect of autoregressive neighborhood conditioning (ANC) on cross-block visual and spatial continuity under GPT-5.5 and human evaluation. The best results are in bold.
| Configuration | Continuity AQS↑ | RDR↑ | ||
|---|---|---|---|---|
| GPT | Hum. | GPT | Hum. | |
| Full Model | 8.25 | 7.75 | 24.04 | 21.97 |
| w/o ANC | 7.25 | 5.38 | 16.66 | 19.07 |








































Click the “Demo” button to explore the interactive HoloWorld demo.
@article{huang2026holoworld,
title={To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation},
author={Huang, Xiaobin and Huang, Zilong and Luo, Yang and Fan, Hongchao and Chen, Yiping and Han, Ting},
journal={arXiv preprint arXiv:2608.05879},
year={2026}
}