Region-aware test-time scaling for compositional image generation
File(s) ECCV_2026 (1).pdf (22.85 MB)
Accepted version
Author(s)
Type
Conference Paper
Abstract
Test-time scaling (TTS) has emerged as a promising paradigm for improving the performance of large-scale models. However, existing vision TTS methods largely rely on global exploration—e.g., resampling noise or rewriting prompts—and often struggle to efficiently search the vast compositional space. As a result, they can exhibit a “scaling plateau,” where additional computation yields diminishing returns in semantic alignment. In this paper, we propose Region-Aware Scaling (RAS), a framework that bridges region-aware generation and test-time scaling. By treating regional decomposition as a powerful and previously overlooked scaling axis, RAS converts complex compositional prompts into coordinated regional sub-tasks, effectively increasing the density of valid candidates during inference. At its core, RAS builds on a training-free Region-Aware Generation (RAG). Unlike many layout-based methods that incur substantial overhead, RAG injects regional guidance only during early denoising, enabling precise attribute binding while preserving global structural coherence. We evaluate RAS on the GenEval benchmark and observe consistent improvements in scaling efficiency across diverse compositional challenges. RAS achieves an overall score of 0.85 with only 2 samples, matching a 32-sample noise-scaling baseline; with 4 samples, it reaches 0.88, surpassing the combination of 32-sample noise and prompt scaling. Overall, our results suggest that structuring the search space via regional decomposition provides a principled and computationally efficient direction for scaling compositional alignment.
Date Acceptance
2026-06-18
Citation
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
ISSN
0302-9743
Publisher
Springer
Journal / Book Title
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Copyright Statement
Subject to copyright. This paper is embargoed until publication. Once published the author’s accepted manuscript will be made available under a CC-BY License in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy).
License URL
Source
The 19th European Conference on Computer Vision -- ECCV 2026
Publication Status
Accepted
Start Date
2026-09-08
Finish Date
2026-09-12
Coverage Spatial
Malmö, Sweden
