DiTs Unlock Precise Regional Image Control
New 'appearance pointers' enable Diffusion Transformers to precisely control image generation regionally, matching SOTA performance without retraining.

Visual TL;DR
creative professionals struggle with precise regional control over image generation
From the articleCreative professionals grapple with the limitations of text-only prompting for precise image generation.
Diffusion Transformers could not direct tokens to specific spatial influences previously
From the article 2 mentionsAchieving granular control over specific regions, dictating materials, identities, and spatial layouts, remains a significant hurdle.
novel compact tokens guide DiTs to apply specific appearance cues spatially
From the article 3 mentionsA novel solution emerges from researchers who have introduced appearance pointers, compact tokens designed to guide DiTs.
powers the mechanism, aligning text/image inputs with user-defined masks
From the articleThis mechanism is powered by a region correspondence network and refined through spatial aggregation, allowing for multiple regional descriptions without a substantial increase in computational load.
DiTs can now precisely control image generation regionally, dictating materials and identities
From the articleThis mechanism is powered by a region correspondence network and refined through spatial aggregation, allowing for multiple regional descriptions without a substantial increase in computational load.
first solution that works with both text and image inputs for regional control
From the articleThis innovation represents the first modality-agnostic interface for localized multimodal control within a DiT, crucially without requiring a full model retraining from scratch.
achieves state-of-the-art results without requiring model retraining
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.