Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation
Updated 4 h ago · first seen 11 Sept 2026
paper_01M294H26HQ4TECDXB8WMNA8TQ
- Published
- 11 Sept 2026
- T1 · 4 h ago
- arXiv
- 2609.11507
- T1 · 4 h ago
- Category
- cs.CV
- T1 · 4 h ago
Abstract
Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Diffusion Transformers (DiTs). We found that certain attention blocks naturally form an Intrinsic Spatial Grounding Map (ISGM) that precisely locates reference subjects. Building on this insight, we propose Dual-phase Intrinsic Attention Leveraging (DIAL), a framework that uses these internal signals for both training and inference. In low-noise stages, we use ISGM to guide the attention mechanism, allowing precise control over fidelity strength during inference without retraining. In high-noise stages, we use these same maps to automatically build preference pairs at no additional cost for Reinforcement Learning (RL). This RL procedure effectively anchors the model's attention to reference subjects and mitigates semantic drift. Extensive experiments show that DIAL significantly outperforms baseline models on the OpenS2V-Eval benchmark, consistently improving identity consistency and enabling controllable fidelity strength.
Authors 8
Niange Yu, Ye Tian, Biaolong Chen, Miao Lu, Aixi Zhang, Hao Jiang, Yunhai Tong, Pipei Huang
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- arXiv id
- 2609.11507
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Categories
- cs.CV
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Primary category
- cs.CV
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 4 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
4 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Niange Yu, Ye Tian, Biaolong Chen
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Published
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 11 Sept 2026 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperHarnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation
New paper: Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CV | feed | T1· Official | 2 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.