Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Updated 1 h ago · first seen 11 Sept 2026
paper_01M29ANNTSTA5J06TZ1PH49Z7T
- Published
- 8 Sept 2026
- T2 · 4 h ago
- arXiv
- 2609.09143
- T2 · 4 h ago
Abstract
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.
Authors 5
Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen, Yang Liu
Specification
- Official page
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium
- arXiv id
- 2609.09143
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium
- Github repo
- amazon-far/Tokenizer_UMM
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium
- Hf paper url
- https://huggingface.co/papers/2609.09143
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium
- Github stars
- 0
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 1 h agomedium
- Hf comments
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 1 h agomedium
- Upvotes
- 3
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 1 h agomedium
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium
- Published
- 8 Sept 2026
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 4 h agomedium
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
11
Source tiers
T211
Freshest observation
1 h ago
Conflicts
None
No models linked to this paper yet.
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Github stars
Github starsmetric.github_stars1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 0 | → current | current | Hugging Face Hub (public pages, model cards, papers)T2 | medium | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
New paper: Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
huggingface
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 1 h ago | 4 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.