Skip to content
AI Atlas
PaperActive

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

arxiv.org/abs/2609.09143

quality59

Updated 5 min ago · first seen 11 Sept 2026

paper_01M29ANNTSTA5J06TZ1PH49Z7T

Published
8 Sept 2026
T2 · 3 h ago
arXiv
2609.09143
T2 · 3 h ago

Abstract

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.

Authors 5

Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen, Yang Liu

Specification

arXiv id
2609.09143

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium

Github repo
amazon-far/Tokenizer_UMM

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium

Hf paper url
https://huggingface.co/papers/2609.09143

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium

Github stars
0

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 5 min agomedium

Hf comments
2

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 5 min agomedium

Upvotes
3

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 5 min agomedium

Published
8 Sept 2026

Source:Hugging Face Hub (public pages, model cards, papers)T2observed 3 h agomedium

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

11

Source tiers

T211

Freshest observation

5 min ago

Conflicts

None