Skip to content
AI Atlas
PaperActive

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

arxiv.org/abs/2609.09143

quality59

Updated 2 h ago · first seen 11 Sept 2026

paper_01M29ANNTSTA5J06TZ1PH49Z7T

Published
8 Sept 2026
T2 · 5 h ago
arXiv
2609.09143
T2 · 5 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history

13 claims · 11 properties

Official pageofficial_url1

Claim history for Official page
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/abs/2609.09143currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

arXiv idarxiv_id1

Claim history for arXiv id
ValueValid from → toStatusSourceConfidenceExtractor
2609.09143currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Authorsauthors1

Claim history for Authors
ValueValid from → toStatusSourceConfidenceExtractor
Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen, Yang LiucurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github repogithub_repo1

Claim history for Github repo
ValueValid from → toStatusSourceConfidenceExtractor
amazon-far/Tokenizer_UMMcurrentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf paper urlhf_paper_url1

Claim history for Hf paper url
ValueValid from → toStatusSourceConfidenceExtractor
https://huggingface.co/papers/2609.09143currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Github starsmetric.github_stars1

Claim history for Github stars
ValueValid from → toStatusSourceConfidenceExtractor
0currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Hf commentsmetric.hf_comments2

Claim history for Hf comments
ValueValid from → toStatusSourceConfidenceExtractor
2currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
1supersededHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Upvotesmetric.upvotes2

Claim history for Upvotes
ValueValid from → toStatusSourceConfidenceExtractor
3currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic
1supersededHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

PDFpdf_url1

Claim history for PDF
ValueValid from → toStatusSourceConfidenceExtractor
https://arxiv.org/pdf/2609.09143currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Publishedpublished_at1

Claim history for Published
ValueValid from → toStatusSourceConfidenceExtractor
8 Sept 2026currentcurrentHugging Face Hub (public pages, model cards, papers)T2mediumdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →