LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Published 16 Sept 2026arXiv:2609.13287
Updated 12 h ago · first seen 9 Sept 2026
paper_01M2HN8YAKDED22HZFWSRB6QMQ
Abstract
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question. We present LLaDA-UI, a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent. LLaDA-UI follows a two-stage training pipeline: general multimodal pre-training aligns a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, followed by GUI-agent supervised fine-tuning on diverse mobile, desktop, web, and grounding data. Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks. These results establish block-wise diffusion as a practical generative paradigm for multimodal GUI agents.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxivLLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxivLLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxivLLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents: published at changed from 2026-09-09T00:00:00+00:00 to 2026-09-15T04:00:00+00:00
Published9 Sept 2026→15 Sept 2026arxiv
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.