MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation
Updated 2 h ago · first seen 11 Sept 2026
paper_01M294H31M65DG1T4D5JFYC91C
- Published
- 11 Sept 2026
- T1 · 2 h ago
- arXiv
- 2511.13135
- T1 · 2 h ago
- Category
- cs.CV
- T1 · 2 h ago
Abstract
Medical vision-language models (VLMs) are increasingly expected to support clinical workflows through diagnostic text and relevant medical images. However, current medical visual benchmarks have three recurring limitations: query-image misalignment from queries weakly grounded in specific image instances, closed-ended formats that narrow answer space and encourage shortcut-based prediction, and text-centric output paradigms that limit evaluation of image-generation and image-editing capabilities. We introduce MedGEN-Bench, a benchmark for open-ended multimodal medical generation. The evaluation snapshot reported in this manuscript comprises 6,422 image-text pairs reviewed by clinical experts and models, spanning 6 canonical imaging modalities, 15 clinical tasks, and 27 named subtasks. It includes 1,100 Visual Question Answering (VQA) pairs, 3,872 Image Editing pairs, and 1,450 Contextual Multimodal Generation pairs. MedGEN-Bench centers on contextual entanglement: dependence of an instruction's intended output on the particular image instance rather than on task wording alone. The benchmark operationalizes this concept through image-grounded instructions and extends evaluation to open-ended multimodal outputs. Its tiered evaluation protocol combines reproducible reference-based fidelity and similarity measures with a structured, checklist-guided assessment by a medical VLM judge. We evaluate 10 compositional frameworks, 2 dedicated image-editing models, 3 unified models, and 5 VLMs. The results show image-output tasks remain unsaturated. Contextual augmentation increases mean image-instruction similarity from 0.273 to 0.372, while a 1,000-case medical-expert audit shows moderate agreement between judge scores and clinician ratings. Source code and dataset are available at https://yangjj007.github.io/medgen.
Authors 15
Junjie Yang, Yuhao Yan, Gang Wu, Rui Qian, Zhisheng Chen, Haijiang Li, Yuhe Wu, Qichao Zhao, Dawen Tian, Xiang Wan, Fenglei Fan, Wenjian Qin, Yongquan Zhang, Feiwei Qin, Changmiao Wang
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- arXiv id
- 2511.13135
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Categories
- cs.CV
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Primary category
- cs.CV
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 2 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
2 h ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Junjie Yang, Yuhao Yan, Gang Wu
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Authors
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperMedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation
New paper: MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CV | feed | T1· Official | 2 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.