Research
Papers
Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages; model links come from model cards citing the paper.
510 papers
| Title | Authors | Organization | Published | Introduces Models (and artifacts) whose model card or documentation cites this paper — inbound described_by relations. | Datasets | Benchmarks | Code |
|---|---|---|---|---|---|---|---|
| Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical TasksarXiv:2505.03380cs.CV | Haonan Wang, Jiaji Mao, Lehan Wang +2 | — | 11 Sept 2026 | — | — | — | — |
| SloMoDeblur: A Large-Scale Smartphone Image Deblurring DatasetarXiv:2506.19445cs.CV | Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder +2 | — | 11 Sept 2026 | — | — | — | — |
| Instance-Aware Algorithm Selection for Maximum Clique via a Dual-Channel Graph Neural ArchitecturearXiv:2508.08005cs.LG | Xiang Li, Shanshan Wang, Chenglong Xiao | — | 11 Sept 2026 | — | — | — | — |
| RAU: Reference-based Anatomical Understanding with Vision Language ModelsarXiv:2509.22404cs.CV | Yiwei Li, Yikang Liu, Jiaqi Guo +2 | — | 11 Sept 2026 | — | — | — | — |
| MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data GenerationarXiv:2510.05124cs.CL | Mingjin Li, Yu Liu, Huayi Liu +2 | — | 11 Sept 2026 | — | — | — | — |
| Generative AI for AnalystsarXiv:2512.19705q-fin.ST | Jian Xue, Qian Zhang, Wu Zhu | — | 11 Sept 2026 | — | — | — | — |
| Meta-RL with Bayesian Linear Task ModelsarXiv:2512.20974cs.LG | Jingyang You, Hanna Kurniawati | — | 11 Sept 2026 | — | — | — | — |
| From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM JudgesarXiv:2601.08654cs.CL | Yihan Hong, Huaiyuan Yao, Bolin Shen +2 | — | 11 Sept 2026 | — | — | — | — |
| Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space VisualizationarXiv:2601.15295cs.HC | Yi Wang, John Joon Young Chung, Melissa Roemmele +2 | — | 11 Sept 2026 | — | — | — | — |
| Toward Learning POMDPs Beyond Full-Rank Actions and State ObservabilityarXiv:2601.18930cs.LG | Seiji Shaw, Travis Manderson, Chad Kessens +1 | — | 11 Sept 2026 | — | — | — | — |
| Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft WristarXiv:2601.19275cs.RO | Tatsuya Kamijo, Mai Nishimura, Nodoka Shibasaki +2 | — | 11 Sept 2026 | — | — | — | — |
| Revisiting the Shape Convention of Transformer Language ModelsarXiv:2602.06471cs.CL | Feng-Ting Liao, Guan-Ting Yi, Tzu-Quan Lin +2 | — | 11 Sept 2026 | — | — | — | — |
| False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UKarXiv:2602.13047cs.CL | Madhurananda Pahar, Caitlin Illingworth, Dorota Braun +2 | — | 11 Sept 2026 | — | — | — | — |
| City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial ModificationarXiv:2602.19326cs.MA | Rui Liu, Steven Jige Quan, Zhong-Ren Peng +2 | — | 11 Sept 2026 | — | — | — | — |
| Spec-Harness: Measuring and Improving Behavioral Adequacy of LLM-Synthesized Formal SpecificationsarXiv:2604.00280cs.SE | Md Rakib Hossain Misu, Iris Ma, Cristina V. Lopes | — | 11 Sept 2026 | — | — | — | — |
| Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement LearningarXiv:2604.10701cs.LG | Zikang Shan, Han Zhong, Liwei Wang +1 | — | 11 Sept 2026 | — | — | — | — |
| Where is the Mind? Persona Vectors and LLM IndividuationarXiv:2604.17031cs.CL | Pierre Beckmann, Patrick Butlin | — | 11 Sept 2026 | — | — | — | — |
| The Biggest Risk of Embodied AI is Governance LagarXiv:2604.21938cs.CY | Shaoshan Liu | — | 11 Sept 2026 | — | — | — | — |
| Dont Just Teach, Explain! A Gamified 20Q Recommender for Cybersecurity EducationarXiv:2604.26964cs.CY | Mary Nusrat, Sarfuddin Bhuiyan, Gahangir Hossain | — | 11 Sept 2026 | — | — | — | — |
| "What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer UsearXiv:2605.00497cs.HC | Shardul Sapkota, Matthew J\"orke, Zane Sabbagh +2 | — | 11 Sept 2026 | — | — | — | — |
| EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsarXiv:2605.13841cs.SD | Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz +2 | — | 11 Sept 2026 | — | — | — | — |
| Complementing reinforcement learning with SFT through logit averaging in the post training of LLMsarXiv:2605.20555cs.LG | Xingwei Gan, Ying Zhu | — | 11 Sept 2026 | — | — | — | — |
| SpecBench: Measuring Reward Hacking in Long-Horizon Coding AgentsarXiv:2605.21384cs.SE | Bingchen Zhao, Dhruv Srikanth, Yuxiang Wu +1 | — | 11 Sept 2026 | — | — | — | — |
| Tracing Computation Density in LLMsarXiv:2605.27033cs.CL | Corentin Kervadec, Iuliia Lysova, Iuri Macocco +2 | — | 11 Sept 2026 | — | — | — | — |
| BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti LanguagearXiv:2606.03504cs.CL | Muhammad Ali | — | 11 Sept 2026 | — | — | — | — |
Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Linked models, datasets, benchmarks and code come from stated relations only; a dash means no source stated one.