Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages; model links come from model cards citing the paper.

1,189 papers

Papers
TitleAuthorsOrganizationPublishedIntroduces Models (and artifacts) whose model card or documentation cites this paper — inbound described_by relations.DatasetsBenchmarksCode
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical ConsultationarXiv:2601.15645cs.CLZhiyao Ren, Yibing Zhan, Siyuan Liang +211 Sept 2026
What Language is This? Ask Your TokenizerarXiv:2602.17655cs.CLClara Meister, Ahmetcan Yavuz, Pietro Lesci +111 Sept 2026
Streaming Translation and Transcription Through Speech-to-Text Causal AlignmentarXiv:2603.11578cs.CLRoman Koshkin, Jeon Haesung, Lianbo Liu +211 Sept 2026
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social InteractionarXiv:2603.17094cs.CLRyo Kamoi, Ameya Godbole, Binglin Zhou +211 Sept 2026
Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and StudyarXiv:2603.24125cs.CLNour Bouchouchi, Thibault Laugel, Xavier Renard +211 Sept 2026
Timing is Everything: Temporal Scaffolding of Semantic Surprise in HumorarXiv:2605.00143cs.CLYuxi Ma, Yongqian Peng, Junchen Lyu +211 Sept 2026
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and DistillationarXiv:2605.12227cs.CLMiguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins11 Sept 2026
Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accountsarXiv:2605.21049cs.CLNi Yang, Rui He, Philipp Homan +211 Sept 2026
MERIT: Matching Expertise via Rubric-Informed Training for Reviewer AssignmentarXiv:2605.27865cs.CLZixuan Yang, Yibo Zhao, Weicong Liu +111 Sept 2026
Characterizing Narrative Content in Web-scale LLM Pretraining DataarXiv:2606.19468cs.CLTeagan Johnson, Elliott Ash, Andrew Piper +111 Sept 2026
Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI DialoguearXiv:2606.21844cs.CLWilliam Hager, Ishika Rathi, Masum Hasan +111 Sept 2026
A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar BooksarXiv:2607.22376cs.CLVarun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager +111 Sept 2026
Predicting Startup Exit from Textual Descriptors - A Computational Linguistics FrameworkarXiv:2608.00045cs.CLAlberto M. G. Saruggia, Sebastien Germano11 Sept 2026
Causal Episodic Memory for Feedback-Driven Agent RepairarXiv:2608.05906cs.CLKhang Nhat Hoang Vo, Tam Minh Chu, Anh Trac Duc Dinh +211 Sept 2026
VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool UsearXiv:2608.08477cs.CLJuan S. Santillana11 Sept 2026
Self-Evolving Embodied Agents via Skill-Harness EvolutionarXiv:2608.11350cs.CLPeidong Wang, Zhiming Ma, Ying Chang +211 Sept 2026
Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization AnalysisarXiv:2608.16379cs.CLHiwa Asadpour11 Sept 2026
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based ModerationarXiv:2608.22230cs.CLJunyu Lu, Kaiyuan Liu, Kaichun Wang +211 Sept 2026
DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database CompletionarXiv:2608.22770cs.CLXuan Yao, Shuping Li, Yang Dai +211 Sept 2026
Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case StudyarXiv:2608.29170cs.CLZijie Zhang, Tan Lee, Yong Cao +111 Sept 2026
Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation RerankingarXiv:2609.00588cs.CLGuangyu Chen, Boxuan Lyu, Hidetaka Kamigaito +211 Sept 2026
OUTLETS: Output-Length Prediction from Speculative Decoding BackbonesarXiv:2609.01068cs.CLWeihuang Wen, Yingying Liu, Yichuan Liu +211 Sept 2026
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE InferencearXiv:2609.04895cs.CLZhenhe Wu, Yaping Jin, Qinghua Xing +211 Sept 2026
CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn PrescriptionarXiv:2609.09766cs.CLMinjoo Kim, Sangjin Park, Seung Hwan Cho11 Sept 2026
SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better TeachersarXiv:2609.09999cs.CLXixian Liao, Maite Melero11 Sept 2026

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Linked models, datasets, benchmarks and code come from stated relations only; a dash means no source stated one.