Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages.

1,028 total

Reset
Papers
TitleAuthorsPublishedCategoriesOrganization / venueQuality
Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary ModelsarXiv:2609.10996Yu-Chung Hsiao11 Sept 2026cs.CL89
K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language ModelsarXiv:2609.11020Yu Sun, Mengyin Lu, Cong Feng +211 Sept 2026cs.CL89
ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute MediationarXiv:2609.11101Zesheng Wei, Mengfan Li, Wenhao Liu +211 Sept 2026cs.CL89
Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific PapersarXiv:2609.11117Hanhua Hong, Yizhi Li, Luu Gia Huy +211 Sept 2026cs.CL89
From Repetition to Recognition: Inductive Discovery of Disinformation NarrativesarXiv:2609.11128Max Upravitelev, Veronika Solopova, Jing Yang +211 Sept 2026cs.CL89
Rubric-Aligned Disentangled Evaluation of Human Simultaneous InterpretingarXiv:2609.11131Ziyu Zhang, Satoshi Nakamura11 Sept 2026cs.CL89
Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema NormalizationarXiv:2609.11141Dong-Jae Koh, Huisu Kim, SeongHwan Yoon +211 Sept 2026cs.CL89
FlexComp: One Model for Every Ratio in Context CompressionarXiv:2609.11192Kaiyan Zhao, Zhongtao Miao, Akiko Aizawa +111 Sept 2026cs.CL89
Automated Identification of Competing Narratives in Political Discourse on Social MediaarXiv:2609.11202Sergej Wildemann, Erick Elejalde11 Sept 2026cs.CL89
OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language ModelsarXiv:2609.11244Jianjiang Yang, Peihang Li, Shanqing Xu +211 Sept 2026cs.CL89
Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case StudyarXiv:2609.11246Hiwa Asadpour11 Sept 2026cs.CL89
The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based MethodsarXiv:2609.11247Ioanna Kaffeza, Efthymios Georgiou, Alexandros Potamianos11 Sept 2026cs.CL89
Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper AdaptationarXiv:2609.11302Maria Frangiadaki, Dimitrios Damianos, Kosmas Kritsis +111 Sept 2026cs.CL89
MultiHuSE: A Multimodal Dataset for Humour Styles and EmotionsarXiv:2609.11322Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat11 Sept 2026cs.CL89
On the Impact of Anonymization on the Performance of Large Language ModelsarXiv:2609.11335Tobias Deu{\ss}er, Max Hahnb\"uck, Lorenz Sparrenberg +211 Sept 2026cs.CL89
SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQarXiv:2609.11355Huy Hoang Le, Long-Bao Nguyen, Minh Tri Dao11 Sept 2026cs.CL89
TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model OutputsarXiv:2609.11399Shenbin Qian, Yves Scherrer11 Sept 2026cs.CL89
SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model ConversationsarXiv:2609.11414Yu Wang, Yuchen Li, Rui Kong +211 Sept 2026cs.CL89
Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language StudyarXiv:2609.11450\'Alvaro Rey-Blanes, Francisco J. Moreno-Barea, Francisco J. Veredas11 Sept 2026cs.CL89
ReGround: Grounding Reviewer Comments in Multimodal EvidencearXiv:2609.11460Serwar Basch, Lizhen Qu, Iryna Gurevych11 Sept 2026cs.CL89
Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-SpeecharXiv:2609.11545Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +211 Sept 2026cs.CL89
A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC FilingsarXiv:2609.11620Jean-Fran\c{c}ois Delpech11 Sept 2026cs.CL89
Structured Transforms for Low-Overhead Quantization of Language ModelsarXiv:2609.11687Daria Cherniuk, Alexander Rudikov, Boris Kashin +111 Sept 2026cs.CL89
The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challengearXiv:2609.11724Jordi Luque, Lorenzo Concina, Marco Matassoni +211 Sept 2026cs.CL89
RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM SafetyarXiv:2609.11758Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser11 Sept 2026cs.CL89
Component-Aware Differential Privacy for Federated Multilingual Speech-LLMsarXiv:2609.11762Jordi Luque, Fernando L\'opez, Aleix Sant11 Sept 2026cs.CL89
Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News FramingarXiv:2609.11769Yi Liu11 Sept 2026cs.CL89
The widening evaluation gap in medical large language model research 2023 to 2026arXiv:2609.11770Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif11 Sept 2026cs.CL89
Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched SpeecharXiv:2609.11786Chibuzor Okocha, Christan Earl Grant11 Sept 2026cs.CL89
Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation modelsarXiv:2609.11838Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif +211 Sept 2026cs.CL89
IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-MixingarXiv:2609.11851Pruthwik Mishra, Rudra Trivedi, Avi Patel +211 Sept 2026cs.CL89
Epistemic orientation predicts legislative effectiveness among members of the US CongressarXiv:2609.11865Segun Aroyehun, Stephan Lewandowsky, David Garcia11 Sept 2026cs.CL89
Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language ModelarXiv:2609.11870Lisa Bylinina11 Sept 2026cs.CL89
Nuha-Speech: Building General-Purpose Arabic Speech-LLMsarXiv:2609.11892Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi11 Sept 2026cs.CL89
Distance generalization in transformers: why bother with positional encoding?arXiv:2609.11913Daniel Henrik Nevermann, Claudius Gros11 Sept 2026cs.CL89
More than half of recent astronomy papers are written with language-model assistancearXiv:2609.10664Serat M. Saad, Yuan-Sen Ting11 Sept 2026astro-ph.IM89
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question GenerationarXiv:2609.10815Karish Gupta, Matthew Alex, Alex Li +211 Sept 2026cs.CV89
KuaiRP Series Role-playing Models Technical ReportarXiv:2609.11127Yipeng Wang, Ziwei Zhang, Jiahui Zhang +211 Sept 2026cs.AI89
Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial SentimentarXiv:2609.11144AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani11 Sept 2026cs.AI89
(Whose defaults?) Is artificial intelligence reorienting archaeological methods?arXiv:2609.11198Lorenzo Cardarelli, Roberto Ragno11 Sept 2026cs.CY89

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Each paper page shows the abstract, related models and every source snapshot.