Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages.

210 total

Reset
Papers
TitleAuthorsPublishedCategoriesOrganization / venueQuality
RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition BiasesarXiv:2609.10092Yingqian Wu, Jingcong Liang, Siyuan Wang +211 Sept 2026cs.AI89
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric ReasoningarXiv:2609.10335Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou +211 Sept 2026cs.AI89
ConvMem: Convolutional Memory for Long-Context ReasoningarXiv:2609.10441Hongming Zhang, Zhaozhen Gu, Fengshuo Bai +211 Sept 2026cs.AI89
Quantifying Logical Consistency in Transformers via Query-Key AlignmentarXiv:2502.17017Eduard Tulchinskii, Anastasia Voznyuk, Laida Kushnareva +211 Sept 2026cs.CL89
From Plausible to Actionable: A Position on LLM Self-ExplanationsarXiv:2607.15957Elize Herrewijnen, Benedetta Muscato, Gizem Gezici +111 Sept 2026cs.CL89
AgenticGen: Reward-Guided Agentic Video Generation for AdvertisingarXiv:2609.09187Xingyuan Bu, Chengru Song, Hao Zhou +211 Sept 2026cs.CV89
Distribution-Consistent Inference for Dynamic Sparse Mixture-of-ExpertsarXiv:2609.09241Dohyeon Kim, Bedionita Soro, Sung Ju Hwang11 Sept 2026cs.LG89
In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document PoisoningarXiv:2609.09243Iliano Fasolino11 Sept 2026cs.CR89
Auditable Emergency Triage for Maternal and Newborn Care in IndiaarXiv:2609.09356Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini +211 Sept 2026cs.CL89
From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function CallsarXiv:2609.09476Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi Nia11 Sept 2026cs.LG89
Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QAarXiv:2609.09684Yuexin Wu, Dayou Yu, Vasile Rus11 Sept 2026cs.CL89
Looped GPT-BERT: Trading Parameters for Computation in Small Language ModelingarXiv:2609.09691Tingshuo Fan, Hongtao Mu, Tianyu Zhou +211 Sept 2026cs.CL89
When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document ContaminationarXiv:2609.09696Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar +111 Sept 2026cs.CL89
Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?arXiv:2609.09768Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi11 Sept 2026cs.LG89
LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial ScenariosarXiv:2609.09790Hanjing Zhou, Mingze Yin, Ying Lian +211 Sept 2026cs.CV89
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoEarXiv:2609.09793Yi Shi, Tanyu Chen, Kai Shen11 Sept 2026cs.CR89
Strangers to Themselves: What Language Models Say About Themselves Is GenericarXiv:2609.09899Phil Blandfort, Urja Pawar11 Sept 2026cs.LG89
Improving Cross-Lingual Token Representations by Adding a Pinch of SALTarXiv:2609.09953Guillem Ram\'irez11 Sept 2026cs.CL89
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk RuntimesarXiv:2609.10016Remco Hendriks (Continker)11 Sept 2026cs.LG71
Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-TrainingarXiv:2609.10052Junwon Ko, Dong-Jae Lee, Minchan Kwon +211 Sept 2026cs.CL89
NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic EnvironmentsarXiv:2609.10058Niramay M. Patel, Bibek Behera, Raksha Sharma11 Sept 2026cs.SD89
Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-TuningarXiv:2609.10142Jing Guan, Yachao Yang, Zhaoliang Liu +211 Sept 2026cs.CL89
LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented GenerationarXiv:2609.10239Daniel Alejandro Coll Tejeda, Pedro Garc\'ia L\'opez, Daniel Barcelona-Pons11 Sept 2026cs.IR89
DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMsarXiv:2609.10253Bhuvan Arora, Devesh Saraogi, Sravya Varada +111 Sept 2026cs.CL89
GANDR: Claim Auditing for Verifiable Legal Answer GenerationarXiv:2609.10293Chen Qian, Yimeng Wang, Yu Chen +211 Sept 2026cs.CL89
RiLM: Parameter-Efficient Language Modeling via Geodesic DecodingarXiv:2609.10305Fang Li11 Sept 2026cs.CL89
Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy OperationalizationarXiv:2609.10410Ayan Majumdar, Shounak Paul, Pushpdeep Singh +211 Sept 2026cs.CL89
IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model IdentifierarXiv:2609.10494Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan11 Sept 2026cs.CL89
Cultural Binding Heads in Language ModelsarXiv:2605.28543Avrile Floro, Luca Benedetto11 Sept 2026cs.AI89
FrontierChallenge: Evaluating Scientific Workflow CompletionarXiv:2608.24979Liangcai Su, Zhaopeng Feng, Zhuo Chen +211 Sept 2026cs.AI89
BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language ModelsarXiv:2408.10608Yongxin Deng (University of Technology Sydney), Xiaoyu Tan (National University of Singapore), Jing Pan (Monash University) +211 Sept 2026cs.CL89
MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data GenerationarXiv:2510.05124Mingjin Li, Yu Liu, Huayi Liu +211 Sept 2026cs.CL89
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM JudgesarXiv:2601.08654Yihan Hong, Huaiyuan Yao, Bolin Shen +211 Sept 2026cs.CL89
Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space VisualizationarXiv:2601.15295Yi Wang, John Joon Young Chung, Melissa Roemmele +211 Sept 2026cs.HC89
Revisiting the Shape Convention of Transformer Language ModelsarXiv:2602.06471Feng-Ting Liao, Guan-Ting Yi, Tzu-Quan Lin +211 Sept 2026cs.CL89
False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UKarXiv:2602.13047Madhurananda Pahar, Caitlin Illingworth, Dorota Braun +211 Sept 2026cs.CL89
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement LearningarXiv:2604.10701Zikang Shan, Han Zhong, Liwei Wang +111 Sept 2026cs.LG89
Where is the Mind? Persona Vectors and LLM IndividuationarXiv:2604.17031Pierre Beckmann, Patrick Butlin11 Sept 2026cs.CL89
"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer UsearXiv:2605.00497Shardul Sapkota, Matthew J\"orke, Zane Sabbagh +211 Sept 2026cs.HC89
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsarXiv:2605.13841Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz +211 Sept 2026cs.SD89

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Each paper page shows the abstract, related models and every source snapshot.