| RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition BiasesarXiv:2609.10092 | Yingqian Wu, Jingcong Liang, Siyuan Wang +2 | 11 Sept 2026 | cs.AI | — | 89 |
| From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric ReasoningarXiv:2609.10335 | Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou +2 | 11 Sept 2026 | cs.AI | — | 89 |
| ConvMem: Convolutional Memory for Long-Context ReasoningarXiv:2609.10441 | Hongming Zhang, Zhaozhen Gu, Fengshuo Bai +2 | 11 Sept 2026 | cs.AI | — | 89 |
| Quantifying Logical Consistency in Transformers via Query-Key AlignmentarXiv:2502.17017 | Eduard Tulchinskii, Anastasia Voznyuk, Laida Kushnareva +2 | 11 Sept 2026 | cs.CL | — | 89 |
| From Plausible to Actionable: A Position on LLM Self-ExplanationsarXiv:2607.15957 | Elize Herrewijnen, Benedetta Muscato, Gizem Gezici +1 | 11 Sept 2026 | cs.CL | — | 89 |
| AgenticGen: Reward-Guided Agentic Video Generation for AdvertisingarXiv:2609.09187 | Xingyuan Bu, Chengru Song, Hao Zhou +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Distribution-Consistent Inference for Dynamic Sparse Mixture-of-ExpertsarXiv:2609.09241 | Dohyeon Kim, Bedionita Soro, Sung Ju Hwang | 11 Sept 2026 | cs.LG | — | 89 |
| In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document PoisoningarXiv:2609.09243 | Iliano Fasolino | 11 Sept 2026 | cs.CR | — | 89 |
| Auditable Emergency Triage for Maternal and Newborn Care in IndiaarXiv:2609.09356 | Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini +2 | 11 Sept 2026 | cs.CL | — | 89 |
| From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function CallsarXiv:2609.09476 | Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi Nia | 11 Sept 2026 | cs.LG | — | 89 |
| Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QAarXiv:2609.09684 | Yuexin Wu, Dayou Yu, Vasile Rus | 11 Sept 2026 | cs.CL | — | 89 |
| Looped GPT-BERT: Trading Parameters for Computation in Small Language ModelingarXiv:2609.09691 | Tingshuo Fan, Hongtao Mu, Tianyu Zhou +2 | 11 Sept 2026 | cs.CL | — | 89 |
| When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document ContaminationarXiv:2609.09696 | Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar +1 | 11 Sept 2026 | cs.CL | — | 89 |
| Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?arXiv:2609.09768 | Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi | 11 Sept 2026 | cs.LG | — | 89 |
| LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial ScenariosarXiv:2609.09790 | Hanjing Zhou, Mingze Yin, Ying Lian +2 | 11 Sept 2026 | cs.CV | — | 89 |
| How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoEarXiv:2609.09793 | Yi Shi, Tanyu Chen, Kai Shen | 11 Sept 2026 | cs.CR | — | 89 |
| Strangers to Themselves: What Language Models Say About Themselves Is GenericarXiv:2609.09899 | Phil Blandfort, Urja Pawar | 11 Sept 2026 | cs.LG | — | 89 |
| Improving Cross-Lingual Token Representations by Adding a Pinch of SALTarXiv:2609.09953 | Guillem Ram\'irez | 11 Sept 2026 | cs.CL | — | 89 |
| MetroLLM-Bench: Evaluating Language Models as Transit Kiosk RuntimesarXiv:2609.10016 | Remco Hendriks (Continker) | 11 Sept 2026 | cs.LG | — | 71 |
| Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-TrainingarXiv:2609.10052 | Junwon Ko, Dong-Jae Lee, Minchan Kwon +2 | 11 Sept 2026 | cs.CL | — | 89 |
| NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic EnvironmentsarXiv:2609.10058 | Niramay M. Patel, Bibek Behera, Raksha Sharma | 11 Sept 2026 | cs.SD | — | 89 |
| Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-TuningarXiv:2609.10142 | Jing Guan, Yachao Yang, Zhaoliang Liu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented GenerationarXiv:2609.10239 | Daniel Alejandro Coll Tejeda, Pedro Garc\'ia L\'opez, Daniel Barcelona-Pons | 11 Sept 2026 | cs.IR | — | 89 |
| DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMsarXiv:2609.10253 | Bhuvan Arora, Devesh Saraogi, Sravya Varada +1 | 11 Sept 2026 | cs.CL | — | 89 |
| GANDR: Claim Auditing for Verifiable Legal Answer GenerationarXiv:2609.10293 | Chen Qian, Yimeng Wang, Yu Chen +2 | 11 Sept 2026 | cs.CL | — | 89 |
| RiLM: Parameter-Efficient Language Modeling via Geodesic DecodingarXiv:2609.10305 | Fang Li | 11 Sept 2026 | cs.CL | — | 89 |
| Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy OperationalizationarXiv:2609.10410 | Ayan Majumdar, Shounak Paul, Pushpdeep Singh +2 | 11 Sept 2026 | cs.CL | — | 89 |
| IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model IdentifierarXiv:2609.10494 | Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan | 11 Sept 2026 | cs.CL | — | 89 |
| Cultural Binding Heads in Language ModelsarXiv:2605.28543 | Avrile Floro, Luca Benedetto | 11 Sept 2026 | cs.AI | — | 89 |
| FrontierChallenge: Evaluating Scientific Workflow CompletionarXiv:2608.24979 | Liangcai Su, Zhaopeng Feng, Zhuo Chen +2 | 11 Sept 2026 | cs.AI | — | 89 |
| BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language ModelsarXiv:2408.10608 | Yongxin Deng (University of Technology Sydney), Xiaoyu Tan (National University of Singapore), Jing Pan (Monash University) +2 | 11 Sept 2026 | cs.CL | — | 89 |
| MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data GenerationarXiv:2510.05124 | Mingjin Li, Yu Liu, Huayi Liu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM JudgesarXiv:2601.08654 | Yihan Hong, Huaiyuan Yao, Bolin Shen +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space VisualizationarXiv:2601.15295 | Yi Wang, John Joon Young Chung, Melissa Roemmele +2 | 11 Sept 2026 | cs.HC | — | 89 |
| Revisiting the Shape Convention of Transformer Language ModelsarXiv:2602.06471 | Feng-Ting Liao, Guan-Ting Yi, Tzu-Quan Lin +2 | 11 Sept 2026 | cs.CL | — | 89 |
| False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UKarXiv:2602.13047 | Madhurananda Pahar, Caitlin Illingworth, Dorota Braun +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement LearningarXiv:2604.10701 | Zikang Shan, Han Zhong, Liwei Wang +1 | 11 Sept 2026 | cs.LG | — | 89 |
| Where is the Mind? Persona Vectors and LLM IndividuationarXiv:2604.17031 | Pierre Beckmann, Patrick Butlin | 11 Sept 2026 | cs.CL | — | 89 |
| "What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer UsearXiv:2605.00497 | Shardul Sapkota, Matthew J\"orke, Zane Sabbagh +2 | 11 Sept 2026 | cs.HC | — | 89 |
| EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsarXiv:2605.13841 | Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz +2 | 11 Sept 2026 | cs.SD | — | 89 |