| How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoEarXiv:2609.09793cs.CR | Yi Shi, Tanyu Chen, Kai Shen | — | 11 Sept 2026 | — | — | — | — |
| Strangers to Themselves: What Language Models Say About Themselves Is GenericarXiv:2609.09899cs.LG | Phil Blandfort, Urja Pawar | — | 11 Sept 2026 | — | — | — | — |
| Improving Cross-Lingual Token Representations by Adding a Pinch of SALTarXiv:2609.09953cs.CL | Guillem Ram\'irez | — | 11 Sept 2026 | — | — | — | — |
| MetroLLM-Bench: Evaluating Language Models as Transit Kiosk RuntimesarXiv:2609.10016cs.LG | Remco Hendriks (Continker) | — | 11 Sept 2026 | — | — | — | — |
| Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-TrainingarXiv:2609.10052cs.CL | Junwon Ko, Dong-Jae Lee, Minchan Kwon +2 | — | 11 Sept 2026 | — | — | — | — |
| NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic EnvironmentsarXiv:2609.10058cs.SD | Niramay M. Patel, Bibek Behera, Raksha Sharma | — | 11 Sept 2026 | — | — | — | — |
| Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-TuningarXiv:2609.10142cs.CL | Jing Guan, Yachao Yang, Zhaoliang Liu +2 | — | 11 Sept 2026 | — | — | — | — |
| LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented GenerationarXiv:2609.10239cs.IR | Daniel Alejandro Coll Tejeda, Pedro Garc\'ia L\'opez, Daniel Barcelona-Pons | — | 11 Sept 2026 | — | — | — | — |
| DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMsarXiv:2609.10253cs.CL | Bhuvan Arora, Devesh Saraogi, Sravya Varada +1 | — | 11 Sept 2026 | — | — | — | — |
| GANDR: Claim Auditing for Verifiable Legal Answer GenerationarXiv:2609.10293cs.CL | Chen Qian, Yimeng Wang, Yu Chen +2 | — | 11 Sept 2026 | — | — | — | — |
| RiLM: Parameter-Efficient Language Modeling via Geodesic DecodingarXiv:2609.10305cs.CL | Fang Li | — | 11 Sept 2026 | — | — | — | — |
| Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy OperationalizationarXiv:2609.10410cs.CL | Ayan Majumdar, Shounak Paul, Pushpdeep Singh +2 | — | 11 Sept 2026 | — | — | — | — |
| IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model IdentifierarXiv:2609.10494cs.CL | Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan | — | 11 Sept 2026 | — | — | — | — |
| Cultural Binding Heads in Language ModelsarXiv:2605.28543cs.AI | Avrile Floro, Luca Benedetto | — | 11 Sept 2026 | — | — | — | — |
| FrontierChallenge: Evaluating Scientific Workflow CompletionarXiv:2608.24979cs.AI | Liangcai Su, Zhaopeng Feng, Zhuo Chen +2 | — | 11 Sept 2026 | — | — | — | — |
| BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language ModelsarXiv:2408.10608cs.CL | Yongxin Deng (University of Technology Sydney), Xiaoyu Tan (National University of Singapore), Jing Pan (Monash University) +2 | — | 11 Sept 2026 | — | — | — | — |
| MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data GenerationarXiv:2510.05124cs.CL | Mingjin Li, Yu Liu, Huayi Liu +2 | — | 11 Sept 2026 | — | — | — | — |
| From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM JudgesarXiv:2601.08654cs.CL | Yihan Hong, Huaiyuan Yao, Bolin Shen +2 | — | 11 Sept 2026 | — | — | — | — |
| Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space VisualizationarXiv:2601.15295cs.HC | Yi Wang, John Joon Young Chung, Melissa Roemmele +2 | — | 11 Sept 2026 | — | — | — | — |
| Revisiting the Shape Convention of Transformer Language ModelsarXiv:2602.06471cs.CL | Feng-Ting Liao, Guan-Ting Yi, Tzu-Quan Lin +2 | — | 11 Sept 2026 | — | — | — | — |
| False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UKarXiv:2602.13047cs.CL | Madhurananda Pahar, Caitlin Illingworth, Dorota Braun +2 | — | 11 Sept 2026 | — | — | — | — |
| Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement LearningarXiv:2604.10701cs.LG | Zikang Shan, Han Zhong, Liwei Wang +1 | — | 11 Sept 2026 | — | — | — | — |
| Where is the Mind? Persona Vectors and LLM IndividuationarXiv:2604.17031cs.CL | Pierre Beckmann, Patrick Butlin | — | 11 Sept 2026 | — | — | — | — |
| "What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer UsearXiv:2605.00497cs.HC | Shardul Sapkota, Matthew J\"orke, Zane Sabbagh +2 | — | 11 Sept 2026 | — | — | — | — |
| EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsarXiv:2605.13841cs.SD | Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz +2 | — | 11 Sept 2026 | — | — | — | — |