| The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM SocietiesarXiv:2509.18052 | Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech DetectionarXiv:2510.15685 | Joshua Wolfe Brook, Ilia Markov | 11 Sept 2026 | cs.CL | — | 89 |
| Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive FactorsarXiv:2511.17036 | Gyuwon Park, Hyounghun Kim | 11 Sept 2026 | cs.CL | — | 89 |
| DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert ReportsarXiv:2601.08536 | Ruizhe Li, Mingxuan Du, Benfeng Xu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical ConsultationarXiv:2601.15645 | Zhiyao Ren, Yibing Zhan, Siyuan Liang +2 | 11 Sept 2026 | cs.CL | — | 89 |
| What Language is This? Ask Your TokenizerarXiv:2602.17655 | Clara Meister, Ahmetcan Yavuz, Pietro Lesci +1 | 11 Sept 2026 | cs.CL | — | 89 |
| Streaming Translation and Transcription Through Speech-to-Text Causal AlignmentarXiv:2603.11578 | Roman Koshkin, Jeon Haesung, Lianbo Liu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social InteractionarXiv:2603.17094 | Ryo Kamoi, Ameya Godbole, Binglin Zhou +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and StudyarXiv:2603.24125 | Nour Bouchouchi, Thibault Laugel, Xavier Renard +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Timing is Everything: Temporal Scaffolding of Semantic Surprise in HumorarXiv:2605.00143 | Yuxi Ma, Yongqian Peng, Junchen Lyu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and DistillationarXiv:2605.12227 | Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins | 11 Sept 2026 | cs.CL | — | 89 |
| Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accountsarXiv:2605.21049 | Ni Yang, Rui He, Philipp Homan +2 | 11 Sept 2026 | cs.CL | — | 89 |
| MERIT: Matching Expertise via Rubric-Informed Training for Reviewer AssignmentarXiv:2605.27865 | Zixuan Yang, Yibo Zhao, Weicong Liu +1 | 11 Sept 2026 | cs.CL | — | 89 |
| Characterizing Narrative Content in Web-scale LLM Pretraining DataarXiv:2606.19468 | Teagan Johnson, Elliott Ash, Andrew Piper +1 | 11 Sept 2026 | cs.CL | — | 89 |
| Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI DialoguearXiv:2606.21844 | William Hager, Ishika Rathi, Masum Hasan +1 | 11 Sept 2026 | cs.CL | — | 89 |
| A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar BooksarXiv:2607.22376 | Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager +1 | 11 Sept 2026 | cs.CL | — | 89 |
| Predicting Startup Exit from Textual Descriptors - A Computational Linguistics FrameworkarXiv:2608.00045 | Alberto M. G. Saruggia, Sebastien Germano | 11 Sept 2026 | cs.CL | — | 89 |
| Causal Episodic Memory for Feedback-Driven Agent RepairarXiv:2608.05906 | Khang Nhat Hoang Vo, Tam Minh Chu, Anh Trac Duc Dinh +2 | 11 Sept 2026 | cs.CL | — | 89 |
| VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool UsearXiv:2608.08477 | Juan S. Santillana | 11 Sept 2026 | cs.CL | — | 89 |
| Self-Evolving Embodied Agents via Skill-Harness EvolutionarXiv:2608.11350 | Peidong Wang, Zhiming Ma, Ying Chang +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization AnalysisarXiv:2608.16379 | Hiwa Asadpour | 11 Sept 2026 | cs.CL | — | 89 |
| Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based ModerationarXiv:2608.22230 | Junyu Lu, Kaiyuan Liu, Kaichun Wang +2 | 11 Sept 2026 | cs.CL | — | 89 |
| DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database CompletionarXiv:2608.22770 | Xuan Yao, Shuping Li, Yang Dai +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case StudyarXiv:2608.29170 | Zijie Zhang, Tan Lee, Yong Cao +1 | 11 Sept 2026 | cs.CL | — | 89 |
| Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation RerankingarXiv:2609.00588 | Guangyu Chen, Boxuan Lyu, Hidetaka Kamigaito +2 | 11 Sept 2026 | cs.CL | — | 89 |
| OUTLETS: Output-Length Prediction from Speculative Decoding BackbonesarXiv:2609.01068 | Weihuang Wen, Yingying Liu, Yichuan Liu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| Cache-Aware Joint Router Adaptation for Memory-Efficient MoE InferencearXiv:2609.04895 | Zhenhe Wu, Yaping Jin, Qinghua Xing +2 | 11 Sept 2026 | cs.CL | — | 89 |
| CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn PrescriptionarXiv:2609.09766 | Minjoo Kim, Sangjin Park, Seung Hwan Cho | 11 Sept 2026 | cs.CL | — | 89 |
| SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better TeachersarXiv:2609.09999 | Xixian Liao, Maite Melero | 11 Sept 2026 | cs.CL | — | 89 |
| Emergent Risks in Generative Multi-Agent SystemsarXiv:2603.27771 | Yue Huang, Yu Jiang, Wenjie Wang +2 | 11 Sept 2026 | cs.MA | — | 89 |
| MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math TeachersarXiv:2604.04036 | Zhihan Guo, Yuting Lu, Jionghao Lin | 11 Sept 2026 | cs.IR | — | 89 |
| Formalizing building-up constructions of self-dual codes through isotropic lines in LeanarXiv:2604.08485 | Jae-Hyun Baek, Jon-Lark Kim | 11 Sept 2026 | cs.IT | — | 89 |
| LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial DomainsarXiv:2604.16379 | Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda | 11 Sept 2026 | cs.IR | — | 89 |
| Strategic Type SpacesarXiv:2606.08297 | Olivier Gossner, Rafael Veiel | 11 Sept 2026 | econ.TH | — | 89 |
| A Group-Based Resource Allocation Model for the Fractional Knapsack ProblemarXiv:2609.06470 | Abhinaba Chakraborty | 11 Sept 2026 | cs.DS | — | 89 |
| Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic TasksarXiv:2609.09233 | Wasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth +2 | 11 Sept 2026 | cs.AI | — | 89 |
| CityPlanner: A Sandbox Agent for Executable Urban PlanningarXiv:2609.09578 | Wentao Zhang, Jingyuan Wang, Zetong Zhou +2 | 11 Sept 2026 | cs.AI | — | 89 |
| Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online SafetyarXiv:2609.09735 | Hamed Jelodar, Amir Firouzi, Yen-Wu Lo +2 | 11 Sept 2026 | cs.AI | — | 89 |
| UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a ModelarXiv:2609.09815 | Xing Zhang, Guanghui Wang, Yanwei Cui +2 | 11 Sept 2026 | cs.AI | — | 89 |
| OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology NormalizationarXiv:2609.10055 | Jie Song, Zhichuan Xu, Ziyu Lu +2 | 11 Sept 2026 | cs.AI | — | 89 |