| Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model ExplorationarXiv:2609.09418cs.AI | Yiran Qiao, Feng Wang, Jing Ma | — | 11 Sept 2026 | — | — | — | — |
| XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?arXiv:2609.09428cs.AI | Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +1 | — | 11 Sept 2026 | — | — | — | — |
| Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal RepresentationsarXiv:2609.09448cs.AI | Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla | — | 11 Sept 2026 | — | — | — | — |
| ContractEval: Query-Conditioned Execution Matching for Procedural Instruction ConformancearXiv:2609.09458cs.AI | Praphul Singh, Shanu Kumar, Akshat Agarwal +1 | — | 11 Sept 2026 | — | — | — | — |
| Multi-Agent Agentic Graph Learning via Structural SignaturesarXiv:2609.09565cs.AI | Liang Qu, Jianxin Li, Hua Wang | — | 11 Sept 2026 | — | — | — | — |
| CityPlanner: A Sandbox Agent for Executable Urban PlanningarXiv:2609.09578cs.AI | Wentao Zhang, Jingyuan Wang, Zetong Zhou +2 | — | 11 Sept 2026 | — | — | — | — |
| A Function-Space Approach to the Statistical Mechanics of Learning DynamicsarXiv:2609.09589cs.AI | Yizhou Zhang, Weichen Wu, Lun Du +1 | — | 11 Sept 2026 | — | — | — | — |
| From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital TwinsarXiv:2609.09625cs.AI | Haoran Gao, An Li, Zhen Li +1 | — | 11 Sept 2026 | — | — | — | — |
| Seven Sources of Physical AI Capability FormationarXiv:2609.09627cs.AI | Gang Chen | — | 11 Sept 2026 | — | — | — | — |
| RobustSGPO: Search-Space Control for Agent Harness EvolutionarXiv:2609.09646cs.AI | Zibo Zhao, Jijun Shi, Mo Zhou +2 | — | 11 Sept 2026 | — | — | — | — |
| Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk DiscoveryarXiv:2609.09647cs.AI | Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa +2 | — | 11 Sept 2026 | — | — | — | — |
| RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation SystemsarXiv:2609.09657cs.AI | Haichuan Hu, Yang Xiao, Mingni Tang +2 | — | 11 Sept 2026 | — | — | — | — |
| PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong ConversationsarXiv:2609.09664cs.AI | Hyojeong Yu, Hyukhun Koh, Minsung Kim +2 | — | 11 Sept 2026 | — | — | — | — |
| Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis AgentsarXiv:2609.09678cs.AI | Yuexin Wu, Vasile Rus | — | 11 Sept 2026 | — | — | — | — |
| Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher DistillationarXiv:2609.09702cs.AI | Xiaofei Feng | — | 11 Sept 2026 | — | — | — | — |
| Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical ReasoningarXiv:2609.09707cs.AI | Yaning Jia, Chunhui Zhang, Wenxuan Xu +2 | — | 11 Sept 2026 | — | — | — | — |
| Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online SafetyarXiv:2609.09735cs.AI | Hamed Jelodar, Amir Firouzi, Yen-Wu Lo +2 | — | 11 Sept 2026 | — | — | — | — |
| LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal AgentsarXiv:2609.09754cs.AI | Yujin Zhou, Mingxuan Zheng, Chuxue Cao +2 | — | 11 Sept 2026 | — | — | — | — |
| Procedural Memory Under Change: Reuse and Interference in Controlled Web TasksarXiv:2609.09774cs.AI | Yanze Cao | — | 11 Sept 2026 | — | — | — | — |
| Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled RewardarXiv:2609.09776cs.AI | Eshwar Reddy M, Sourav Karmakar | — | 11 Sept 2026 | — | — | — | — |
| UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a ModelarXiv:2609.09815cs.AI | Xing Zhang, Guanghui Wang, Yanwei Cui +2 | — | 11 Sept 2026 | — | — | — | — |
| The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM AgentsarXiv:2609.09853cs.AI | Benjamin Gruenbaum, Doron Porat, Assaf Natanzon +2 | — | 11 Sept 2026 | — | — | — | — |
| Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction FieldsarXiv:2609.09864cs.AI | Cy Gorman, Yihang Yao | — | 11 Sept 2026 | — | — | — | — |
| AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI AgentsarXiv:2609.09875cs.AI | Shrey Nag, Sachita, Abhishek Kumar Singh +2 | — | 11 Sept 2026 | — | — | — | — |
| Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation FormatarXiv:2609.09882cs.AI | Touchapon Kraisingkorn, Krittin Pachtrachai, Wachiravit Modecrua | — | 11 Sept 2026 | — | — | — | — |