Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages.

378 total

Reset
Papers
TitleAuthorsPublishedCategoriesOrganization / venueQuality
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model ExplorationarXiv:2609.09418Yiran Qiao, Feng Wang, Jing Ma11 Sept 2026cs.AI89
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?arXiv:2609.09428Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +111 Sept 2026cs.AI89
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal RepresentationsarXiv:2609.09448Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla11 Sept 2026cs.AI89
ContractEval: Query-Conditioned Execution Matching for Procedural Instruction ConformancearXiv:2609.09458Praphul Singh, Shanu Kumar, Akshat Agarwal +111 Sept 2026cs.AI89
Multi-Agent Agentic Graph Learning via Structural SignaturesarXiv:2609.09565Liang Qu, Jianxin Li, Hua Wang11 Sept 2026cs.AI89
CityPlanner: A Sandbox Agent for Executable Urban PlanningarXiv:2609.09578Wentao Zhang, Jingyuan Wang, Zetong Zhou +211 Sept 2026cs.AI89
A Function-Space Approach to the Statistical Mechanics of Learning DynamicsarXiv:2609.09589Yizhou Zhang, Weichen Wu, Lun Du +111 Sept 2026cs.AI89
From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital TwinsarXiv:2609.09625Haoran Gao, An Li, Zhen Li +111 Sept 2026cs.AI89
Seven Sources of Physical AI Capability FormationarXiv:2609.09627Gang Chen11 Sept 2026cs.AI89
RobustSGPO: Search-Space Control for Agent Harness EvolutionarXiv:2609.09646Zibo Zhao, Jijun Shi, Mo Zhou +211 Sept 2026cs.AI89
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk DiscoveryarXiv:2609.09647Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa +211 Sept 2026cs.AI89
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation SystemsarXiv:2609.09657Haichuan Hu, Yang Xiao, Mingni Tang +211 Sept 2026cs.AI89
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong ConversationsarXiv:2609.09664Hyojeong Yu, Hyukhun Koh, Minsung Kim +211 Sept 2026cs.AI89
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis AgentsarXiv:2609.09678Yuexin Wu, Vasile Rus11 Sept 2026cs.AI89
Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher DistillationarXiv:2609.09702Xiaofei Feng11 Sept 2026cs.AI89
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical ReasoningarXiv:2609.09707Yaning Jia, Chunhui Zhang, Wenxuan Xu +211 Sept 2026cs.AI89
Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online SafetyarXiv:2609.09735Hamed Jelodar, Amir Firouzi, Yen-Wu Lo +211 Sept 2026cs.AI89
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal AgentsarXiv:2609.09754Yujin Zhou, Mingxuan Zheng, Chuxue Cao +211 Sept 2026cs.AI89
Procedural Memory Under Change: Reuse and Interference in Controlled Web TasksarXiv:2609.09774Yanze Cao11 Sept 2026cs.AI89
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled RewardarXiv:2609.09776Eshwar Reddy M, Sourav Karmakar11 Sept 2026cs.AI89
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a ModelarXiv:2609.09815Xing Zhang, Guanghui Wang, Yanwei Cui +211 Sept 2026cs.AI89
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM AgentsarXiv:2609.09853Benjamin Gruenbaum, Doron Porat, Assaf Natanzon +211 Sept 2026cs.AI89
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction FieldsarXiv:2609.09864Cy Gorman, Yihang Yao11 Sept 2026cs.AI89
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI AgentsarXiv:2609.09875Shrey Nag, Sachita, Abhishek Kumar Singh +211 Sept 2026cs.AI89
Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation FormatarXiv:2609.09882Touchapon Kraisingkorn, Krittin Pachtrachai, Wachiravit Modecrua11 Sept 2026cs.AI89
Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D CommunicationsarXiv:2609.09885Yaxuan Liu11 Sept 2026cs.AI89
Grounded Evaluation and Repair for NL-to-PDDL Problem GenerationarXiv:2609.09898Joana Rosa, Pedro Santos, Valdemar Oliveira +211 Sept 2026cs.AI89
Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action ModelsarXiv:2609.09925Shengye Dong, Haochen Niu, Hao Liu +211 Sept 2026cs.AI89
Structural Process Supervision for Latent Chain-of-Thought ReasoningarXiv:2609.09928Yiqi Li, Xu Chen, Chen Ju +211 Sept 2026cs.AI89
Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial ObservabilityarXiv:2609.10036Arnab Chattopadhayay, Debdipta Halder11 Sept 2026cs.AI89
OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology NormalizationarXiv:2609.10055Jie Song, Zhichuan Xu, Ziyu Lu +211 Sept 2026cs.AI89
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden StatesarXiv:2609.10060Marek Jeli\'nski, Jan Dubi\'nski, Maciej Chrabaszcz +111 Sept 2026cs.AI89
RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition BiasesarXiv:2609.10092Yingqian Wu, Jingcong Liang, Siyuan Wang +211 Sept 2026cs.AI89
Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather AlertsarXiv:2609.10135Shuai Yan, Yang Xu, Shan He11 Sept 2026cs.AI89
Kernel-Managed Shared Memory for System-Wide PersonalizationarXiv:2609.10144Ryan Lum, Yongfeng Zhang11 Sept 2026cs.AI89
Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context LearningarXiv:2609.10177Mingbo Yang, Wenqiang Wang, Zhaolu Kang +211 Sept 2026cs.AI89
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool SelectionarXiv:2609.10221Haoyue Liu, Xiaoyu Ma, Ye Chen +211 Sept 2026cs.AI89
What Should an Agent Forget? Separating What Is Stored from What Is UsedarXiv:2609.10263Yuhang Li, Yuchen Li11 Sept 2026cs.AI89
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized RewardsarXiv:2609.10315Rui Sun, Zhan Shi, Bing He11 Sept 2026cs.AI89
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric ReasoningarXiv:2609.10335Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou +211 Sept 2026cs.AI89

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Each paper page shows the abstract, related models and every source snapshot.