Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages; model links come from model cards citing the paper.

510 papers

Papers
TitleAuthorsOrganizationPublishedIntroduces Models (and artifacts) whose model card or documentation cites this paper — inbound described_by relations.DatasetsBenchmarksCode
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model ExplorationarXiv:2609.09418cs.AIYiran Qiao, Feng Wang, Jing Ma11 Sept 2026
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?arXiv:2609.09428cs.AIYanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +111 Sept 2026
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal RepresentationsarXiv:2609.09448cs.AIPriyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla11 Sept 2026
ContractEval: Query-Conditioned Execution Matching for Procedural Instruction ConformancearXiv:2609.09458cs.AIPraphul Singh, Shanu Kumar, Akshat Agarwal +111 Sept 2026
Multi-Agent Agentic Graph Learning via Structural SignaturesarXiv:2609.09565cs.AILiang Qu, Jianxin Li, Hua Wang11 Sept 2026
CityPlanner: A Sandbox Agent for Executable Urban PlanningarXiv:2609.09578cs.AIWentao Zhang, Jingyuan Wang, Zetong Zhou +211 Sept 2026
A Function-Space Approach to the Statistical Mechanics of Learning DynamicsarXiv:2609.09589cs.AIYizhou Zhang, Weichen Wu, Lun Du +111 Sept 2026
From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital TwinsarXiv:2609.09625cs.AIHaoran Gao, An Li, Zhen Li +111 Sept 2026
Seven Sources of Physical AI Capability FormationarXiv:2609.09627cs.AIGang Chen11 Sept 2026
RobustSGPO: Search-Space Control for Agent Harness EvolutionarXiv:2609.09646cs.AIZibo Zhao, Jijun Shi, Mo Zhou +211 Sept 2026
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk DiscoveryarXiv:2609.09647cs.AIDivyanshu Kumar, Nitin Aravind Birur, Tanay Baswa +211 Sept 2026
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation SystemsarXiv:2609.09657cs.AIHaichuan Hu, Yang Xiao, Mingni Tang +211 Sept 2026
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong ConversationsarXiv:2609.09664cs.AIHyojeong Yu, Hyukhun Koh, Minsung Kim +211 Sept 2026
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis AgentsarXiv:2609.09678cs.AIYuexin Wu, Vasile Rus11 Sept 2026
Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher DistillationarXiv:2609.09702cs.AIXiaofei Feng11 Sept 2026
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical ReasoningarXiv:2609.09707cs.AIYaning Jia, Chunhui Zhang, Wenxuan Xu +211 Sept 2026
Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online SafetyarXiv:2609.09735cs.AIHamed Jelodar, Amir Firouzi, Yen-Wu Lo +211 Sept 2026
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal AgentsarXiv:2609.09754cs.AIYujin Zhou, Mingxuan Zheng, Chuxue Cao +211 Sept 2026
Procedural Memory Under Change: Reuse and Interference in Controlled Web TasksarXiv:2609.09774cs.AIYanze Cao11 Sept 2026
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled RewardarXiv:2609.09776cs.AIEshwar Reddy M, Sourav Karmakar11 Sept 2026
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a ModelarXiv:2609.09815cs.AIXing Zhang, Guanghui Wang, Yanwei Cui +211 Sept 2026
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM AgentsarXiv:2609.09853cs.AIBenjamin Gruenbaum, Doron Porat, Assaf Natanzon +211 Sept 2026
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction FieldsarXiv:2609.09864cs.AICy Gorman, Yihang Yao11 Sept 2026
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI AgentsarXiv:2609.09875cs.AIShrey Nag, Sachita, Abhishek Kumar Singh +211 Sept 2026
Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation FormatarXiv:2609.09882cs.AITouchapon Kraisingkorn, Krittin Pachtrachai, Wachiravit Modecrua11 Sept 2026

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Linked models, datasets, benchmarks and code come from stated relations only; a dash means no source stated one.