Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages.

210 total

Reset
Papers
TitleAuthorsPublishedCategoriesOrganization / venueQuality
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM SocietiesarXiv:2509.18052Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou +211 Sept 2026cs.CL89
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech DetectionarXiv:2510.15685Joshua Wolfe Brook, Ilia Markov11 Sept 2026cs.CL89
Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive FactorsarXiv:2511.17036Gyuwon Park, Hyounghun Kim11 Sept 2026cs.CL89
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert ReportsarXiv:2601.08536Ruizhe Li, Mingxuan Du, Benfeng Xu +211 Sept 2026cs.CL89
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical ConsultationarXiv:2601.15645Zhiyao Ren, Yibing Zhan, Siyuan Liang +211 Sept 2026cs.CL89
What Language is This? Ask Your TokenizerarXiv:2602.17655Clara Meister, Ahmetcan Yavuz, Pietro Lesci +111 Sept 2026cs.CL89
Streaming Translation and Transcription Through Speech-to-Text Causal AlignmentarXiv:2603.11578Roman Koshkin, Jeon Haesung, Lianbo Liu +211 Sept 2026cs.CL89
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social InteractionarXiv:2603.17094Ryo Kamoi, Ameya Godbole, Binglin Zhou +211 Sept 2026cs.CL89
Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and StudyarXiv:2603.24125Nour Bouchouchi, Thibault Laugel, Xavier Renard +211 Sept 2026cs.CL89
Timing is Everything: Temporal Scaffolding of Semantic Surprise in HumorarXiv:2605.00143Yuxi Ma, Yongqian Peng, Junchen Lyu +211 Sept 2026cs.CL89
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and DistillationarXiv:2605.12227Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins11 Sept 2026cs.CL89
Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accountsarXiv:2605.21049Ni Yang, Rui He, Philipp Homan +211 Sept 2026cs.CL89
MERIT: Matching Expertise via Rubric-Informed Training for Reviewer AssignmentarXiv:2605.27865Zixuan Yang, Yibo Zhao, Weicong Liu +111 Sept 2026cs.CL89
Characterizing Narrative Content in Web-scale LLM Pretraining DataarXiv:2606.19468Teagan Johnson, Elliott Ash, Andrew Piper +111 Sept 2026cs.CL89
Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI DialoguearXiv:2606.21844William Hager, Ishika Rathi, Masum Hasan +111 Sept 2026cs.CL89
A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar BooksarXiv:2607.22376Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager +111 Sept 2026cs.CL89
Predicting Startup Exit from Textual Descriptors - A Computational Linguistics FrameworkarXiv:2608.00045Alberto M. G. Saruggia, Sebastien Germano11 Sept 2026cs.CL89
Causal Episodic Memory for Feedback-Driven Agent RepairarXiv:2608.05906Khang Nhat Hoang Vo, Tam Minh Chu, Anh Trac Duc Dinh +211 Sept 2026cs.CL89
VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool UsearXiv:2608.08477Juan S. Santillana11 Sept 2026cs.CL89
Self-Evolving Embodied Agents via Skill-Harness EvolutionarXiv:2608.11350Peidong Wang, Zhiming Ma, Ying Chang +211 Sept 2026cs.CL89
Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization AnalysisarXiv:2608.16379Hiwa Asadpour11 Sept 2026cs.CL89
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based ModerationarXiv:2608.22230Junyu Lu, Kaiyuan Liu, Kaichun Wang +211 Sept 2026cs.CL89
DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database CompletionarXiv:2608.22770Xuan Yao, Shuping Li, Yang Dai +211 Sept 2026cs.CL89
Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case StudyarXiv:2608.29170Zijie Zhang, Tan Lee, Yong Cao +111 Sept 2026cs.CL89
Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation RerankingarXiv:2609.00588Guangyu Chen, Boxuan Lyu, Hidetaka Kamigaito +211 Sept 2026cs.CL89
OUTLETS: Output-Length Prediction from Speculative Decoding BackbonesarXiv:2609.01068Weihuang Wen, Yingying Liu, Yichuan Liu +211 Sept 2026cs.CL89
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE InferencearXiv:2609.04895Zhenhe Wu, Yaping Jin, Qinghua Xing +211 Sept 2026cs.CL89
CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn PrescriptionarXiv:2609.09766Minjoo Kim, Sangjin Park, Seung Hwan Cho11 Sept 2026cs.CL89
SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better TeachersarXiv:2609.09999Xixian Liao, Maite Melero11 Sept 2026cs.CL89
Emergent Risks in Generative Multi-Agent SystemsarXiv:2603.27771Yue Huang, Yu Jiang, Wenjie Wang +211 Sept 2026cs.MA89
MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math TeachersarXiv:2604.04036Zhihan Guo, Yuting Lu, Jionghao Lin11 Sept 2026cs.IR89
Formalizing building-up constructions of self-dual codes through isotropic lines in LeanarXiv:2604.08485Jae-Hyun Baek, Jon-Lark Kim11 Sept 2026cs.IT89
LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial DomainsarXiv:2604.16379Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda11 Sept 2026cs.IR89
Strategic Type SpacesarXiv:2606.08297Olivier Gossner, Rafael Veiel11 Sept 2026econ.TH89
A Group-Based Resource Allocation Model for the Fractional Knapsack ProblemarXiv:2609.06470Abhinaba Chakraborty11 Sept 2026cs.DS89
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic TasksarXiv:2609.09233Wasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth +211 Sept 2026cs.AI89
CityPlanner: A Sandbox Agent for Executable Urban PlanningarXiv:2609.09578Wentao Zhang, Jingyuan Wang, Zetong Zhou +211 Sept 2026cs.AI89
Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online SafetyarXiv:2609.09735Hamed Jelodar, Amir Firouzi, Yen-Wu Lo +211 Sept 2026cs.AI89
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a ModelarXiv:2609.09815Xing Zhang, Guanghui Wang, Yanwei Cui +211 Sept 2026cs.AI89
OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology NormalizationarXiv:2609.10055Jie Song, Zhichuan Xu, Ziyu Lu +211 Sept 2026cs.AI89

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Each paper page shows the abstract, related models and every source snapshot.