Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages; model links come from model cards citing the paper.

210 papers

Papers
TitleAuthorsOrganizationPublishedIntroduces Models (and artifacts) whose model card or documentation cites this paper — inbound described_by relations.DatasetsBenchmarksCode
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoEarXiv:2609.09793cs.CRYi Shi, Tanyu Chen, Kai Shen11 Sept 2026
Strangers to Themselves: What Language Models Say About Themselves Is GenericarXiv:2609.09899cs.LGPhil Blandfort, Urja Pawar11 Sept 2026
Improving Cross-Lingual Token Representations by Adding a Pinch of SALTarXiv:2609.09953cs.CLGuillem Ram\'irez11 Sept 2026
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk RuntimesarXiv:2609.10016cs.LGRemco Hendriks (Continker)11 Sept 2026
Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-TrainingarXiv:2609.10052cs.CLJunwon Ko, Dong-Jae Lee, Minchan Kwon +211 Sept 2026
NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic EnvironmentsarXiv:2609.10058cs.SDNiramay M. Patel, Bibek Behera, Raksha Sharma11 Sept 2026
Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-TuningarXiv:2609.10142cs.CLJing Guan, Yachao Yang, Zhaoliang Liu +211 Sept 2026
LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented GenerationarXiv:2609.10239cs.IRDaniel Alejandro Coll Tejeda, Pedro Garc\'ia L\'opez, Daniel Barcelona-Pons11 Sept 2026
DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMsarXiv:2609.10253cs.CLBhuvan Arora, Devesh Saraogi, Sravya Varada +111 Sept 2026
GANDR: Claim Auditing for Verifiable Legal Answer GenerationarXiv:2609.10293cs.CLChen Qian, Yimeng Wang, Yu Chen +211 Sept 2026
RiLM: Parameter-Efficient Language Modeling via Geodesic DecodingarXiv:2609.10305cs.CLFang Li11 Sept 2026
Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy OperationalizationarXiv:2609.10410cs.CLAyan Majumdar, Shounak Paul, Pushpdeep Singh +211 Sept 2026
IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model IdentifierarXiv:2609.10494cs.CLBlake Stenstrom, Charangan Vasantharajan, Brian Sathianathan11 Sept 2026
Cultural Binding Heads in Language ModelsarXiv:2605.28543cs.AIAvrile Floro, Luca Benedetto11 Sept 2026
FrontierChallenge: Evaluating Scientific Workflow CompletionarXiv:2608.24979cs.AILiangcai Su, Zhaopeng Feng, Zhuo Chen +211 Sept 2026
BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language ModelsarXiv:2408.10608cs.CLYongxin Deng (University of Technology Sydney), Xiaoyu Tan (National University of Singapore), Jing Pan (Monash University) +211 Sept 2026
MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data GenerationarXiv:2510.05124cs.CLMingjin Li, Yu Liu, Huayi Liu +211 Sept 2026
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM JudgesarXiv:2601.08654cs.CLYihan Hong, Huaiyuan Yao, Bolin Shen +211 Sept 2026
Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space VisualizationarXiv:2601.15295cs.HCYi Wang, John Joon Young Chung, Melissa Roemmele +211 Sept 2026
Revisiting the Shape Convention of Transformer Language ModelsarXiv:2602.06471cs.CLFeng-Ting Liao, Guan-Ting Yi, Tzu-Quan Lin +211 Sept 2026
False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UKarXiv:2602.13047cs.CLMadhurananda Pahar, Caitlin Illingworth, Dorota Braun +211 Sept 2026
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement LearningarXiv:2604.10701cs.LGZikang Shan, Han Zhong, Liwei Wang +111 Sept 2026
Where is the Mind? Persona Vectors and LLM IndividuationarXiv:2604.17031cs.CLPierre Beckmann, Patrick Butlin11 Sept 2026
"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer UsearXiv:2605.00497cs.HCShardul Sapkota, Matthew J\"orke, Zane Sabbagh +211 Sept 2026
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsarXiv:2605.13841cs.SDTara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz +211 Sept 2026

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Linked models, datasets, benchmarks and code come from stated relations only; a dash means no source stated one.