Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages; model links come from model cards citing the paper.

210 papers

Papers
TitleAuthorsOrganizationPublishedIntroduces Models (and artifacts) whose model card or documentation cites this paper — inbound described_by relations.DatasetsBenchmarksCode
SpecBench: Measuring Reward Hacking in Long-Horizon Coding AgentsarXiv:2605.21384cs.SEBingchen Zhao, Dhruv Srikanth, Yuxiang Wu +111 Sept 2026
Tracing Computation Density in LLMsarXiv:2605.27033cs.CLCorentin Kervadec, Iuliia Lysova, Iuri Macocco +211 Sept 2026
BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti LanguagearXiv:2606.03504cs.CLMuhammad Ali11 Sept 2026
Expert-Level Crisis Detection in Mental Health ConversationsarXiv:2606.10380cs.CLGrace Byun, Abigail Lott, Rebecca Lipschutz +211 Sept 2026
DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL GenerationarXiv:2608.11889cs.DBAnik Pramanik, Murat Kantarcioglu, Vincent Oria +111 Sept 2026
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language ModelsarXiv:2608.15129cs.CLVarvara Arzt, Allan Hanbury, Terra Blevins11 Sept 2026
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-TuningarXiv:2608.16620cs.CLPeng Du, Kiran Kamble, Rakshith Vasudev +211 Sept 2026
'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech DetectionarXiv:2608.24191cs.CLFawzia Zehra (Fuzzy), Kara-Isitt, Sonal Khosla +111 Sept 2026
AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLParXiv:2608.30107cs.CLJoan Nwatu, Tsedeniya Solomon Amare, Longju Bai +211 Sept 2026
Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese WebarXiv:2609.07699cs.CLGon\c{c}alo Vinagre, Rui Pedro Guerra, Pedro Gomes +211 Sept 2026

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Linked models, datasets, benchmarks and code come from stated relations only; a dash means no source stated one.