Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages.

18 total

Reset
Papers
TitleAuthorsPublishedCategoriesOrganization / venueQuality
DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy PredictionarXiv:2609.10796Yingfan Xu, Tieming Liu, Ye Liang11 Sept 2026cs.LG
Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System DevelopmentarXiv:2609.11331Johannes M\"akelburg, Tim Schwabe, Maribel Acosta11 Sept 2026cs.LG
DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow SynthesisarXiv:2609.11504Abhinav Rajeev Kumar, Harshit Arora, Varun Singh +111 Sept 2026cs.LG
Optimizing AI Inference Across the Deployment StackarXiv:2609.10550Tejinder Singh, John Pflueger, Jeebak Mitra +211 Sept 2026cs.SE
On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical StudyarXiv:2609.10610Marius Mignard (CRIStAL), Steven Costiou (CRIStAL), Anne Etien (CRIStAL +111 Sept 2026cs.SE
Numbat: Building and Verifying a Self-Contained Machine-Learning StackarXiv:2609.10632Thang Tran (CloudKites AI Lab, New South Wales, Australia) +211 Sept 2026cs.SE
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research AgentsarXiv:2609.09219Jingjie Ning, Shanshan Zhong, Xiaochuan Li +111 Sept 2026cs.MA
Talking to Itself While Coding: What Makes Comments Help Code Generation?arXiv:2609.09242Dangfeng Pan, Zhensu Sun, Cenyuan Zhang +211 Sept 2026cs.SE
The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible AdoptionarXiv:2609.09560Sales G. Aribe Jr., Louie Jay S. Labastida11 Sept 2026cs.SE
Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database BranchesarXiv:2609.09671Kevin Hartman11 Sept 2026cs.SE
Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering designarXiv:2609.10132Vinicius Kaster Marini, Petter Krus11 Sept 2026eess.SY
A-JIT: Agentic Just-In-Time Software ConstructionarXiv:2609.10248Mark Marron, Earl T. Barr11 Sept 2026cs.SE
FrontierChallenge: Evaluating Scientific Workflow CompletionarXiv:2608.24979Liangcai Su, Zhaopeng Feng, Zhuo Chen +211 Sept 2026cs.AI
EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language RenderingarXiv:2609.08435Ziliang Zhao, Zenan Xu, Shuting Wang +211 Sept 2026cs.AI
A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision ModelarXiv:2408.02920Jingwen Zhou, Qinghua Lu, Jieshan Chen +211 Sept 2026cs.SE
Spec-Harness: Measuring and Improving Behavioral Adequacy of LLM-Synthesized Formal SpecificationsarXiv:2604.00280Md Rakib Hossain Misu, Iris Ma, Cristina V. Lopes11 Sept 2026cs.SE
SpecBench: Measuring Reward Hacking in Long-Horizon Coding AgentsarXiv:2605.21384Bingchen Zhao, Dhruv Srikanth, Yuxiang Wu +111 Sept 2026cs.SE
Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test SoftwarearXiv:2607.03215Mohamed Chahine Ghanem11 Sept 2026cs.CR

18 results

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Each paper page shows the abstract, related models and every source snapshot.