Skip to content
AI Atlas
PaperActive

An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc

arxiv.org/abs/2603.15976

quality89

Updated 2 h ago · first seen 12 Sept 2026

paper_01M29X35267DMDXBXBPTV013HA

Published
12 Sept 2026
T1 · 2 h ago
arXiv
2603.15976
T1 · 2 h ago
Category
cs.AI
T1 · 2 h ago

Abstract

While LLMs have accelerated scientific code generation, comprehensively evaluating generated code remains challenging. Many benchmarks emphasize functional correctness or task completion, which is insufficient for code built on production HPC libraries, where solver selection, API conventions, memory management, parallel awareness, and performance also matter. We introduce PETSCAgent-Bench, a multidimensional benchmark and agent-based framework for assessing whether AI-generated scientific code uses a production HPC library as an expert would. A tool-augmented evaluator compiles, executes, and measures code and combines deterministic checks with LLM-based assessments in a 14-evaluator pipeline spanning five categories: correctness, performance, code quality, algorithmic appropriateness, and library-specific conventions. A2A and MCP enable black-box evaluation of compatible coding agents. Across realistic PETSc problems, frontier models generate readable, well-structured code but struggle with correctness on challenging problems and with library-specific conventions even when code compiles and runs---limitations that conventional pass/fail evaluation does not capture.

Authors 7

Barry Smith, Hong Zhang, Junchao Zhang, Le Chen, Lois Curfman McInnes, Murat Keceli, Satish Balay

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2603.15976

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None