Skip to content
AI Atlas
PaperActive

DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

arxiv.org/abs/2609.11504

Updated 30 min ago · first seen 11 Sept 2026

paper_01M294FP7MV8MMFB3HCTVNGWS7

Published
11 Sept 2026
T1 · 30 min ago
arXiv
2609.11504
T1 · 30 min ago
Category
cs.LG
T1 · 30 min ago

Abstract

A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a fixed 5% price-impact cap. A slippage bound derived from a quote does not prevent the price impact of the order itself. We propose Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters. On 75 held-out workflow prompts, its hybrid variant scores 0.67 on the static safety proxy, compared with 0.33 for the best baseline. Koan-Safe records no unsafe executions on the saved benchmark outputs. A matched-candidate ablation produces 14-17 unsafe executions when enforcement is disabled. Additional tests expose the limits of default injection: permissive existing thresholds can still authorize unsafe trades. A separately evaluated policy cap addresses this failure on a 36-case diagnostic grid. These results support explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.

Authors 4

Abhinav Rajeev Kumar, Harshit Arora, Varun Singh, Manikandan Nanjappan

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

arXiv id
2609.11504

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Categories
cs.LG, cs.SE

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 30 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

30 min ago

Conflicts

None