Skip to content
AI Atlas
PaperActive

Statistical analysis of Inverse Entropy-regularized Reinforcement Learning

arxiv.org/abs/2512.06956

quality89

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294FSYMAQYJQNEMQAW7Z04A

Published
11 Sept 2026
T1 · 3 h ago
arXiv
2512.06956
T1 · 3 h ago
Category
stat.ML
T1 · 3 h ago

Abstract

-cross Abstract: Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward: many reward functions can induce the same optimal policy, rendering the inverse problem ill-posed. In this paper, we develop a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves this ambiguity by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. This combination yields a unique and well-defined so-called least-squares reward consistent with the expert policy. We model the expert demonstrations as a Markov chain with the invariant distribution defined by an unknown expert policy $\pi^\star$ and estimate the policy by a penalized maximum-likelihood procedure over a class of conditional distributions on the action space. We establish high-probability bounds for the excess Kullback--Leibler divergence between the estimated policy and the expert policy, accounting for statistical complexity through covering numbers of the policy class. These results lead to non-asymptotic minimax optimal convergence rates for the least-squares reward function, revealing the interplay between smoothing (entropy regularization), model complexity, and sample size. Our analysis bridges the gap between behavior cloning, inverse reinforcement learning, and modern statistical learning theory.

Authors 4

Denis Belomestny, Alexey Naumov, Artemy Rubtsov, Sergey Samsonov

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

arXiv id
2512.06956

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Categories
stat.ML, cs.LG, math.ST, stat.TH

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Primary category
stat.ML

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

3 h ago

Conflicts

None