Updated 29 h ago · first seen 15 Sept 2026
paper_01M2JK193F9VF1QQPY9NAQMD09
Abstract
The agent-centric general value function (ACGVF) construction of \citet{tasse2026goal} lets the agent make two decisions that are normally imposed by the environment or agent designer: which goal to pursue and when to declare a goal as finished (in addition to choosing the action). This is a very general framework that subsumes almost all prior work on reinforcement learning, control and planning, as well as more general formalisms proposed in the cognitive sciences. However, it assumes the environment is fully observed, i.e., that the observation is a sufficient statistic. In \citet{murphy2025rl}, a general agent design was proposed where the policy is based on an internal belief state $z_t$ and an internal goal; however, the goals were assumed to be externally provided. In this note, we unify and extend these two approaches using the formalism of hierarchical hidden Markov models (HHMM) \citep{murphy2001hhmm}.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 1
New paper: A note on goal-based hierarchical RL
arxiv
Sources
Sources 1
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.