Skip to content
AI Atlas
PaperActive

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

arxiv.org/abs/2609.10036

quality89

Updated 3 h ago · first seen 11 Sept 2026

paper_01M294GKDWVDPP1BSQJ9VFH2FR

Published
11 Sept 2026
T1 · 3 h ago
arXiv
2609.10036
T1 · 3 h ago
Category
cs.AI
T1 · 3 h ago

Abstract

Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation can collapse their uncertainty onto the wrong hypothesis. Policies drift as the history grows. We trace these symptoms to a common structural cause. An LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state. We propose an architectural fix. The Belief-State Engine (BSE) is an inference module placed outside the LLM. It maintains a Bayesian posterior over the latent states of a given POMDP (Partially Observable Markov Decision Process) model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is not shown. We set out a minimal four-axiom specification of what a belief-consistent internal state must satisfy, and prove that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP. It therefore inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history. We evaluate the architecture on the Tiger POMDP and a red-team attack-graph task, against six baselines: a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP. Across both domains, the BSE-augmented agent improves task return, belief calibration, and decision consistency. Ten targeted ablations isolate the contribution of each architectural choice confirms that the effect is not specific to any one model. Code, environment specifications, prompt templates, and seed logs accompany this paper.

Authors 2

Arnab Chattopadhayay, Debdipta Halder

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

arXiv id
2609.10036

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Categories
cs.AI, cs.LG, cs.RO

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 3 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

3 h ago

Conflicts

None