Skip to content
AI Atlas

An Evolutionary Computation Framework for Multi-Agent Q-Learning with Mean-Field Environmental Feedback

Published 16 Sept 2026arXiv:2609.13253

data quality89

Updated 12 h ago · first seen 15 Sept 2026

paper_01M2JK197AXG3RCDRBYYDKXRMY

Abstract

Multi-agent reinforcement learning in networked populations is governed by the interaction between individual adaptation, local encounters, and changing environmental conditions. To study this interaction, we formulate a coupled learning--environment model in which agents update stateless $Q$-values on a fixed graph, while their population-average behavior drives an environmental variable that dynamically modifies the payoff matrix. Under a first-order mean-field closure, we derive a deterministic transport equation for the population distribution of $Q$-values and couple it with a projected discrete update for the environmental state. The resulting model is evaluated against finite-network Monte Carlo simulations on random regular, Erd\H{o}s--R\'enyi, Barab\'asi--Albert, and random geometric graphs. Across the tested parameter ranges, the mean-field system reproduces the main macroscopic cooperation and environmental trajectories, and the trajectory-level root-mean-square error generally decreases with population size and average degree. The analysis further shows that environmental feedback reshapes the learned action-value ordering, while reinforcing feedback can produce pronounced dependence on the initial learning bias and resource level. The environmental timescale also plays an important role: a rapid response can drive the resource state to a boundary before learning adapts, whereas a slower response preserves the interaction between behavioral learning and environmental recovery. These results provide a population-level description of coupled reinforcement learning and environmental dynamics and characterize the performance of the mean-field approximation within the tested network and parameter ranges.

Authors

Authors 3

Lichen WangLinjie LiuShijia Hua

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 0

No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 2

Full timeline →

Sources

Sources 1

Source documents
SourceDocumentTypeTierLast observedSnapshots
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.AI feedT1· Official12 h ago5

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.