Skip to content
AI Atlas
PaperActive

UBCL: A Reinforcement Learning Framework for Controllable and Diverse Player Behaviors

arxiv.org/abs/2512.10835

quality89

Updated 4 h ago · first seen 11 Sept 2026

paper_01M294FRXCBS46JJAM3VSRA5XJ

Published
11 Sept 2026
T1 · 4 h ago
arXiv
2512.10835
T1 · 4 h ago
Category
cs.LG
T1 · 4 h ago

Abstract

This paper introduces a reinforcement learning framework that enables controllable and diverse player behaviors without relying on human gameplay data. Existing approaches often require large-scale player trajectories, train separate models for different player types, or provide no direct mapping between interpretable behavioral parameters and the learned policy, limiting their scalability and controllability. We define player behavior in an N-dimensional continuous space and uniformly sample target behavior vectors from a region that encompasses the subset representing real human styles. During training, each agent receives both its current and target behavior vectors as input, and the reward is based on the normalized reduction in distance between them. This allows the policy to learn how actions influence behavioral statistics, enabling smooth control over attributes such as aggressiveness, mobility, and cooperativeness. A single PPO-based multi-agent policy can reproduce new or unseen play styles without retraining. Experiments conducted in a custom multi-player Unity game show that the proposed framework produces significantly greater behavioral diversity than a win-only baseline and reliably matches specified behavior vectors across diverse targets. The method offers a scalable solution for automated playtesting, game balancing, human-like behavior simulation, and replacing disconnected players in online games.

Authors 2

Atahan Cilan, Atay \"Ozg\"ovde

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

arXiv id
2512.10835

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Categories
cs.LG

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

DOI
10.1109/TG.2026.3703355

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Primary category
cs.LG

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 4 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

10

Source tiers

T110

Freshest observation

4 h ago

Conflicts

None