Skip to content
AI Atlas
PaperActive

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

arxiv.org/abs/2609.10815

Updated 22 min ago · first seen 11 Sept 2026

paper_01M294G51H9MTWD811DDZHSJT0

Published
11 Sept 2026
T1 · 23 min ago
arXiv
2609.10815
T1 · 23 min ago
Category
cs.CV
T1 · 23 min ago

Abstract

Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data remains a significant challenge due to its multimodal video format. BWC videos, in many cases, comprise chaotic scenes with low visual quality, rapid movement/interactions, and high-noise audio that make visual understanding a challenge for even SOTA multimodal models. Current Vision-Language Models (VLMs) frequently overlook critical forensic details, such as the presence of valuable evidence or the latent nuances of suspect-officer interactions, which are vital for fair legal outcomes and civilian/officer safety. To address these limitations, we propose an Adaptive Visual Question Answering (VQA) framework engineered for high-stakes law enforcement. Our framework employs a structured reasoning approach to extract fine-grained visual evidence that traditional captioning systems fail to capture. We experiment with multiple question generation models, including foundation models and fine-tuned open-weight models, to observe performance variation among question generation model implementations. Our results demonstrate that this VQA-driven architecture provides a more reliable, objective, and detailed record of enforcement events, ultimately serving as a powerful tool to protect both law enforcement officers and the public through AI-assisted forensic clarity.

Authors 9

Karish Gupta, Matthew Alex, Alex Li, Yang Wu, Yun-Wei Chu, Kashif Munir, Xiaotian Zhou, Zhengping Ji, Xiaozhong Liu

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 22 min agohigh

arXiv id
2609.10815

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Categories
cs.CV, cs.CL

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Primary category
cs.CV

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

22 min ago

Conflicts

None