TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation
Published 18 Sept 2026arXiv:2609.20414
Updated 4 h ago · first seen 18 Sept 2026
paper_01M2SEH09PBB2KAKHYZDXKHQXS
Abstract
Tactile signals provide direct contact and force measurements that are essential for understanding physical interactions and enabling dexterous robotic manipulation. However, tactile sensing requires direct measurement at contact interfaces, making large-scale data collection reliant on intrusive, costly, and restrictive instrumentation. We present TouchSight, a monocular egocentric vision framework for dense full-hand contact force prediction that leverages 500 hours of pressure-glove recordings and extensive hand-object interaction (HOI) data. To address the appearance gap between gloved training data and bare-hand real-world scenarios, we construct TwinTouch-20H: 20 hours of paired visual data in which generative video models re-render gloved recordings as bare-hand observations against new backgrounds while preserving the original measured tactile labels. TouchSight predicts dense force from both gloved and generated bare-hand videos, outperforms prior contact prediction methods on OakInk2, qualitatively generalizes to natural bare-hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales. These results demonstrate that dense tactile signals can be recovered from egocentric vision alone, without tactile instrumentation at capture time.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperTouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation
TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperTouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation
New paper: TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.