Skip to content
AI Atlas
PaperActive

HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition

arxiv.org/abs/2609.10917

Updated 51 min ago · first seen 11 Sept 2026

paper_01M294H1R6AFHA679NM9C9JK5W

Published
11 Sept 2026
T1 · 51 min ago
arXiv
2609.10917
T1 · 51 min ago
Category
cs.CV
T1 · 51 min ago

Abstract

Texture recognition remains challenging for modern vision models because discriminative evidence is often carried by higher-order spatial statistics rather than by object shape alone. While Vision Transformers provide strong long-range modeling capacity, their standard object-centric representations do not explicitly expose such statistical structure, which limits texture sensitivity in fine-grained recognition settings. We present HiPerViT, a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based recognition pipeline. The method combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. This design enables direct interaction between spatial tokens and second-order feature co-occurrence statistics, providing the model with explicit access to texture-relevant information without requiring multimodal pretraining or ensemble construction. Across six texture recognition benchmarks, HiPerViT achieves consistent improvements over strong vision-only baselines under the reported evaluation protocols, including gains of +3.05 percentage points on DTD, +10.48 on GTOS-Mobile, and +10.10 on 1200Tex. Beyond benchmark performance, our analyses show that these gains are largely invariant to the backbone depth used to extract second-order statistics and to the ordering of interaction and distillation stages. This pattern suggests that the primary source of improvement is not a specific fusion topology, but the explicit availability of second-order statistical information as a first-class representational signal. These results support explicit statistical tokenization as an effective and robust design principle for texture-centric visual recognition.

Authors 2

Jo\~ao Pedro C. A. de S\'a, Odemir Martinez Bruno

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

arXiv id
2609.10917

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Categories
cs.CV

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Primary category
cs.CV

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 51 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

51 min ago

Conflicts

None