Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices
Published 18 Sept 2026arXiv:2609.19243
Updated 4 h ago · first seen 18 Sept 2026
paper_01M2SEG2NRJWZQDBS8MX9G3Q5S
Abstract
Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This paper presents two randomized approximations for normalized spectral co-clustering of bipartite text data when the numbers of document and word clusters may differ. The first method uses randomized SVD through random projection, while the second combines partial SVD with element-wise random sampling. Across real-world and synthetic datasets, both methods reduce runtime relative to the full-SVD baseline, but their behavior depends on matrix sparsity. The random projection method is the more reliable approximation across the tested settings, whereas the sampling-based method is most useful on denser matrices and provides limited benefit on already sparse text data. These results show that randomized approximations for spectral co-clustering should be selected according to the underlying structure of the data.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperRandomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices
Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv New paper: Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.