Skip to content
AI Atlas
PaperActive

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

arxiv.org/abs/2609.10862

quality89

Updated 2 h ago · first seen 11 Sept 2026

paper_01M294FQ8C7FKQR7DEY0VF34WC

Published
11 Sept 2026
T1 · 2 h ago
arXiv
2609.10862
T1 · 2 h ago
Category
cs.SD
T1 · 2 h ago

Abstract

This report presents results from Project Qualia, an ongoing effort to determine whether experiential similarity between songs, a structure not captured by genre or metadata taxonomies, can be recovered from real listening behavior. We constructed a large-scale dataset of listening sessions, comprising 1.29 billion scrobbles collected from 9,396 users via the Last.fm API and reduced through a preprocessing pipeline to 531.6 million training scrobbles across 28.6 million sessions. On this corpus, we trained a skip-gram Word2Vec model (Song2Vec), treating each session as a sentence and each track as a token. As anticipated, the resulting embedding space was dominated by artist identity, a consequence of single-artist runs within sessions. To test for a subtler, artist-independent signal, we developed an artist-residual procedure: subtracting each artist's centroid from its tracks' embeddings and evaluating whether the remainder retained structure. Mean cross-artist cosine similarity fell from 0.2487 in raw embedding space to 0.0005 in residual space, yet 4,577 cross-artist track pairs retained cosine similarity $\ge 0.70$ in residual space, forming coherent genre- and era-based clusters, including trip-hop, 1990s grunge, 2020 mainstream pop, and cross-composer classical piano pairs at cosine similarity up to 0.95. These results confirm that the training data contains experiential structure independent of artist identity, establishing an empirical basis for an architecture designed to learn this experiential layer directly.

Authors 3

Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2609.10862

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.SD, cs.IR, cs.LG

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.SD

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None