NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
Updated 49 min ago · first seen 11 Sept 2026
paper_01M294G4AZZMH57A5M561GXW5D
- Published
- 11 Sept 2026
- T1 · 5 h ago
- arXiv
- 2609.10715
- T1 · 5 h ago
- Category
- cs.CL
- T1 · 5 h ago
Abstract
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.
Authors 28
NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng, Xu Cheng, Beiya Dai, Yufan Feng, Kewen Ge, Ruijun Ge, Jiayi Huang, Yang Jiao, Dahua Lin, Zhouhan Lin, Yifan Liu, Yuliang Liu, Biqing Qi, Mowen Ruan, Junzhe Shen, Yunchong Song, Hao Sun, Zhongbo Tian, Yixuan Wang, Rubin Wei, Jiaxin Xiong, Kangyu Yang, Qian Yao, Qi Zhang, Bowen Zhou
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- arXiv id
- 2609.10715
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Categories
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Hf paper url
- https://huggingface.co/papers/2609.10715
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 5 h agomedium
- Hf comments
- 2
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 49 min agomedium
- Upvotes
- 177
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 49 min agomedium
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 5 h agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
12
Source tiers
T1T29 / 3
Freshest observation
49 min ago
Conflicts
4 flagged
No models linked to this paper yet.
- Authors
- NCP Team, Jiaqi Cao, Chiyu Chen
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · arXiv id
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.10715 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperNCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
New paper: NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 4 h ago | 1 |
| Hugging Face Hub (public pages, model cards, papers) | huggingface.co/papers | listing | T2· Quality secondary | 49 min ago | 4 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.