Updated 42 min ago · first seen 11 Sept 2026
paper_01M294G5CAWGFDPM5X597E8BC2
- Published
- 12 Sept 2026
- T1 · 42 min ago
- arXiv
- 2609.11864
- T1 · 8 h ago
- Category
- eess.AS
- T1 · 8 h ago
Abstract
Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance reasoning capabilities without inducing prohibitive delays, an inherent accuracy-latency trade-off persists. In this paper, we investigate whether a streaming SpeechLLM can dynamically revise its reasoning traces on the fly. We introduce RetroThinker, a multi-stage post-training framework that equips the Moshi model to self-verify and forward-correct CoT steps during inference. RetroThinker combines supervised fine-tuning (SFT) on curated retrospective thinking data with length-based direct preference optimization (DPO) to optimize retrospective during early reasoning (i.e., reasoning concurrently while the user speaks). Evaluated on the GSM8K benchmark, RetroThinker significantly improves the accuracy-latency trade-off over non-retrospective baselines, achieving an 11% absolute accuracy gain at a comparable latency.
Authors 4
Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Arxiv announce type
- cross
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- arXiv id
- 2609.11864
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Categories
- eess.AS, cs.AI, cs.CL
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Primary category
- eess.AS
Source:arXiv (Atom API + RSS)T1observed 8 h agohigh
- Published
- 12 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 42 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
42 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Arxiv announce type
Arxiv announce typearxiv_announce_type1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| cross | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
RetroThinker: Enabling Retrospective Thinking in Speech LLMs: published at changed from 2026-09-11T04:00:00+00:00 to 2026-09-12T04:00:00+00:00
Published11 Sept 2026→12 Sept 2026arxivNew paper: RetroThinker: Enabling Retrospective Thinking in Speech LLMs
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 42 min ago | 2 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 43 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.