PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
Updated 54 min ago · first seen 11 Sept 2026
paper_01M294GNGP3DBQ6KCWHCCA4GE2
- Published
- 11 Sept 2026
- T1 · 55 min ago
- arXiv
- 2609.10372
- T1 · 55 min ago
- Category
- cs.CV
- T1 · 55 min ago
Abstract
We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a QoE objective and minimizes it under quality/cost constraints. Unlike prior work on cascaded routing, semantic caching, or adaptive retrieval, PACE jointly controls which answer source composes the response and what fills the waiting window. Deployed on a humanoid-robot sales service, it combines three mechanisms: a load-adaptive cascading router, a joint path-filler controller, and volatility-aware cache admission. On 75k CarQA requests, the cascade halves pure-LLM PTFR at P95 (0.29 vs 0.53s at c16). The adaptive controller reaches 0.41s P95, outperforming RAG by 2.4 times at high load with equal quality. The filler controller cuts calls by 94% with zero conflict. Volatility-aware admission reduces stale answers from 86% to 0%. A gating rule ensures the controller never worse than the baseline, with exposure bounded by one hold period. This is the first quantification of filler-answer conflict risk in deployed services.
Authors 7
Lin Huang, Yujuan Tan, Weisheng Li, Lixiang Zeng, Kun Yang, Yongzong Wang, Suihan Xiao
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Arxiv announce type
- replace
Source:arXiv (Atom API + RSS)T1observed 54 min agohigh
- arXiv id
- 2609.10372
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Categories
- cs.CV, cs.AI, cs.RO
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Primary category
- cs.CV
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 55 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
54 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Lin Huang, Yujuan Tan, Weisheng Li
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · arXiv id
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.10372 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- Property changedPaperPACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving: arxiv announce type changed from cross to replace
Arxiv announce typecross→replacearxiv New paper: PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CV | feed | T1· Official | 54 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.AI | feed | T1· Official | 54 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.