Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss
Updated 52 min ago · first seen 11 Sept 2026
paper_01M294FQGCGE0MEDTNDEW4V67Z
- Published
- 11 Sept 2026
- T1 · 53 min ago
- arXiv
- 2609.11029
- T1 · 53 min ago
- Category
- cs.CL
- T1 · 53 min ago
Abstract
Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the tendency to memorize surface-level text spans. To address this, we present an information-weighted cross-entropy loss that rescales token-level contributions using TF-IDF statistics, emphasizing semantically informative tokens while down-weighting ubiquitous ones. Experiments on five decoder-only LLMs ranging from 1.1B to 13B parameters show consistent reductions in memorized substring length while preserving perplexity and downstream task performance. Under LoRA fine-tuning, TF-IDF reduces average substring memorization length by 14% across all five models. Under full-weight fine-tuning on TinyLLaMA 1.1B, the reduction reaches 58%. Our approach is architecture-agnostic and can be incorporated into existing training pipelines with less than 3% computational overhead, offering a lightweight and principled way to mitigate memorization without disrupting standard training dynamics.
Authors 3
Zhijian Li, Stefan Larson, Kevin Leach
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 52 min agohigh
- arXiv id
- 2609.11029
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Categories
- cs.CL, cs.LG
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Primary category
- cs.CL
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 53 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
52 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Zhijian Li, Stefan Larson, Kevin Leach
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · arXiv id
arXiv idarxiv_id1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 2609.11029 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- Property changedPaperRebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss
Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperRebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss
New paper: Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.CL | feed | T1· Official | 52 min ago | 1 |
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 52 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.