Skip to content
AI Atlas
PaperActive

Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss

arxiv.org/abs/2609.11029

Updated 23 min ago · first seen 11 Sept 2026

paper_01M294FQGCGE0MEDTNDEW4V67Z

Published
11 Sept 2026
T1 · 23 min ago
arXiv
2609.11029
T1 · 23 min ago
Category
cs.CL
T1 · 23 min ago

Abstract

Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the tendency to memorize surface-level text spans. To address this, we present an information-weighted cross-entropy loss that rescales token-level contributions using TF-IDF statistics, emphasizing semantically informative tokens while down-weighting ubiquitous ones. Experiments on five decoder-only LLMs ranging from 1.1B to 13B parameters show consistent reductions in memorized substring length while preserving perplexity and downstream task performance. Under LoRA fine-tuning, TF-IDF reduces average substring memorization length by 14% across all five models. Under full-weight fine-tuning on TinyLLaMA 1.1B, the reduction reaches 58%. Our approach is architecture-agnostic and can be incorporated into existing training pipelines with less than 3% computational overhead, offering a lightweight and principled way to mitigate memorization without disrupting standard training dynamics.

Authors 3

Zhijian Li, Stefan Larson, Kevin Leach

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

arXiv id
2609.11029

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Categories
cs.CL, cs.LG

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

23 min ago

Conflicts

None