Back to search

Article

Foresight Is Not Enough: Sentence-Level Future Signals, Self-Loop Hard Negatives, and the Calibration Gap in Small Language Models

2026-07-14

Abstract excerpt

<title>Abstract</title> <p>A language model can learn a representation that appears to anticipate future text, but such a representation may still fail to control generation behavior. We study this representation-to-behavior gap in ForesightLM-v2, a DistilGPT-2-scale language model trained with a sentence-level future-prediction head and self-loop hard negatives. Across three WikiText-103 seeds, the held-out hard...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
b70fef49-8059-5aa0-9344-af8d8f7a2197
DOI
10.21203/rs.3.rs-10331341/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Foresight Is Not Enough: Sentence-Level Future Signals, Self-Loop Hard Negatives, and the Calibration Gap in Small Language ModelsDOI 10.21203/rs.3.rs-10331341/v1
Select a neighboring publication to make it the new centre.