News/Dataset & Benchmark
We introduce L-HET (Long-Horizon Emotional Trajectory), a corpus of two-person spoken dialogues recorded from scratch with multi-level affect annotation, and L-HESS, a benchmark built on it. Together they evaluate whether a model can track how a speaker's emotional state evolves over a full conversation — through shifts, gradual drift and masking — rather than classify emotion at a single moment.
At a glance
Sample Access
L-HET samples are available for evaluation on our password-protected preview site. Submit a request and we will send the link and password by email after review.
Emotion recognition is usually evaluated one utterance at a time. In real conversation, however, a speaker's emotional state is revised as the dialogue unfolds: it shifts in response to events, drifts gradually, and is at times deliberately masked. A model that labels isolated utterances correctly may still fail to maintain a coherent account of the speaker over time.
L-HET and L-HESS move evaluation from static classification to dynamic state tracking and revision, under incomplete and conflicting evidence. They target four capabilities:
| Capability | Definition |
|---|---|
| Initial affect judgment | Estimating the speaker's affective state from the opening of the dialogue |
| Dynamic revision | Updating that estimate as new evidence arrives |
| Masking recognition | Detecting when expressed and underlying emotion diverge |
| Holistic state modeling | Integrating the whole dialogue into one consistent account of the speaker |
Dialogues are recorded from scratch using an improvisational heuristic design. Instead of a fixed script, participants receive a scenario background, an initial emotional state and a set of trigger events. This keeps the speech natural while making emotional arcs controllable and annotatable.
Each dialogue is kept whole — 6–10 minutes, typically 10–15 turns — and delivered as full, unsegmented audio with transcripts and timestamps. Preserving the complete temporal context prevents models from relying on shortcuts available in isolated utterances.
Affect is annotated at three levels, supporting interpretable error analysis and trajectory-consistency scoring.
| Level | Scope |
|---|---|
| Utterance / turn | Affect expressed in each utterance or turn |
| Trajectory | How affect evolves across turns |
| Global | The speaker's overall affective state across the dialogue |
The model receives the complete, unsegmented dialogue audio and generates a natural, empathetic and context-consistent spoken response to the final utterance.
The model is never prompted to output emotion labels. An appropriate response is only possible if the speaker's state has been tracked across the whole dialogue, so the task measures long-horizon affect tracking implicitly.
Tiers allow results to be decomposed by capability and support curriculum-style evaluation.
| Tier | Characteristic |
|---|---|
| A | Explicit emotional shift |
| B | Natural drift, with vocal leakage |
| C | Long-term masking and multiple transitions |
L-HET and L-HESS are produced through a six-stage pipeline. The same pipeline is reusable for targeted data generation and for expanding evaluation sets.
01
Scenario design
02
Audio recording
03
Transcription & timestamps
04
Multi-level annotation
05
Benchmark construction
06
Model evaluation
L-HET and L-HESS are intended for research and evaluation in:
In each case, the question is whether a system genuinely updates its understanding of the user over time.