News/Dataset & Benchmark

Introducing L-HET and L-HESS: A Dataset and Benchmark for Long-Horizon Emotional Understanding in Spoken Dialogue

We introduce L-HET (Long-Horizon Emotional Trajectory), a corpus of two-person spoken dialogues recorded from scratch with multi-level affect annotation, and L-HESS, a benchmark built on it. Together they evaluate whether a model can track how a speaker's emotional state evolves over a full conversation — through shifts, gradual drift and masking — rather than classify emotion at a single moment.

At a glance

Components
L-HET (dataset) · L-HESS (benchmark)
Data
Two-person natural spoken dialogues; full unsegmented audio with transcripts and timestamps
Dialogue length
6–10 minutes, typically 10–15 turns
Scale (Phase 1)
150 dialogues, approx. 20–25 hours (planned)
Annotation
Three levels: utterance/turn, trajectory, global
Status
Phase 1 in progress; first 50 dialogues recorded and annotated
Access
Samples available on request

Sample Access

Request access to the samples

L-HET samples are available for evaluation on our password-protected preview site. Submit a request and we will send the link and password by email after review.

Access is granted per request after review. Your details are used only to respond to this request.

Motivation

Emotion recognition is usually evaluated one utterance at a time. In real conversation, however, a speaker's emotional state is revised as the dialogue unfolds: it shifts in response to events, drifts gradually, and is at times deliberately masked. A model that labels isolated utterances correctly may still fail to maintain a coherent account of the speaker over time.

L-HET and L-HESS move evaluation from static classification to dynamic state tracking and revision, under incomplete and conflicting evidence. They target four capabilities:

CapabilityDefinition
Initial affect judgmentEstimating the speaker's affective state from the opening of the dialogue
Dynamic revisionUpdating that estimate as new evidence arrives
Masking recognitionDetecting when expressed and underlying emotion diverge
Holistic state modelingIntegrating the whole dialogue into one consistent account of the speaker

Dataset: L-HET

Collection protocol

Dialogues are recorded from scratch using an improvisational heuristic design. Instead of a fixed script, participants receive a scenario background, an initial emotional state and a set of trigger events. This keeps the speech natural while making emotional arcs controllable and annotatable.

Structure

Each dialogue is kept whole — 6–10 minutes, typically 10–15 turns — and delivered as full, unsegmented audio with transcripts and timestamps. Preserving the complete temporal context prevents models from relying on shortcuts available in isolated utterances.

Annotation

Affect is annotated at three levels, supporting interpretable error analysis and trajectory-consistency scoring.

LevelScope
Utterance / turnAffect expressed in each utterance or turn
TrajectoryHow affect evolves across turns
GlobalThe speaker's overall affective state across the dialogue

Benchmark: L-HESS

Task

The model receives the complete, unsegmented dialogue audio and generates a natural, empathetic and context-consistent spoken response to the final utterance.

Design rationale

The model is never prompted to output emotion labels. An appropriate response is only possible if the speaker's state has been tracked across the whole dialogue, so the task measures long-horizon affect tracking implicitly.

Difficulty tiers

Tiers allow results to be decomposed by capability and support curriculum-style evaluation.

TierCharacteristic
AExplicit emotional shift
BNatural drift, with vocal leakage
CLong-term masking and multiple transitions

Construction

L-HET and L-HESS are produced through a six-stage pipeline. The same pipeline is reusable for targeted data generation and for expanding evaluation sets.

  1. 01

    Scenario design

  2. 02

    Audio recording

  3. 03

    Transcription & timestamps

  4. 04

    Multi-level annotation

  5. 05

    Benchmark construction

  6. 06

    Model evaluation

Intended Use

L-HET and L-HESS are intended for research and evaluation in:

  • Emotional agents
  • Voice companions
  • Supportive dialogue research
  • Real-time multimodal assistants
  • Affect-aware evaluation

In each case, the question is whether a system genuinely updates its understanding of the user over time.