News
L-HET is a corpus of two-person spoken dialogues recorded from scratch, with affect annotated at utterance, trajectory and dialogue level. L-HESS, built on it, evaluates whether a model tracks a speaker's emotional state across a full conversation — including gradual drift and masking. Phase 1 is in progress; samples are available on request.
We identify meaningful human actions in large-scale video, filter for clear hand–object interactions, and turn them into training-ready data for imitation learning and manipulation foundation models.
Bidirectional audio with per-speaker channels and timestamped transcripts across French, German, Italian, Japanese, Korean, Portuguese and more — built for speech-to-speech and spoken dialogue models.
Gameplay recordings with action–state alignment, temporal event annotation and player behavior data, for learning state transitions, planning and long-horizon reasoning.