Applied AI Research Lab
Data, Benchmarks & Evaluations
We co-research with frontier AI labs to find where models fail to understand humans, and build the data, benchmarks and evaluations that fix it.
50B+
Data Assets
30+
Languages
30+
AI Enterprise Clients
Our Mission
Models still misread people — what they mean, how they speak, how they act. Working alongside frontier AI labs, we find where that understanding breaks down, then build the data that closes the gap and the benchmarks and evaluations that show it has closed.
01
Find
Co-research with frontier AI labs to locate where models fail to understand humans
02
Measure
Turn each failure into a benchmark, so the gap can be measured
03
Build
Collect and annotate targeted data across 30+ languages and real-world environments
04
Evaluate
Verify on held-out evaluations that the gap has actually closed
AI Data
The data behind our research falls into two lines: how people speak and see, and how they act in the physical world.
01
What the world looks, sounds and reads like.
Audio, image and video at web scale — for LLMs, multimodal and generative models.
02
How the world responds to action.
Real-world interaction data — for embodied AI and world models.
What We Do
01
Joint programmes with frontier AI labs to find where models fail to understand humans.
02
Benchmarks and evaluation sets that make human-understanding gaps measurable.
03
Off-the-shelf data across both data lines, with samples and a datasheet for evaluation.
04
Targeted data, collected and annotated for the gap you need to close.
Why SEMO AI
Finding a failure is only half the work. We also build what fixes it.
01
We work alongside research teams at frontier AI labs, on the failures that matter to the models they are training now.
02
Our work centres on the human side of AI: intent, speech, conversation, emotion and physical behaviour.
03
Research findings become training-ready data on the same infrastructure that serves 30+ AI companies.
News
L-HET is a corpus of two-person spoken dialogues recorded from scratch, with affect annotated at utterance, trajectory and dialogue level. L-HESS, built on it, evaluates whether a model tracks a speaker's emotional state across a full conversation — including gradual drift and masking. Phase 1 is in progress; samples are available on request.
We identify meaningful human actions in large-scale video, filter for clear hand–object interactions, and turn them into training-ready data for imitation learning and manipulation foundation models.
Bidirectional audio with per-speaker channels and timestamped transcripts across French, German, Italian, Japanese, Korean, Portuguese and more — built for speech-to-speech and spoken dialogue models.
Gameplay recordings with action–state alignment, temporal event annotation and player behavior data, for learning state transitions, planning and long-horizon reasoning.
Tell us where your models fall short, or what data you need. Our team will get back to you within 24 hours.