Skip to content

Matha Labs / Dataset directory

Human activity. Machine understanding.

A curated starting point for models that see, hear, understand, and act. Explore original dataset sources for egocentric activity recognition, robotics, images, audio, and language.

See the working example
15curated sources
5research areas

Links reviewed

These are independently published datasets linked from their original sources. Matha Labs curates this directory; access and usage rights come from each publisher.

Executed workflow / Public robot sample

One real episode. An inspectable data workflow.

We imported a published robot recording, aligned its camera frames with action/state records, generated candidate labels, and exported a separate synthetic signal replay. Watch each step and inspect the resulting files.

This example uses third-party, fixed-camera robot footage. It demonstrates the data workflow; it is not HumanHUD capture or a trained activity model.

real-data records
303
generated records
151
second source episode
10.1
camera views
2
01

Import a real recording

Recorded source · No new robot capture

A published pick-and-place episode with two fixed RGB cameras and synchronized six-dimensional robot action/state signals.

Watch the second camera
02

Review candidate labels

Heuristic labels · No SAM or learned inference

A simple color heuristic marks the visible pink object; changes in robot signals produce candidate motion labels. These candidates still need human review.

Download curated records
03

Export a synthetic replay

Generated signals · Separate provenance

Interpolate recorded signals into a separate generated dataset and display a schematic arm replay. This is not a calibrated robot model or physics simulation.

Download generated records

Source

LeRobot / svla_so101_pickplace

Publisher task: pink lego brick into the transparent box

Read the publisher’s dataset card

Dataflow checks

  • Frame counts: Verified
  • Timestamp alignment: Verified
  • Finite action/state values: Verified
  • Source file hashes: Verified
  • Separate generated provenance: Verified
Scope and limitations of this example
  • Third-party real robot sample; not captured by HumanHUD and not egocentric video.
  • No audio. No learned activity model, SAM inference, policy training or real robot execution occurred.
  • Dataset title says SO101; publisher robot_type metadata says so100_follower; discrepancy preserved.
  • Publisher supplies Apache-2.0 license but no separate participant-consent statement in the card.
  • Color and motion labels are unreviewed heuristic candidates; task text comes from the publisher.
  • Synthetic rows are interpolated native-unit signals; schematic video is not calibrated or physics validated.
  • One episode validates dataflow only; it does not measure model accuracy or generalization.

One episode checks the dataflow. It does not establish model accuracy, robot-policy quality, or generalization to other activities.

The directory

Find the data for your next model.

Filter by research area

15 of 15 sources shown

Egocentric

External source

Ego4D Consortium

Ego4D

First-person recordings of daily life with narrations and benchmarks for activity understanding, memory, hands and objects, and audiovisual interaction.

  • Egocentric video
  • Audio (subset)
  • Narrations
  • Activity annotations
Dataset scope
3,600+ hours of video
Useful for
Activity recognition · Episodic memory · Action anticipation
Access & use
Custom Ego4D agreements
Access and commercial-use detailsAccept publisher agreements, obtain credentials, then select subsets with the official downloader.Defined commercial model and product development is permitted under signed agreements; dataset resale and third-party redistribution are restricted.

Egocentric

External source

Ego-Exo4D Consortium

Ego-Exo4D

Synchronized first- and third-person recordings of skilled activities, paired with expert commentary, pose and procedural annotations.

  • Egocentric video
  • Exocentric video
  • Audio
  • Gaze
  • IMU
  • 3D pose
Dataset scope
V2: 1,286.3 video hours, including 221.26 egocentric hours
Useful for
Skill understanding · Activity segmentation · Cross-view learning
Access & use
Custom Ego-Exo4D agreements
Access and commercial-use detailsSign a separate Ego-Exo4D agreement and use the publisher's CLI with issued credentials.Commercial model development is covered by the agreement; redistribution, sublicensing and resale of the underlying data are restricted.

Egocentric

External source

EPIC-KITCHENS team

EPIC-KITCHENS-100

Unscripted kitchen activity viewed through a head-mounted camera, with temporal action labels and narrated interactions with everyday objects.

  • Egocentric video
  • Audio
  • Action labels
  • Narrations
Dataset scope
100 hours; 45 kitchens; about 90,000 action segments
Useful for
Action recognition · Action anticipation · Temporal detection
Access & use
CC BY-NC 4.0
Access and commercial-use detailsUse the official page's data links and download scripts. Commercial licensing enquiries go to the publisher.The public license excludes commercial use. A separate commercial license is required from the EPIC-KITCHENS team.

Egocentric

External source

University of Bristol and collaborators

HD-EPIC

Detailed everyday cooking recordings connect actions, audio events, gaze and object movements to reconstructed kitchen scenes.

  • Egocentric video
  • Audio
  • Gaze
  • 3D scenes
  • Object tracks
Dataset scope
41 hours; 59,454 actions; 50,968 audio annotations
Useful for
Fine-grained activity understanding · Video question answering · Object interaction
Access & use
CC BY-NC 4.0
Access and commercial-use detailsOfficial links provide videos, annotations, audio, gaze and scene data separately.Noncommercial under the public license; contact the EPIC-KITCHENS team for separate commercial rights.

Egocentric

External source

EgoLife / EvolvingLMMs Lab

EgoLife

Week-long daily-life recordings support long-context activity understanding, event recall and questions about personal routines.

  • Egocentric video
  • Audio
  • Captions
  • Transcriptions
  • Multiview recordings
Dataset scope
300 hours across six participants
Useful for
Long-context video understanding · Activity memory · Life-oriented question answering
Access & use
Conflicting publisher labels: MIT / S-Lab 1.0
Access and commercial-use detailsThe official Hugging Face dataset card labels the release MIT; the linked project repository uses S-Lab 1.0. Confirm terms for the specific files before use.Publisher clarification required: the S-Lab project license limits use to noncommercial purposes despite the MIT dataset-card label.

Robotics

External source

DROID collaboration

DROID

Human-teleoperated manipulation demonstrations across varied real environments, pairing camera observations with robot actions and language instructions.

  • Robot video
  • Robot state
  • Actions
  • Language instructions
Dataset scope
76,000 trajectories; 350 hours; 564 scenes
Useful for
Imitation learning · Robot manipulation · Policy generalization
Access & use
CC BY 4.0 (dataset)
Access and commercial-use detailsThe project page includes a visualizer and TensorFlow Datasets/Colab quickstart. The paper specifies the dataset license.Commercial use permitted under CC BY 4.0 with attribution and applicable license conditions.

Robotics

External source

UC Berkeley RAIL and collaborators

BridgeData V2

Robot manipulation trajectories with natural-language task instructions, useful for learning household object interactions across scenes.

  • Robot images
  • Robot state
  • Actions
  • Language instructions
Dataset scope
60,096 trajectories across 24 environments
Useful for
Language-conditioned control · Imitation learning · Goal-conditioned policies
Access & use
CC BY 4.0
Access and commercial-use detailsPublisher-hosted archives separate teleoperated demonstrations from scripted policy rollouts; training code is linked from the project.Commercial use permitted with attribution under CC BY 4.0.

Robotics

External source

Open X-Embodiment collaboration / Google DeepMind

Open X-Embodiment

A shared access format for robot-learning datasets from multiple institutions, supporting comparisons and learning across robot platforms.

  • Robot observations
  • Actions
  • Language (varies)
Dataset scope
Multiple contributed robot datasets in RLDS format
Useful for
Cross-robot learning · Policy pretraining · Dataset interoperability
Access & use
Repository materials: CC BY 4.0; code: Apache 2.0
Access and commercial-use detailsUse the official dataset spreadsheet and Colab. Select a constituent dataset and review its source terms.Repository materials allow commercial use under their licenses; verify the specific contributed dataset's terms before building a training mixture.

Audio

External source

Universities of Bristol and Oxford

EPIC-SOUNDS

Time-aligned sound-event annotations for everyday kitchen interactions, designed to identify actions from what they sound like.

  • Audio
  • Temporal sound labels
  • Paired egocentric video
Dataset scope
44 audio classes
Useful for
Sound event recognition · Audio event detection · Audio-visual activity understanding
Access & use
CC BY-NC 4.0
Access and commercial-use detailsDownload annotations from the official repository; recordings come from EPIC-KITCHENS-100.Public release is noncommercial. Contact the publisher for a separate commercial license.

Audio

External source

Google Research

AudioSet

A broad vocabulary of everyday sounds with human-labeled audio segments, useful for environmental sound and activity cues.

  • Audio features
  • Sound labels
  • YouTube segment references
Dataset scope
About 2.1 million labeled segments; 527 annotated classes
Useful for
Sound event classification · Audio representation learning · Environmental sound detection
Access & use
CC BY 4.0 data; CC BY-SA 4.0 ontology
Access and commercial-use detailsOfficial downloads provide segment IDs, labels and extracted features, rather than a freely licensed raw-audio archive.Released data/features permit commercial use with attribution. Underlying YouTube media has separate rights and terms; the ontology has share-alike conditions.

Images

External source

COCO Consortium

COCO

Everyday scenes annotated with objects, segmentation masks, captions and person keypoints for visual understanding.

  • Images
  • Object boxes
  • Segmentation masks
  • Captions
  • Keypoints
Dataset scope
330,000 images; over 200,000 labeled
Useful for
Object detection · Instance segmentation · Image captioning
Access & use
CC BY 4.0 annotations; image rights vary
Access and commercial-use detailsDownload image splits and annotations from the official site. Image copyrights remain with their owners.Annotations allow commercial use with attribution; image usage depends on individual image rights and Flickr terms.

Images

External source

Google Research and collaborators

Open Images V7

Diverse images with object labels, boxes, masks, visual relationships and localized descriptions for scene understanding.

  • Images
  • Object labels
  • Bounding boxes
  • Masks
  • Visual relationships
Dataset scope
About 9 million images
Useful for
Object detection · Segmentation · Visual relationship understanding
Access & use
CC BY 4.0 annotations; images listed as CC BY 2.0
Access and commercial-use detailsUse publisher download tools or select subsets. The publisher asks users to verify each image's license.Attribution licenses permit commercial use, but the publisher does not warrant image license status; verify selected images.

Language

External source

Hugging Face

FineWeb

Filtered and deduplicated English web text for language-model pretraining, with source metadata and a reproducible processing pipeline.

  • Text
  • Source metadata
Dataset scope
More than 18.5 trillion tokens in the current card
Useful for
LLM pretraining · Text representation learning · Data filtering research
Access & use
ODC-By 1.0; source terms also apply
Access and commercial-use detailsHugging Face supports dataset subsets, downloads and streaming. The card also binds use to Common Crawl terms.Database rights allow commercial use with attribution; ODC-By does not clear rights in every source text.

Language

External source

Allen Institute for AI (Ai2)

Dolma

A documented mix of web text, publications, code and reference material used for open language-model pretraining research.

  • Text
  • Code
  • Source metadata
Dataset scope
Trillion-token corpus; multiple versioned releases
Useful for
LLM pretraining · Corpus curation · Data mixture research
Access & use
ODC-By 1.0; source terms also apply
Access and commercial-use detailsChoose a version and use the publisher's URL manifests. Original source licenses and terms remain applicable.The database license permits commercial use with attribution, subject to the rights and terms of the original sources.

Language

External source

OpenAssistant / LAION and community

OpenAssistant OASST1

Human-written assistant conversations with quality ratings and reply trees for instruction tuning and learning from human feedback.

  • Conversation text
  • Quality ratings
  • Reply trees
Dataset scope
161,443 messages across 35 languages; 88,838 in ready export
Useful for
Instruction tuning · Reward modeling · Dialogue evaluation
Access & use
Apache 2.0
Access and commercial-use detailsHugging Face offers ready-for-export subsets and full message-tree files; preserve conversation structure when selecting training samples.Commercial use permitted under Apache 2.0, including its notice and redistribution conditions.

Match the data to the model

One directory. Different learning tasks.

Activity + world models

Start with video sequences, audio context, and temporal activity labels. The right inputs depend on your model’s architecture and training objective.

Robotics + vision

Look for demonstrations, action and state data, object labels, and camera views that match the robot or perception task you want to support.

Language + decisions

Use text corpora for language models, and labeled examples to evaluate classification and structured decisions. Training and evaluation require different data splits.

A note on Jev

TypeSafe describes Jev as a text-only decision model without customer fine-tuning. Labeled text and structured activity descriptions can support decision evaluation; raw video and audio are not direct Jev inputs.

Read TypeSafe’s model documentation

Custom collection / HumanHUD

Enterprise data, built around useful activity.

Our HumanHUD direction connects first-person activity understanding with useful assistance. We’re developing custom, permissioned collections around activities, objects, intent, and outcomes that matter to people, alongside a planned synthetic-data workflow.

  • Target activities and environments
  • Audio, video, and annotation needs
  • Participant permissions and usage rights
  • Delivery format, quality criteria, and scope

Planned collection workflows

Segment and label

Use Meta’s SAM 3.1 for prompted object masks and tracks, then combine these with activity labels and human review. Segmentation supports the labeling workflow; it does not identify every activity on its own.

Simulate and vary

Reconstruct captured human actions as tasks for simulated robot arms, then repeat and vary them to produce synthetic examples. Simulation and generated collections are in development.

Connect action to outcome

Through the planned robot SDK, connect the first-person view and instruction with a compatible robot’s permitted action and result, with participant permissions.

Meta’s SAM 3.1 project

Custom and synthetic collections are scoped with your team. The downloadable example above demonstrates a small data workflow; production collections have their own scope and quality criteria.