Human first-person household demonstrations with hand-action semantics, task decomposition and object-interaction labels. Controlled shooting + teleoperation, consent-based and desensitized.
1,200h+Hours
60+Task families
RLDSFormat
First-person kitchen operations across prep / cooking / cleaning task chains, action sequences aligned with language instructions.
8 类Scene families
AlignedLang-aligned
LeRobotFormat
First-person indoor movement with spatial-relation labels, path semantics and obstacle events — for spatial understanding and navigation training.
Multi-homeEnvironments
SpatialRelations
RLDSFormat
Cinema-grade clips with dense scene / visual / narrative / feature / sentiment labels, key-frame accuracy >90%* — first-choice corpus for video foundation models.
200万+Clips
DenseLabels
Licensed3-layer license
Documentary-grade real-physics footage covering natural motion, material interaction and lighting — high-fidelity corpus for world models.
500万+Clips
RealReal physics
IP 链Traceable
Real human actions curated from a 2.3B+ clip pool with action semantics and temporal labels — scaled ammunition for the video-distillation track.
2.3B+Pool size
120+Task families
时序动作Labels
Multi-cam synchronized 4D video, cm-level spatial alignment*, with action semantics and spatial-relation labels.
50,000h+Reserve
cm-levelAlignment*
DaaSDaaS
Real-robot teleoperation end-effector trajectories aligned frame-level with vision — for VLA post-training and imitation learning.
60+Task families
Frame-levelVision-aligned
RLDSFormat
Creator, post and comment data across 10 regions with sentiment labels, ASR transcripts and OCR fields — fully PII-desensitized.
30+Languages
10 区域Platforms
100%PII Desensitized
SKU-level image-text alignment across major platforms — for multimodal retrieval and generative commerce.
2.1B+SKU images
Img-textPairs
ParquetFormat
Product, review and search corpora across major cross-border platforms, 30+ languages, SFT / RAG ready.
30+Languages
200+Countries
JSONLFormat
High-quality review corpora curated from a 23B+ text pool, human × AI cleaned, with stance and sentiment labels.
23B+Pool size
SentimentLabels
JSONLFormat
Can't find the dataset you need?
Beyond the catalog, ENDATA customizes capture and processing by scene family, task family and format — first delivery in as fast as 6 weeks.
* Figures are product metrics · internal evaluation, illustrative, as of 2026Q1; every dataset ships with licensing docs and a free sample.