1,200h+ human first-person household operations with hand-keypoint and task-decomposition labels.
1,200h+Hours
60+Task families
RLDSFormat
First-person kitchen operations across prep / cooking / cleaning chains, actions aligned with language instructions.
420h+Hours
8Scene families
LeRobotFormat
First-person indoor movement with spatial-relation labels, path semantics and obstacle events.
380h+Hours
Multi-homeEnvironments
RLDSFormat
First-person assembly and inspection on factory lines, with tool semantics, station transitions and safety-event labels.
260h+Hours
12Station types
RLDSFormat
First-person shelf-stocking, replenishment and checkout operations with shelf spatial relations and product-interaction labels.
180h+Hours
6Store formats
RLDSFormat
2M+ cinema-grade clips with dense scene / visual / narrative / feature / sentiment labels, three-layer traceable licensing.
2M+Clips
DenseLabels
Licensed3-layer
5M+ documentary-grade real-physics clips covering natural motion, material interaction and lighting.
5M+Clips
RealPhysics
TraceableIP chain
Real human actions curated from a 2.3B+ clip pool with action semantics and temporal labels.
2.3B+Pool size
120+Task families
TemporalLabels
YouTube long-form clips with ASR transcripts, chapter structure and multilingual subtitles across tutorial / review / documentary genres.
30+Languages
ASRAligned
ChaptersStructured
TikTok short-form content with camera-motion types, effect tags, audio-track metadata and creative-structure labels.
10Regions
CameraMotion tags
AudioMetadata
Full micro-drama episode clips with per-episode narrative beats, emotion curves and character-relation labels.
BeatsAnnotated
EmotionQuantified
LicensedContent
Livestream video segments with host-behavior, product-display and interaction-event labels across commerce / talent / gaming scenarios.
MultiStream types
EventsAnnotated
Long-formContinuous
Multi-cam synchronized 4D video, cm-level spatial alignment*, with action semantics and spatial-relation labels.
50,000h+Reserve
cm-levelAlignment*
DaaSCustom
Real-robot teleoperation end-effector trajectories aligned frame-level with vision, for VLA post-training.
60+Task families
FrameAligned
RLDSFormat
Real-robot operations with force/haptic feedback recording contact forces, torques and slip events for fine-manipulation policy training.
6-DoFForce/torque
ContactEvents
High-rateSampling
Public posts, topics and engagement data from X with sentiment labels, topic clustering and entity recognition. Fully PII-desensitized.
30+Languages
TopicsClustered
100%PII-safe
Three-tier TikTok data — creator profiles, posts and comments — with niche tags, ASR transcripts and OCR fields.
10Regions
3-tierStructure
ASR+ OCR
Instagram creator and image-post data with visual tags, hashtags and engagement metrics for multimodal alignment training.
Img-textPairs
VisualTaxonomy
EngagementMetrics
YouTube comments, subtitles and channel metadata including long-form discussion threads with stance annotation.
Long-formThreads
SubtitlesMultilingual
StanceLabeled
Public Facebook page posts and engagement across brand, media and community accounts with region and language fields.
PublicPages
Multi-regionCoverage
BrandMedia·Community
Aggregated social data across 5 platforms and 10 regions with sentiment labels, ASR transcripts and OCR fields. Fully PII-desensitized.
30+Languages
10Regions
100%PII-safe
Public discussion data from mainstream Chinese social platforms via compliant interfaces and licensed partnerships, with sentiment and topic-heat labels.
PublicDiscussions
SentimentLabeled
TopicHeat index
Amazon structured product attributes, images, text and buyer reviews with price-tracking series and rating distributions.
SKU-levelAttributes
PriceTracking
ReviewsSentiment
Shopee SEA product and review data across Indonesia / Thailand / Vietnam / Philippines with localized multilingual processing.
SEAMulti-site
MultilingualLocalized
Img-textPairs
Lazada product catalogs, category trees and reviews with cross-site price comparison and category-mapping fields.
CategoryTree
Cross-sitePricing
ReviewsStructured
Coupang Korea product and review data with Korean localization, fulfillment tags and category-attribute fields.
KoreanLocalized
CategoryAttributes
ReviewsStructured
TikTok Shop products linked to commerce content, bridging short-video content and conversion-path field structures.
ContentProduct link
ConversionPath fields
Multi-regionSites
SKU-level image-text alignment across major cross-border platforms for multimodal retrieval and product understanding.
2.1B+SKU images
Img-textPairs
ParquetFormat
Product, review and search corpora across major cross-border platforms, 30+ languages, SFT / RAG ready.
30+Languages
200+Countries
JSONLFormat
High-quality review corpora curated from a 23B+ text pool via three-stage human×AI cleaning, with stance and sentiment labels.
23B+Pool size
3-stageCleaning
JSONLFormat
Trainable public-opinion corpora with stance classification, argument-structure and emotion-intensity labels for opinion extraction and stance detection.
Stance4-class
ArgumentStructured
EmotionQuantified
Multilingual news corpora with headline-body-summary triplets, entity linking and timeline annotation for summarization and retrieval training.
TripletsH-B-S
EntitiesLinked
TimelineStructured
Film and TV review instruction corpora with title metadata, review dimensions and instruction-input-output triplets. ENDATA-owned content.
InstructionTriplets
DimensionsTagged
OwnedCopyright
RAG-oriented social knowledge base with document and chunk dual-layer structure, including source, niche and token-count fields.
Dual-layerDoc+Chunk
NicheTagged
TokenCounted
Can't find the dataset you need?
Beyond the catalog, ENDATA customizes capture and processing by scene family, task family and format — first delivery in as fast as 6 weeks.
* Figures are product metrics · internal evaluation, illustrative, as of 2026Q1; every dataset ships with licensing docs and a free sample.