// What’s new

Physical
Intelligence

Structured, real-world datasets for robotic perception, action understanding, and environment-aware decision-making.
// What’s new

CUA SFT
Trajectories

Expert-recorded action and reasoning data for real professional workflows in Blender, Adobe Photoshop, Excel, and PowerPoint.
// What’s new

DeepSWE Tasks

Long-horizon coding tasks mined from live production repos, calibrated for genuine difficulty, and verified through acceptance gates and full failure traceability.
200
+
Total datasets
1M+
Contributors
400
Samples launched
420
Placeholder datapoint
Explore our datasets
Egocentric Factory De-IDed
De-identified factory egocentric data from active production lines, with surrounding faces blurred to protect worker privacy while preserving the hands-on detail of real industrial work.
5 datapoints
Latest
Egocentric factory
Factory egocentric video captured on active production lines, documenting first-person, precision-critical work across assembly, inspection, soldering, testing, packaging, and material handling.
5 datapoints
Latest
Motion capture video samples
Studio-collected motion capture data built for whole-body control: full-body suits capturing clean skeletal ground truth across dexterous and diverse tasks, long-horizon sequences that preserve transitions between subtasks, and inter-agent interaction featuring handoffs, turn-taking, and shared-object coordination
5 datapoints
Latest
Vision Annotation
Covers 9 computer-vision annotation types (2D Bounding Box, Segmentation, Polygon, Lines, Points, Skeleton, 3D Bounding Box, 3D Segmentation, Scene-Level), each with labeling guidelines, a demo GIF, and sample JSON coordinates.
5 datapoints
Latest

Physical Intelligence

Structured, real-world datasets for robotic perception, action understanding, and environment-aware decision-making.
Egocentric Hand Keypoint Annotation
An egocentric dataset spanning diverse tasks and environments, annotated with detailed hand pose keypoints.
5 datapoints
Latest
Egocentric household annotated
Annotated household egocentric video with dense, multi-level labels for real-world domestic activities, from cleaning and cooking to laundry and organization. Each sequence includes temporal action chunks with time-stamped boundaries, 1 FPS natural-language hand–object descriptions, scene classification, and chunk level scene captions.
5 datapoints
Latest
Egocentric video data
High-fidelity egocentric video data capturing the complex realities of physical contact dynamics: diverse object properties, tasks, environments, and lighting, plus the spontaneous corrective behavior that defines dexterous manipulation — across both clean task execution and natural failure-and-recovery trajectories.
5 datapoints
Latest
Multi-Cam
Synchronized multi-camera manipulation data combining a first-person viewpoint with dual wrist-mounted close-ups, capturing both scene-level context and fine-grained hand–object interactions.
5 datapoints
Latest
Preference Pair Trajectory
Expert-curated preference-pair dataset derived from terminal-bench trajectories, containing paired high- and low-quality execution traces annotated through multi-layer evaluation (deterministic scoring, LLM validation, and human review), enabling training of debugging and tool-using agents via DPO and reinforcement learning.
5 datapoints
Latest

Coding Agent Eval

Benchmarks measuring code-generation and agentic coding performance.
terminal bench tasks
Expert-curated Terminal-Bench samples sprawling across diverse taxonomies and model-breaking scenarios, specifically designed to stress-test autonomous agents against ultra-dense, multi-step workflows that expose the limits of current long-horizon reasoning.
5 datapoints
Latest
Multi-Turn Instruction Following · Long Context
A multi-turn instruction-following evaluation set run on hy3 (tencent/hy3-preview), stress-testing how well the model holds and follows instructions across long-context conversations. Each data point is graded by atomic binary verifiers — single, observable facts scored strictly 0 or 1, tagged as Retrieval, Arithmetic, Reasoning, Constraint Adherence, or Completeness.
5 datapoints
Latest

Alignment

Preference and safety data for aligning model behavior with human intent.
Deep Research Agent Study- PCMB
Evaluates Deep Research Agents (DRAs) to function as autonomous Research Associates by bench-marking their ability to extract precise technical data, synthesize complex mechanisms, and logically derive conclusions & future scope.
5 datapoints
Latest

Functional Streams

Expert-curated datasets across diverse functional domains for training, evaluating, and benchmarking AI models.
LongHorizon STEM Expert Tasks
This dataset evaluates frontier models across LongHorizon trajectories in STEM workflows, assessing human-in-the-loop collaboration quality between subject matter experts and AI agents.
5 datapoints
Latest
CUA SFT - Excel Samples
A supervised fine-tuning (SFT) dataset consisting of 10–20 minute long-horizon Excel workflows performed by human operators. Each trajectory begins with opening downloaded datasets from public sources and executing complex Excel operations, including data cleaning, transformation, analysis, formatting, and workbook management.
5 datapoints
Latest
GUI grounding based code generation
Each task presents 4 screenshots of a 2-page React application shown in two states — broken and fixed — across both pages. The model must compare the before and after versions to infer the intended fixes, then generate the complete React code that reproduces the fixed application, including the workflow and navigation between the two pages.
5 datapoints
Latest
Image-gen Thinking traces
We provide high quality, detailed thinking traces and enhanced prompts with support of negative prompts to help CoT models improve on image to image editing tasks across multiple categories.
5 datapoints
Latest
Spatial Counterfactual occlusion and rearrangement evaluation (SCORE)
A benchmark evaluating models on spatial counterfactuals and rearrangement reasoning that covers counterfactual scene understanding, occlusion reasoning, and object rearrangement planning (understanding of how a scene would change after objects are moved, removed, hidden, or visually altered)
5 datapoints
Latest
Visual-Captioning
Improved the quality of AI-generated video captions through a human-in-the-loop annotation workflow. Reviewed and refined captions by verifying visual details, temporal events, and contextual accuracy, while ensuring consistency across the dataset.
5 datapoints
Latest

Multimodal

Curated multimodal datasets spanning images, videos, documents, charts, GUIs, and more for training and evaluating AI models.
Chart-to-Table SFT
Designed a Chart-to-Table SFT dataset, targeting extraction of quantitative data from charts into structured tabular representations. The objective is to enable reasoning and document intelligence capabilities for AI models
5 datapoints
Latest
CUA SFT - Powerpoint Samples
A supervised fine-tuning (SFT) dataset consisting of 15–25 minute long-horizon PowerPoint workflows performed by human operators. Each trajectory begins with acquiring content from public sources, downloading datasets, researching topics, and retrieving media and progresses through building native charts and diagrams, composing custom shapes, applying a cohesive visual theme, and sequencing animations and transitions to produce polished, presentation-ready decks.
5 datapoints
Latest
Image-gen SFT
High-quality text-to-image SFT dataset, purpose-built for training image-generation models on curated prompt–image pairs for better visual accuracy, style control, and instruction following.
5 datapoints
Latest
Multi-image reference to Video generation
Evaluated the SOTA model’s video generation capabilities using multi-image referencing, including elements, backgrounds, inspirations, and other creative inputs. This reflects one of the most common requirements in creative editing workflows. The evaluation was conducted using private, expert-curated datasets and model responses.
5 datapoints
Latest
Video-to-video SFT Dataset
Designed high-quality SFT datasets to improve Video-to-Video (V2V) model performance across controlled transformation, stylization, and structural video editing tasks. Each sample pairs a prompt with a source video and a target video.
5 datapoints
Latest

Frequently Asked Questions

  • Sample datasets are typically delivered within 48 hours, and even sooner for off-the-shelf (OTS) data.

  • All sample data is shared through our data platform and can be downloaded, making it easy for your team to review, analyze, and assess fit for your workflow.

  • Absolutely. We can set up an expert consultation to understand your requirements and create a custom data pack tailored to your needs

  • Yes. We support both standard and highly specialized requests across domains, and can build data pipelines suited to your scope, format, and quality requirements.

  • Explore curated data collections designed to accelerate model development across training, alignment, and evaluation workflows.

Have a specific query? Write to us

Thank you for your interest.
We will reach out to you shortly.
Oops! Something went wrong while submitting the form.