Data and Simulation Infra for Physical AI
Engine-level gameplay, real-world egocentric data, and scalable high-fidelity simulation for world models and robotics.

01 Mission
An applied research lab closing the deployment gap.



The next industrial revolution begins when robots leave the lab.
Physical AI produces impressive lab demos. Yet very few robots make it into production at scale.
Closing that gap requires two components:
Data that captures the diversity and messiness of the real world at the fidelity models need.
Simulation where robots can be evaluated, learn through interaction, and improve at scale.
Midcentury builds both.

02 Datasets
High-Fidelity Data for the Real World
Rights-cleared robotics and world model data, captured in motion with the full signal stack for the next generation of embodied agents.



Conversational Voice
Natural multilingual conversation captured in full-duplex, multi-channel settings. Speaker-separated audio, transcripts, and metadata built for speech and conversational models.
69K+
Hours
25+
Languages
Multi-Channel
Audio
03 Simulation
Midcentury Matrix

Develop, evaluate, and improve physical AI in massively parallel simulation.

01
Develop
Build digital twins, environments, and scenario distributions around real deployment conditions. Combine classical simulation with learned rendering and manipulation physics for higher sim-to-real fidelity.

02
Evaluate
Massively parallel GPU evaluation across thousands of closed-loop scenarios. Sweep conditions, compare builds, replay failures, and isolate regressions before returning to hardware.

03
Learn
Turn failures into new scenarios and training experience. Feed simulation and real-world data back into policies so systems continuously improve through interaction.
04 Research
Our contributions to the frontier of physical intelligence.

01 / 03
STEAM
Closed-loop evaluation for spatial intelligence.
A purpose-built benchmark of interactive 3D environments that measures navigation, object interaction, spatial reasoning, environmental understanding, and planning directly from simulator state.

02 / 03
Giacometti
Action supervision from human video.
A data engine that turns raw egocentric video into training-grade physical interaction data, including metric 3D hand trajectories, camera motion, depth, and task structure.
03 / 03
Neural Renderer
High-fidelity rendering from engine-level supervision.
A learned renderer trained on gameplay data paired with depth, normals, albedo, motion vectors, and camera state to map structured simulator state into realistic visual observations.

