Simulation is best used to expand controlled variation and reduce the cost of rare scenarios; real-world collection is needed to anchor models to the sensor noise, contact physics, human behavior and operational constraints that simulation does not reproduce reliably. Generalization comes from designing the two sources as complementary coverage layers.
“Sim or real?” is the wrong strategic question. The useful question is: which parts of the deployment distribution can we generate safely and systematically, and which parts must be observed in the physical world?
What simulation is unusually good at
Simulation can vary parameters that are slow, expensive or unsafe to vary physically:
- object pose, scale, texture and mass;
- lighting, camera placement and background;
- robot geometry and controller parameters;
- rare initial states and perturbations;
- large numbers of repetitions with exact state labels;
- counterfactual evaluation under controlled changes.
Systems such as MimicGen demonstrate a compelling pattern: use a small number of human demonstrations as structure, then generate much larger datasets across new scene configurations. The project reports more than 50,000 generated demonstrations from fewer than 200 human demonstrations across 18 tasks.
The important idea is not that synthetic data replaces people. It is that human demonstrations can define task structure while simulation multiplies controlled coverage.
What the physical world still owns
Real-world data captures effects that are difficult to enumerate:
- contact with deformable, reflective or transparent objects;
- wear, backlash, vibration and calibration drift;
- motion blur, exposure changes and imperfect depth;
- messy human environments and unplanned object combinations;
- operator adaptation, hesitation and alternative strategies;
- reset cost, safety stops and workflow constraints.
DROID is valuable not only because of its 76,000 trajectories, but because those demonstrations span 564 real scenes and were gathered by 50 collectors across three continents. Scene and collector diversity are part of the signal.
Define the long tail as a coverage cube
“Long tail” often becomes a vague synonym for “more variety.” Make it operational with four axes:
- Object: geometry, material, articulation, texture, state and condition.
- Environment: layout, lighting, clutter, surface, camera and distractors.
- Actor: operator strategy, speed, handedness, expertise and interaction style.
- Outcome: clean success, inefficient success, partial completion, recoverable deviation and terminal failure.
Each episode occupies a cell in this cube. The goal is not uniform coverage of every theoretical combination. It is risk-weighted coverage of combinations likely to appear in deployment or expose model weakness.
A six-step sim–real program
1. Build the task ontology in the real world
Observe people and robots performing the task before constructing a large simulator. Identify subtasks, contact events, object states, allowed strategies and genuine failure modes. Egocentric and exocentric human capture can reveal procedural variation that a lab-designed task script misses.
2. Create a small physical anchor set
Collect high-quality, synchronized episodes on the target or a representative embodiment. This set establishes camera statistics, action ranges, timing, contact signatures and realistic task duration.
3. Generate controlled variation in simulation
Expand factors the simulator represents credibly. Preserve the parameter values for every generated episode so performance can later be analyzed by condition rather than only by average success.
4. Measure the sim–real gap by feature
Do not speak of one global gap. Compare image appearance, depth noise, action dynamics, contact timing, object motion and episode duration separately. Different gaps require different remedies.
5. Collect real residuals
Deploy the policy in controlled physical trials and collect the conditions where simulation-trained behavior fails. These residuals are the highest-value real-world collection targets because they represent evidence, not speculation.
6. Rebalance and version
Keep real, synthetic, human and robot data identifiable. Set sampling weights intentionally and run ablations. If lineage is lost, the team cannot determine which source improved—or harmed—the model.
Diversity is not randomness
Randomizing every available parameter can create physically implausible scenes and waste training capacity. Diversity should follow hypotheses. If the model fails on reflective packaging, vary material and illumination while holding unrelated factors stable enough to interpret results. If it fails when the camera moves, create a structured camera-pose sweep.
Research on robotic manipulation scaling laws suggests that environment and object diversity can matter more than adding repetitions within the same condition. This supports a practical rule: once a cell is stable, spend the next episode on a new cell or a known failure mode.
Where human data adds unique value
Humans do more than provide trajectories. They expose task intent, shortcuts, recovery strategies and the range of “acceptable” outcomes. Ego-Exo4D combines synchronized views with narrations, action descriptions and expert commentary, illustrating how language and skill context can turn video into more than pixels.
For a commercial embodied AI program, human capture can support:
- task discovery before the robot is ready;
- semantic and procedural pretraining;
- rare environment coverage;
- expert-strategy comparison;
- identification of safe recovery options.
A buyer’s rule for budget allocation
Fund simulation where the state is controllable and the model of the world is credible. Fund real collection where uncertainty, contact, sensors or human environments dominate. Fund targeted collection whenever evaluation reveals a repeatable failure cluster.
MovraWorks’ egocentric, UMI and custom multimodal programs can create the physical anchor and residual datasets around a simulation pipeline. The point is not to win an argument between synthetic and real data. It is to make every source responsible for the part of generalization it can actually support.
