Why Do Robots Train in Simulation Before Entering the Real World?

Imagine teaching a walking robot by letting it fall 100,000 times.

In the real world, every fall has a cost.

The robot has to be reset. Batteries drain. Motors heat up. Gearboxes wear. A cable can snag. A person may need to walk over, stand the robot back up and start the next trial.

In simulation, the same failure can end with one instruction:

Reset.

Simulation is valuable because it turns robot experience into something that can be generated, reset, varied and repeated without consuming the physical robot at the same rate.

Start With the Real Problem: Experience Is Expensive

Robot learning needs experience.

For reinforcement learning, that may mean many sequences of:

observe → act → see what happened → update

On hardware, each sequence consumes real time.

It also creates physical consequences.

A bad exploratory action can:

  • hit a joint limit,
  • drop an object,
  • damage a gearbox,
  • collide with the environment,
  • or create a safety risk for people nearby.

This is the first reason simulation matters.

It makes failure cheap enough to repeat.

Original Asset 1: The Experience Factory

A useful way to think about simulation is not “a virtual robot.”

Think of it as an experience factory.

many environments
×
fast reset
×
controlled variation
×
repeatable measurement
=
high-throughput experience

That combination changes the economics of learning.

A physical lab may have one robot.

A simulator can run many copies of that robot with different starting poses, commands, floor friction or disturbances.

The learning algorithm gathers experience from all of them and improves one policy.

Reset Is More Important Than It Looks

A training episode has to end somewhere.

The robot may fall, reach the goal, time out or enter an invalid state.

Then the system must return to a useful starting condition.

On hardware, reset can be one of the slowest parts of the loop.

In simulation, reset can be nearly immediate and automatic.

Fast reset is not a convenience feature. It is one of the mechanisms that makes large-scale robot learning possible.

Parallel Worlds Change Wall-Clock Time

Now imagine 64 virtual robots.

Each one experiences a different version of the task.

Robot 1 ─┐
Robot 2 ─┤
Robot 3 ─┤
... ├→ experience → one policy
Robot 64 ─┘

NVIDIA Isaac Lab is designed to collect experience from many simulation environments in parallel and train robot policies at scale.[1]

This is one of the reasons GPU-accelerated simulation has become important in modern robot learning.

Original Asset 2: Three Speeds That Are Easy to Confuse

When someone says “our simulator is fast,” ask what they mean.

1. Physics step rate

How quickly can the simulator advance one virtual world?

2. Parallel environment count

How many worlds can advance at the same time?

3. Learning wall-clock time

How many real minutes or hours pass before the policy reaches the required performance?

These are related, but they are not the same.

Isaac Lab's current documentation explicitly notes that increasing the number of environments can improve sample collection only until simulation cost, optimizer structure or GPU memory becomes the bottleneck.[2]

More virtual robots can create more experience per second. They do not guarantee that learning becomes proportionally faster.

Simulation Is More Than a 3D Animation

A picture of a robot moving on a screen is not enough.

A training simulator must predict how the state changes after an action.

For a physical robot, that can require:

  • mass and inertia,
  • joint motion and limits,
  • gravity,
  • contact,
  • friction,
  • actuator behavior,
  • sensor behavior,
  • and timing.

MuJoCo—short for Multi-Joint dynamics with Contact—is an open-source physics engine designed for fast and accurate simulation of articulated systems and contact-rich behavior.[6]

Its value in robot learning is not cinematic graphics.

Its value is useful physical cause and effect.

Original Asset 3: The Simulation Layer Stack

A robot simulator can be thought of as a stack.

geometry / kinematics
↓
mass & inertia
↓
joints / limits
↓
actuator model
↓
contact / friction
↓
sensor model
↓
latency / timing
↓
task environment
↓
reset & randomization

A policy can become dependent on errors in any layer that matters to the task.

Does the Simulator Need to Be Perfect?

No.

In fact, “make everything as realistic as possible” is not a useful engineering requirement.

The better question is:

Which parts of reality can change the observation, action or success of this task?

Those parts deserve the most fidelity.

Original Asset 4: The Fidelity Budget

Different robot skills need different kinds of realism.

Walking

Important effects may include:

  • contact,
  • friction,
  • mass and inertia,
  • actuator strength,
  • backlash,
  • delay,
  • and body geometry.

Vision-based manipulation

Important effects may include:

  • camera calibration,
  • object appearance,
  • lighting,
  • contact and friction,
  • gripper behavior,
  • and object mass.

Navigation

The important fidelity may move toward:

  • geometry,
  • sensor coverage,
  • moving obstacles,
  • localization noise,
  • and communication delay.

A simulator can therefore be visually simple and still be excellent for one learning problem.

It can also be beautiful and still be wrong in the one physical detail that matters.

Photorealism and Physical Fidelity Are Different

If a locomotion policy uses only joint positions and an IMU, perfect shadows may not help it learn to walk.

If a vision policy has to recognize a cup under changing light, visual variation may be critical.

This is why simulator choice should follow the policy's observation and action contract rather than a general ranking of graphics quality.

Actuators Deserve More Attention Than Beginners Usually Give Them

A simulator can model a joint as if it instantly obeys a command.

A real motor does not.

It has:

  • torque limits,
  • speed limits,
  • control gains,
  • friction,
  • backlash,
  • electrical dynamics,
  • and thermal behavior.

MuJoCo's 3.7.0 release in April 2026 added a DC-motor actuator model with optional electrical dynamics, cogging torque, temperature-dependent resistance and LuGre friction.[8]

That is a useful reminder: simulation fidelity is not only about rigid-body geometry.

Why Model Sensor Noise and Delay?

The policy sees observations, not perfect physical truth.

A real sensor has noise, bias, sampling rate and delay.

If the simulator gives the policy perfect information that the real robot will never have, the policy can learn to depend on an impossible advantage.

The same applies to timing.

If actions arrive instantly in simulation but are delayed on hardware, the closed-loop behavior changes.

What Is Domain Randomization?

A real robot is never described by one perfect set of numbers.

Friction changes. Payload changes. Motor behavior changes. Calibration is imperfect.

Domain randomization means deliberately varying selected simulator parameters during training so the policy experiences a family of plausible worlds instead of one exact virtual world.

Original Asset 5: The Uncertainty Envelope

best estimate
±
plausible uncertainty
↓
family of training worlds

Current Isaac Lab tooling can randomize physical properties such as friction, mass, center of mass, actuator gains, joint parameters, gravity, external forces and collider properties, as well as visual properties for perception tasks.[4]

But there is an important discipline here.

Domain randomization does not mean “randomize everything.” It means train across uncertainty that is plausible and relevant.

Why Too Much Randomization Can Be a Problem

If the training worlds become physically implausible, the policy may spend capacity solving problems the real robot will never see.

Randomization ranges therefore need engineering judgment.

This is where another idea enters: system identification.

System Identification and Domain Randomization Are Different Jobs

System identification uses real measurements to estimate parameters of the physical system—for example motor response, friction or inertia.

Its job is to make the nominal simulator more representative of the real robot.

Domain randomization then exposes the policy to uncertainty that remains.

Original Asset 6: Nominal Model + Uncertainty

real measurements
↓
system identification
↓
better nominal simulator
±
domain randomization
↓
robust training distribution

The two techniques complement each other.

A Simulator Can Be Wrong Because of Code, Not Physics

This is an important expert-level check.

A simulator is software.

Configuration mistakes can quietly become training data.

Pollen Robotics' current Microduck RL repository documents several examples from its own sim-to-real engineering playbook:

  • observation normalization can appear correct inside simulation but be missing after export,
  • adding action filtering only on one side of training/deployment can break transfer,
  • a randomizer that accumulated changes across resets once degraded long runs,
  • and a reward can become inconsistent if it measures a different signal from the one the policy observes.[9]

A simulator bug is not merely a software bug. During learning, it can become part of the world the policy believes is real.

Reproducibility Is Harder Than “Set the Seed”

Experiments should be repeatable enough that engineers can tell whether a change helped.

Modern GPU simulation complicates this.

Isaac Lab supports deterministic seeds, but its documentation notes that GPU work scheduling and floating-point execution order can produce small numerical differences that may diverge over thousands of environments and simulation steps.[5]

For a practitioner, the lesson is simple:

reproducibility is an engineering property of the full simulation-and-training stack, not just one random seed.

Training Simulator vs. Digital Twin

These terms are often mixed together.

A digital twin usually aims to represent a particular real asset or system closely enough for monitoring, prediction or engineering analysis.

A training simulator may intentionally create many versions of the robot and environment.

Original Asset 7: Different Goals

Training simulatorDigital twin
Main goalGenerate useful experienceRepresent a specific real system
VariationOften intentionalUsually tries to track actual state/parameters
Success testDoes the policy learn robust behavior?Does the model predict/represent the real asset well?

The two can overlap, but they are not automatically the same thing.

A Concrete Example: Microduck

Pollen Robotics' current microduck_rl stack trains the small biped in MuJoCo Warp with PPO at 50 Hz and includes actuator physics, domain randomization, backlash simulation and ONNX export.[9]

The learning path is:

virtual robot
↓
physics + actuator model
↓
parallel experience
↓
policy training
↓
ONNX policy
↓
real runtime

An Even More Useful Engineering Idea: Keep the Software the Same

Microduck's current simulation design can run the real robot daemons against a MuJoCo body.

Above that seam, the same control loop, policy, safety, fall detection, kinematics and IPC interfaces can run as they do on the physical robot.[10][11]

This reduces one source of mismatch:

the simulation path and hardware path do not need completely different application software.

Original Asset 8: Simulator Fitness Test

Before trusting a training simulator, ask:

  1. Skill: What behavior is being trained?
  2. Physics: Which physical effects can change success?
  3. Model: Are those effects represented well enough?
  4. Actuator: Does the simulated actuator behave like the real one where it matters?
  5. Observation: Does the policy receive information available on hardware?
  6. Timing: Are relevant sensor and action delays represented?
  7. Reset: Can failures be reset cleanly and automatically?
  8. Variation: Can important uncertainty be randomized deliberately?
  9. Throughput: Can enough experience be generated fast enough?
  10. Reproducibility: Can configurations, seeds and results be traced?
  11. Validation: Has the simulator been compared with real measurements?
  12. Boundary: What still must be tested on hardware?

What Simulation Cannot Give You

Simulation can make experience cheaper.

It cannot prove that the real world matches the model.

It cannot reveal every hardware defect, calibration drift, cable interaction, lighting condition, surface property or rare human behavior unless those effects are represented or tested elsewhere.

And it cannot eliminate the need for real-hardware validation.

The Simple Idea to Remember

A good simulator is not the one that looks most like reality.

It is the one that reproduces the cause-and-effect relationships the policy needs, fast enough and broadly enough to learn useful behavior.

Simulation is an experience factory, not a replacement for reality.

The next question is where the model and reality disagree—and what to do about it.

Next: What Is Sim-to-Real? Why a Robot That Works in Simulation Can Fail in Reality

Key Vocabulary

simulation
A computational model that predicts how a robot and environment change after actions.

parallel environment
One of many independent simulated worlds collecting experience at the same time.

simulation throughput
How much simulated experience a system can generate per unit of real time.

task-relevant fidelity
Accuracy in the parts of the model that materially affect the policy and task.

domain randomization
Deliberately varying selected simulation properties during training.

system identification
Estimating model parameters from measurements of the real system.

digital twin
A model intended to closely represent a particular real asset or system.

determinism
The property that the same initial conditions and computation produce the same result.

Read the Physical AI Learning Series

  1. How Does a Robot Learn?
  2. What Is a Robot Policy?
  3. Why Train Robots in Simulation? (this article)
  4. What Is Sim-to-Real?
  5. Why Robot Actuators Matter
  6. Why Microduck Matters

Related Articles

Sources

  1. NVIDIA — Isaac Lab, checked October 4, 2026.
  2. Isaac Lab — Reinforcement Learning and scaling, checked October 4, 2026.
  3. Isaac Lab — Multi-GPU and Multi-Node Training, checked October 4, 2026.
  4. Isaac Lab — Environment events and domain randomization, checked October 4, 2026.
  5. Isaac Lab — Reproducibility and Determinism, checked October 4, 2026.
  6. MuJoCo — Advanced Physics Simulation, checked October 4, 2026.
  7. MuJoCo Documentation — Overview, checked October 4, 2026.
  8. MuJoCo — Changelog, including version 3.7.0, April 14, 2026.
  9. Pollen Robotics — microduck_rl, checked October 4, 2026.
  10. Pollen Robotics — Microduck simulation design, checked October 4, 2026.
  11. Pollen Robotics — The simulated duck, checked October 4, 2026.
  12. Muratore et al. — Robot Learning From Randomized Simulations: A Review.

Sources checked through October 4, 2026. Simulation quality is task-dependent. A model can be highly accurate in one property and inadequate in another, so real-hardware validation remains necessary.