A robot walks beautifully in simulation.
The same policy is moved to real hardware.
Now the robot shakes, drifts, misses its steps—or falls.
The obvious explanation is: “the simulator was not realistic enough.”
That is true, but not very useful.
The engineering question is more precise:
Which part of the closed-loop system is different enough to change the robot's behavior?
That question is the heart of sim-to-real.
Quick Answer
Sim-to-real is the process of making a policy or behavior developed in simulation continue to work on physical hardware.
The difficulty is the reality gap: the mismatch between the virtual system used during development and the physical system used during deployment.
But the reality gap is not one error.
It is a collection of different gaps.
Original Asset 1: The Reality Gap Map
| Gap | Simulation might assume | Reality might do |
|---|---|---|
| Dynamics | Known mass, friction and contact | Different mass, surface and impacts |
| Actuator | Ideal joint response | Voltage sag, friction, backlash, saturation |
| Observation | Clean, normalized state | Noise, bias, wrong scale or frame |
| Timing | Immediate sensing and actuation | Sampling, communication and command delay |
| Software | Training preprocessing | Different export, filtering or action scaling |
| Environment | Known floor and objects | Variation, wear, people and disturbances |
This decomposition is useful because different gaps need different fixes.
A Policy Can Be Identical While the Robot System Is Not
Suppose you export exactly the same neural-network weights from simulation to hardware.
It is tempting to say:
“The policy is the same, so the behavior should be the same.”
But a policy is only one part of a feedback loop.
Original Asset 2: Transfer Without the Math
real behavior
depends on
policy
+ real observations
+ real actuator response
+ real timing
+ real environment
+ software contract
If any of those changes enough, the closed-loop system changes.
The same policy can therefore produce different motion.
Gap 1: Dynamics
The simulator needs values for mass, inertia, geometry, friction and contact.
The real robot never matches those values perfectly.
A foot may grip the floor differently.
A backpack or battery can shift mass.
A compliant surface can change impact behavior.
For a slow task, some of these differences may barely matter.
For dynamic walking or jumping, a small difference can arrive at exactly the wrong moment in the gait cycle.
Gap 2: The Actuator
This is one of the most important gaps in physical robots.
The policy does not move the joint by itself.
Its command goes through motors, drivers, gears, control firmware and mechanical transmission.
A real actuator can have:
- torque and speed limits,
- backlash,
- Coulomb and load-dependent friction,
- battery-voltage dependence,
- current limits,
- thermal effects,
- and nonlinear control behavior.
The current Microduck RL stack models its small Dynamixel servos with voltage control, load-dependent friction, randomized battery voltage, load-dependent voltage sag, command delay and explicit backlash variants.[5][7]
Pollen Robotics makes the reason explicit in its repository: at the scale of a small ~800 g biped, actuator fidelity accounts for a large part of the sim-to-real gap.[5]
Gap 3: Observation
A policy does not react to the physical world directly.
It reacts to its observations.
That means a transfer can fail even if the physics is good.
Examples:
- sensor bias,
- different coordinate frames,
- different units,
- missing normalization,
- different filtering,
- or an encoder measuring a slightly different mechanical point.
Microduck's current engineering notes document a particularly important example: observation normalization can look correct during in-simulation testing because the simulator applies it automatically, but a manually exported policy can fail on hardware if that normalizer is not included in the deployed model.[6]
Sometimes the “reality gap” is not physics. It is a data-contract mismatch.
Gap 4: Timing
Physical systems run in time.
A command that is correct now may be wrong 80 milliseconds later.
Delay can come from:
- sensor sampling,
- communication,
- policy inference,
- motor buses,
- and actuator response.
Isaac Lab's current actuator documentation provides a delayed actuator model that can sample command delay at reset. Its example shows delays from 0 to 133 milliseconds making the joint increasingly trail its reference.[4]
This is why latency is part of the dynamics seen by the policy.
Gap 5: Software
Training and deployment are often separate software pipelines.
That creates opportunities for silent mismatch.
Microduck's current sim-to-real playbook warns against:
- different observation normalization after export,
- action filtering used only during training or only during deployment,
- randomization code that accidentally accumulates across resets,
- and rewards that measure a different signal from the one the policy observes.[6]
These are not glamorous failures.
They are also exactly the kind that can waste weeks because the policy looks perfect in the viewer.
Gap 6: The Environment
The real world contains variation that may never have appeared in training.
Surfaces change.
Objects move.
Lighting changes.
People interfere.
Hardware wears.
So sim-to-real is partly a modeling problem and partly a robustness problem.
The Wrong Goal: One Perfect Virtual Copy
Improving simulator accuracy is useful.
But one perfectly calibrated copy of one robot on one day is not the whole target.
The physical system itself varies.
This is why sim-to-real usually combines two ideas:
- make the nominal model better,
- train across the important uncertainty that remains.
System Identification Comes First
System identification means using measurements from the real robot to estimate parameters of the model.
For example:
- motor response,
- joint friction,
- delay,
- mass and inertia,
- or contact parameters.
The result is a better nominal simulator.
Then Domain Randomization Handles Remaining Uncertainty
Domain randomization deliberately varies selected simulation parameters during training.
Instead of learning one exact world, the policy experiences a family of plausible worlds.
The classic intuition from domain-randomization research is that if the training distribution contains enough relevant variation, reality can become one case the policy is already prepared for.[1]
But that idea is easy to oversimplify.
Original Asset 3: Targeted Randomization
Randomize a parameter when three things are true:
- It is uncertain.
- It can change task success.
- The range is physically plausible.
Examples might include:
- floor friction,
- battery voltage,
- motor strength,
- command delay,
- joint friction,
- payload mass,
- sensor bias.
Domain randomization is not “add more randomness.” It is “train across the uncertainty that matters.”
Can Too Much Randomization Hurt?
Yes.
If the range is too wide, the policy may become conservative or spend capacity solving unrealistic versions of the task.
A 2026 review of sim-to-real RL emphasizes the combined role of simulator optimization, actuator modeling and domain randomization rather than treating randomization as a standalone cure.[3]
This leads to a better workflow.
Original Asset 4: The Transfer Loop
measure real robot
↓
system identification
↓
better nominal model
↓
targeted randomization
↓
train policy
↓
deploy
↓
measure mismatch
↺
Real-world failure is not only a failure.
It is new data about which assumption was wrong.
Zero-Shot Transfer Does Not Mean Zero Testing
Zero-shot sim-to-real usually means the task policy is deployed on real hardware without additional task-level learning or fine-tuning.
It does not mean:
- no calibration,
- no safety checks,
- no hardware validation,
- or no model engineering.
Some projects do achieve direct simulation-to-hardware transfer. The important question is what engineering work made that possible.
Original Asset 5: The Transfer Ladder
| Level | What happens |
|---|---|
| 1. Zero-shot | Deploy the simulated policy directly after calibration and safety checks |
| 2. Recalibrate & retrain | Update simulator parameters/ranges, then train again |
| 3. Local adaptation | Adapt selected components such as actuator correction or residual policy |
| 4. Real-world fine-tuning | Continue policy learning on hardware under constraints |
Higher numbers are not automatically better.
The right level depends on safety, data cost, hardware wear and how large the mismatch is.
What Should You Measure on the Real Robot?
The first hardware deployment should be treated as an experiment.
Original Asset 6: Real-Hardware Measurement Pack
Record as much of the following as practical:
- commanded joint position or velocity,
- measured joint position or velocity,
- policy action,
- current or torque proxy,
- battery voltage,
- loop rate and timestamp,
- IMU data,
- contact timing,
- actuator temperature when relevant,
- safety intervention,
- and task-level performance.
Then compare the same quantities in simulation.
The point is not merely to ask “did it fall?”
The point is to ask where the trajectories began to diverge.
Original Asset 7: Sim-to-Real Debug Order
When transfer fails, check in this order:
- Policy contract: same inputs, outputs, units and action scaling?
- Observation: same normalization, frame, bias and filtering?
- Timing: same loop rate and realistic delay?
- Actuator: same response, limits, friction, backlash and voltage behavior?
- Contact/environment: plausible friction, surfaces and geometry?
- Training distribution: did training cover this condition?
- Learning algorithm: only after simpler mismatches have been checked.
This order is not universal, but it prevents a common mistake:
Do not blame the learning algorithm for a unit mismatch, missing normalizer or 80-millisecond delay.
There Is Another Direction: Make Reality More Like the Model
Most sim-to-real work asks:
How do we make simulation better match reality?
A 2026 paper on actuator reality shaping explores the reverse idea: use low-level control to make real actuator behavior match a standardized reference model used during policy training.[8]
The broader idea is useful even if the specific method changes:
simulation → more realistic
and/or
hardware → more standardized
Both can make the interface seen by the policy more consistent.
Why This Leads Directly to Actuators
Many reality-gap problems become visible where software turns into force.
That is the actuator system.
The next article therefore goes one layer deeper.
Next: Why does robot AI need good actuators, not just a good model?
The Simple Idea to Remember
Sim-to-real is not a one-time jump from a virtual robot to a physical robot.
It is an engineering loop.
Measure the mismatch. Decide which gap caused it. Improve the model, training distribution, runtime or hardware interface. Then test again.
Key Vocabulary
sim-to-real
Transferring a behavior or policy developed in simulation to physical hardware.
reality gap
The mismatch between the simulated closed-loop system and the physical one.
system identification
Estimating model parameters from measurements of the real system.
domain randomization
Training across deliberately varied simulation parameters.
zero-shot transfer
Deploying a trained policy without additional task-level learning on the real robot.
actuator model
A model of how motor, gearing, electronics and low-level control turn commands into physical motion or force.
Read the Physical AI Learning Series
- How Does a Robot Learn?
- What Is a Robot Policy?
- Why Train Robots in Simulation?
- What Is Sim-to-Real? (this article)
- Why Robot Actuators Matter
- Why Microduck Matters
Related Articles
- What Is Physical AI? When AI Leaves the Screen and Enters the Real World
- What Is the AI Full Stack?
- AI Training vs. Inference
- Why the Robot Race Is Becoming a Manufacturing Race
Sources
- Tobin et al. — Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World.
- Muratore et al. — Robot Learning From Randomized Simulations: A Review.
- Tiwari, Khapre & Singh — Reinforcement learning in robotic systems: A review on sim-to-real transfer, April 2026.
- Isaac Lab — Actuators and command delay, checked October 4, 2026.
- Pollen Robotics — microduck_rl README, checked October 4, 2026.
- Pollen Robotics — microduck_rl engineering invariants, checked October 4, 2026.
- Pollen Robotics — Microduck sim-to-real inference controls, checked October 4, 2026.
- Yamamori et al. — Actuator Reality Shaping for Zero-Shot Sim-to-Real Robot Learning, July 2, 2026.
Sources checked through October 4, 2026. Sim-to-real performance is task- and hardware-dependent. Domain randomization, actuator modeling and system identification improve transfer only when they address mismatches that materially affect the closed-loop behavior.