What Is a Robot Policy? How AI Turns Sensors Into Movement

A robot can have cameras, joint sensors, an IMU and a powerful AI computer.

None of those parts, by themselves, answer the most immediate question:

What should the robot do next?

A robot policy is one way to answer that question.

A robot policy is the rule that turns what the robot can observe—and what it is trying to do—into its next action.

The One Loop to Remember

If you remember only one diagram from this article, remember this:

Sensors
↓
Observation
↓
Policy
↓
Action
↓
Controller
↓
Actuator
↓
Movement
↓
New sensor data

The robot repeats this loop again and again.

That repetition is what turns a static AI model into part of a moving physical system.

First, Sensors Are Not the Same as Observations

A camera produces pixels.

An IMU measures motion and orientation.

Joint encoders measure where the robot's joints are.

These are sensor measurements.

The observation is the information package actually given to the policy.

It may contain raw measurements, processed features, previous actions, commands, or estimated quantities.

So the path is not always:

sensor → policy

It is often:

sensor → processing / normalization → observation → policy

What Does the Policy Produce?

The output of the policy is called an action.

But “action” does not always mean “send this voltage to the motor.”

A locomotion policy might output:

  • target joint positions,
  • target joint velocities,
  • torque commands,
  • small changes to a target,
  • or several future actions at once.

The hardware below the policy may still need to turn those targets into safe motor commands.

Original Asset 1: Planner → Policy → Controller

Suppose the user says:

“Walk to the charging station.”

The robot may separate that request into layers.

LayerQuestion it answersExample
PlannerWhat should happen, and in what order?Turn left, cross the room, approach charger
PolicyGiven the current situation, what action should happen next?Move these joints toward the next walking target
ControllerHow do we track that target stably?Correct joint error using feedback
ActuatorHow is physical force produced?Motor creates torque and motion

The exact boundaries vary by robot.

Some modern models combine planning and policy more tightly.

But this layered mental model prevents a common mistake:

“The AI decided to move” is not one step. It is a stack of decisions and control loops.

Original Asset 2: The Policy Contract

A useful way to understand any robot policy is to ignore the marketing name and ask five questions.

Policy
=
Input + Output + Rate + Memory + Safety Boundary

1. Input — What can it observe?

Joint positions? Camera images? Language command? Previous action? Contact state?

2. Output — What does it command?

Joint targets? Velocity? Torque? End-effector motion? A chunk of future actions?

3. Rate — How often does it run?

Ten times per second? Fifty? Hundreds?

4. Memory — Does it remember anything?

Only the current observation? A short history? A recurrent hidden state?

5. Safety Boundary — Who can say no?

Can another layer clamp a dangerous joint target, stop after a fall or ignore stale sensor data?

If you know these five things, you understand far more about a robot policy than you do from knowing only the neural-network brand name.

Why the Rate Matters

A physical robot keeps moving while the computer is thinking.

If the policy is supposed to run at 50 Hz, one cycle is about 20 milliseconds.

The system must read sensors, prepare the observation, run inference, process the action and communicate with the actuators within a useful timing budget.

If it repeatedly misses that timing, the robot may respond late to the world.

This is why policy latency is not just a software-performance detail.

It is part of control behavior.

A Concrete Example: Microduck at 50 Hz

Pollen Robotics' current Microduck runtime makes the policy contract unusually visible.

The robot runs a 50-Hz control loop that reads the robot state, chooses joint targets and writes commands back to its servos.[2][4]

Its current recurrent-policy interface uses:

  • a 61-dimensional observation,
  • 14 raw joint actions,
  • either a feed-forward model or an LSTM recurrent model,
  • and the same 50-Hz runtime contract.[3]

That gives us a concrete mental model:

61 numbers describing the current situation
↓
policy inference
↓
14 joint-action values
↓
runtime safety / control
↓
motors move

Why 50 Hz Is Not a Universal Magic Number

A faster policy loop is not automatically better.

Higher rates can improve responsiveness, but they also increase compute demand, communication load and sensitivity to timing jitter.

A slow high-level planner may run at a very different rate from a fast balance controller.

The correct rate depends on:

  • robot dynamics,
  • task speed,
  • sensor rate,
  • actuator bandwidth,
  • inference latency,
  • and what the lower-level controller is already doing.

Feed-Forward Policy: Decide From the Current Snapshot

A feed-forward policy produces an action from its current input without carrying an internal memory from one inference step to the next.

It can still receive history explicitly—for example, the previous action can be included in the observation.

Recurrent Policy: Carry Some Memory Forward

A recurrent policy also carries hidden internal state from one step to the next.

This can help when the current sensor snapshot does not tell the whole story.

Examples include:

  • which part of a gait cycle the robot is in,
  • whether contact happened a moment ago,
  • motion hidden by noisy sensors,
  • or an object that has temporarily disappeared from view.

Original Asset 3: The Memory Test

Ask:

If I froze the robot at this exact instant, would the current observation alone tell me everything needed to choose the next action?

If yes, a memoryless policy may be enough.

If no, history or recurrent memory may help.

But memory adds a new engineering problem: the hidden state must be reset at the right time.

Microduck's current recurrent runtime specifies when policy memory is cleared—for example after controller resets, some pauses, policy changes or fall recovery.[3]

Policy Memory Is Not the Same as Long-Term Robot Memory

A recurrent hidden state is usually short-term computational context.

It is not automatically a database of past experiences or a permanent memory of what happened yesterday.

This distinction becomes important as robots add larger planning and memory systems.

What About Vision-Language-Action Models?

VLA stands for Vision-Language-Action.

The name is a useful shortcut for what the model is trying to do: see the world, understand a language instruction, and turn both into robot actions.

For example, imagine telling a robot, “Pick up the red cup.”

The camera provides the visual scene. The language instruction provides the goal. The VLA uses those inputs—together with the robot's own state—to produce actions that move the robot toward the cup, grasp it and continue the task.

A vision-language-action model is therefore still understandable through the same policy idea, but with richer inputs and often a broader range of tasks.

The inputs become richer.

Instead of only joint state, the model may receive:

  • camera images,
  • natural-language instructions,
  • and the robot's proprioceptive state—its internal sense of joint configuration.

The output is still action.

Google DeepMind's current Gemini Robotics On-Device 2 model card makes this especially clear: its inputs include text, images and robot proprioception, and its outputs are numerical robot actions.[9]

A VLA expands what a policy can understand. It does not remove the need for controllers, actuators, timing and safety.

What Is an Action Chunk?

Small policies often produce one action at a time.

Large modern robot policies may predict several future actions together.

This is called an action chunk.

NVIDIA's current GR00T documentation describes its models as taking image, language and robot-state inputs and producing action chunks—predictive sequences of relative joint motions.[6]

A simple picture is:

current observation
↓
large VLA policy
↓
[action t, action t+1, action t+2, ...]

Predicting ahead can help smooth execution and reduce the need to wait for a large model before every tiny movement.

But it creates another question:

What if the world changes before the whole action chunk is finished?

Modern systems therefore need ways to refresh, interrupt or overlap policy predictions with execution.

Why Big Policies Make Latency More Important

A small locomotion policy can be fast.

A large multimodal VLA may require far more computation.

Kevin Black's 2026 UC Berkeley dissertation identifies real-time control with large, compute-intensive robot policies as a central challenge and studies asynchronous execution so robot motion does not simply wait for every large policy inference to finish.[10]

This reveals a useful principle:

A policy can be intelligent enough to choose a good action and still be unusable if it chooses that action too late.

Where Safety Fits

A learned policy should not automatically be the only authority allowed to move hardware.

Microduck's runtime includes safety logic such as joint limits, fall handling and deadman-style intent logic around the policy/control path.[4]

Google DeepMind similarly describes robotics safety as multiple semantic, physical and operational layers rather than one model acting as the only safety barrier.[11]

The general architecture is:

policy proposes
↓
control & safety check
↓
hardware executes

How Is a Policy Created?

A policy does not belong to one training method.

It can be:

  • hand-designed,
  • derived from classical control,
  • learned from demonstrations,
  • trained with reinforcement learning,
  • fine-tuned from a robot foundation model,
  • or built by combining several of these.

In the previous article, reinforcement learning was one way to train a policy.

The policy is the thing that deployment repeatedly runs.

Original Asset 4: Policy Reality Check

When a company says “our robot uses an AI policy,” ask:

  1. Input: What exactly can the policy observe?
  2. Output: What exactly does it command?
  3. Rate: How often does it run?
  4. Latency: How long does one inference take?
  5. Memory: Feed-forward or recurrent?
  6. Horizon: One action or an action chunk?
  7. Controller: What sits underneath it?
  8. Safety: What can reject or clamp its output?
  9. Failure: What happens after sensor dropout or missed deadlines?
  10. Evaluation: Does it work outside the demo conditions?

This is a better test than asking only how many parameters the model has.

Why This Matters for Physical AI

Physical AI is often described with impressive words: reasoning, autonomy, generalization and intelligence.

But eventually the system has to cross a very practical boundary:

what the robot knows
↓
what the robot does next

The policy lives at that boundary.

And because the action changes a physical system, the policy's timing, memory, output format and safety interface matter as much as its abstract intelligence.

The Simple Idea to Remember

If you forget every technical term, remember this:

A robot policy repeatedly turns “what I can sense now + what I am trying to do” into “what I should do next.”

And if you want to understand a real policy, ask five things:

What goes in?
What comes out?
How often?
What does it remember?
Who can stop it?

Key Vocabulary

policy
A decision rule that maps observations and goals to actions.

observation
The information package actually provided to the policy.

action
The command produced by the policy, such as a joint target, velocity, torque or action chunk.

controller
A lower-level feedback system that tracks targets and manages physical dynamics and limits.

planner
A system that chooses tasks, routes, subtasks or longer-horizon sequences.

proprioception
The robot's internal sense of its own body state, such as joint position or velocity.

recurrent policy
A policy that carries hidden state from one inference step to the next.

action chunk
A sequence of future actions predicted together by a policy.

Read the Physical AI Learning Series

  1. How Does a Robot Learn?
  2. What Is a Robot Policy? (this article)
  3. Why Train Robots in Simulation?
  4. What Is Sim-to-Real?
  5. Why Robot Actuators Matter
  6. Why Microduck Matters

Related Articles

Sources

  1. OpenAI Spinning Up — Key Concepts in RL, checked October 4, 2026.
  2. Pollen Robotics — Microduck runtime, checked October 4, 2026.
  3. Pollen Robotics — Recurrent ONNX policies, checked October 4, 2026.
  4. Pollen Robotics — robotd control-loop design, checked October 4, 2026.
  5. Pollen Robotics — Policy manifest, checked October 4, 2026.
  6. NVIDIA — Isaac GR00T, checked October 4, 2026.
  7. NVIDIA — Develop Humanoid Robot Policies End-to-End with GR00T 1.7, July 7, 2026.
  8. Google DeepMind — Gemini Robotics 2, July 30, 2026.
  9. Google DeepMind — Gemini Robotics On-Device 2 model card, July 2026.
  10. UC Berkeley — General Robot Manipulation with Multi-Modal Vision-Language-Action Models, August 14, 2026.
  11. Google DeepMind — Robotics safety framework, checked October 4, 2026.

Sources checked through October 4, 2026. “Policy” is a broad robotics term and system boundaries vary. A learned policy can output different action representations and may sit above or alongside classical controllers, planners and independent safety systems.