← home

Teaching a pixel robot to stack crates

A 3,268-parameter policy, trained with PPO in a NumPy sandbox, running live in your browser at 60 Hz.

by aadi shah·August 2026·9 min read

The live arm needs a wider screen, so it's hidden on mobile. The full write-up is below.

Nothing about the arm above is keyframed. Every frame, a small neural network reads fourteen numbers from the simulator and outputs four commands: drive the base, turn each of the two joints, and open or close the gripper. It learned everything it does over 32 million simulator steps of trial and error, guided by the reward described below.

01Problem formulation

A task like this is usually framed as a Markov decision process. At each step the agent sees a state, a compact summary of what's going on, and picks an action. The environment moves forward and hands back a reward. That reward is just a score for that one step. It doesn't tell the agent what it should have done.

Reinforcement learning is about learning a policy from that feedback by trial and error. What counts is the total reward over the whole episode, so an action can be worth taking even if it pays off little right away.

policy14 → 48 → 48 → 4simulatorarm, base, cratesaction adrive · shoulder · elbow · gripstate s, reward r14 numbers · one scalar
One step of the loop: the simulator sends back a state and a reward, and the policy answers with an action.

The Markov part means the current state and action are enough to predict what happens next. The policy doesn't need to know how the arm got to its current pose. If that history mattered, it would have to go into the state.

What we're actually learning is the policy, a function that takes a state and returns a distribution over actions:

is the weights of a small neural network. During training the policy samples from that distribution so it keeps trying new things. In the browser it just takes the mean. Both states and actions are continuous:

state
14 numbers: base position along the track (1); sine and cosine of the shoulder and elbow angles (4); how open the gripper is (1); whether it is holding a crate (1); the offset from the gripper to whatever it currently wants (2); that target's offset from the base (1); and the elbow and gripper positions relative to the shoulder (4). Everything is an offset rather than a world coordinate, so the same situation looks the same anywhere along the track.
action
4 continuous rates: drive the base, bend the two joints, open or close the gripper.
reward
One number per step, built from progress toward the current target, gripper state, a time cost, motion penalties, and bonuses or penalties for grasping, placing, or dropping.
environment
A 480 × 60 px world with a wheeled base at , a two-link arm with 16 px and 14 px segments, a gripper, and two crate columns 390 px apart.
episode
Move one crate from the source stack to the destination stack. Capped at 600 steps, and ends early on a place or a drop.

02Learning with PPO

So the setup defines what the policy sees, what it can do, and how it's scored. What it never gets is the right answer for any state.

With nothing to imitate, the policy has to act, see what reward comes back, and adjust. That makes training online: the policy generates its own data in the simulator, and every update changes what the next rollout looks like.

Policy gradients:

We want to maximize the expected return . Here is one rollout and discounts rewards that arrive later. The catch is that we can't backprop through the simulator. The policy gradient theorem gets around this by rewriting the gradient in terms of the policy's log-probability, which we can differentiate with respect to :

Value and advantage:

Intuitively, the update makes actions that went better than expected more likely and actions that went worse less likely. To pin down “expected”: the value of a state is the return the policy expects from there on, and the action-value is the same thing if you commit to one action first. The gap between them is the advantage:

means the action beat the policy's own average from that state, and means it fell short. We don't know or , so a second network, the critic , learns , and is estimated from the critic plus the rewards we actually saw. Subtracting as a baseline doesn't bias the gradient. It just cuts the variance.

In practice you estimate the expectation by averaging over a batch of steps. That's REINFORCE [5] with a baseline. It's simple and unbiased, but noisy enough that big update steps aren't safe.

Advantage estimation:

We never observe the advantage directly, so every step needs an estimate built from the rewards that followed it and the critic's predictions. The simplest version looks one step ahead, and it's called the TD residual:

That's low-variance, but it's only as good as . At the other extreme, Monte Carlo sums the actual rewards to the end of the episode. That's unbiased, but much noisier. GAE [3] blends the two using :

recovers the one-step estimate, the Monte Carlo return minus the baseline, and anything in between is a geometric blend. I used , which keeps variance manageable over 600-step episodes without leaning too hard on a critic that starts out pretty bad.

Clipped objective:

With an advantage for every step, the remaining question is how big a step to take. There's a catch: means the rollouts came from the current policy, so the gradient estimate only holds for the parameters that collected them. That's what makes the method on-policy. After one update the batch describes an older policy, and reusing it without correction makes the estimate worse. So a big update can hurt twice: it can make the policy worse, and it can leave you with data that no longer describes it.

TRPO [2] handles this with a hard limit on the KL divergence between the old and new policy (a measure of how different two distributions are). It enforces that limit with second-order methods, which use curvature as well as slope and are expensive. PPO [1] replaces this with a clipped first-order objective. You can take several gradient steps on the same batch, but there's no extra reward for pushing any one action's probability too far in a single update.

To do that, PPO tracks how much each sampled action's probability has changed:

When the data is collected, . During the update it drifts: above 1 means the action got more likely, and below 1 means less likely. PPO clips the ratio once it moves past in the direction the advantage is pushing it:

How the clip plays out depends on the sign of the advantage.

Positive advantage (). Making the action more likely helps, but only up to . Past that, the clipped term is smaller, the picks it, and since it's constant in , that sample stops contributing gradient.

Negative advantage (). Now the objective wants the action to become less likely. Below the term goes flat, so one bad sample can't push its own probability to zero.

The still picks the unclipped term whenever it's smaller, so if an update moves the ratio the wrong way, there's still a gradient pulling it back. It isn't a hard constraint either: nothing stops from leaving the interval, and other samples keep moving the parameters. Clipping just removes the incentive to keep going in the direction that's already helping.

Total loss:

Putting it together, the loss has three parts:

The value term trains to predict the returns we actually observed, which is what feeds . The entropy term keeps the action variance from collapsing too early, so the policy keeps exploring. I annealed its coefficient from to over training.

In the browser the policy just outputs its mean action and never samples. The Gaussian only matters during training.

03Reward

PPO will optimize whatever reward you write down, loopholes included. My first few reward functions let the policy rack up points without ever stacking a crate, so I kept revising it based on what went wrong.

Sparse and shaped rewards:

My first attempt gave one point for placing a crate and zero otherwise. Nothing happened. A random policy would have to drive over, lower onto a crate, close, drive back, and open at the right height, all by accident, before it ever saw a nonzero reward. Until then there's nothing to learn from.

So I swapped it for a dense reward that pays out for getting closer to the current target. Here are the main terms in the final version:

is the distance from the gripper to its current target: the crate when it's empty, the drop slot when it's carrying one. This term does most of the work.

Failures and fixes:

Every other term is there because the policy did something I didn't want. The left column is what it did, and the right column is the fix.

it never let go
It carried the crate to the slot and just held it there, because holding was safer than letting go in the wrong spot. rewards closing near a crate and opening near the drop slot.
it fluttered
When closing near a crate earned a one-time bonus, it learned to close, back off, reopen, and do it again. Rewarding only the change in makes that round trip net to zero, so the only way to come out ahead is to end up holding the crate. This is the potential-based shaping idea from Ng, Harada and Russell [6], except I use the undiscounted difference .
it windmilled
It stacked the crates while spinning its shoulder about six full turns per episode. The sim squashes the policy's output before turning it into joint velocity, so once the command saturated, the penalty stopped growing even though the raw output kept climbing. Penalizing the raw output keeps pushing back after the joint hits its speed limit.
it hovered
It parked next to a crate and never grabbed it. A small cost on every step makes waiting around expensive.

04Training

Learning the whole task from scratch didn't work within my training budget, so I split it into four stages, each starting from the previous stage's checkpoint:

Curriculum stages:

1 · reach
Reach a random point in the workspace, no gripper. Gets within 2 px about 99% of the time.
2 · grasp
Reach the top crate. The episode ends when a closed gripper touches it.
3 · place
Starts already holding a crate next to the drop column, on either side. It just has to reach the slot and let go.
4 · pick and place
The full sequence from an empty gripper: drive, grasp, carry, place. The gap between columns grows from 90 px to 390 px over training.

Configuration:

policy
14 → 48 → 48 → 4, tanh, state-independent log σ
critic
14 → 64 → 64 → 1, tanh, separate network
rollout
128 parallel envs × 96 steps = 12,288 samples per update
ppo
, , , , 4 epochs per batch, minibatch 3072
update rule
Adam [4], , gradient norm clipped at 0.5
budget
32M environment steps, 407k episodes, about four minutes on a laptop CPU

05Further reading

If you want to go deeper, here's what I'd recommend, roughly in the order I'd read them:

06References

  1. [1]Schulman, Wolski, Dhariwal, Radford, Klimov. Proximal Policy Optimization Algorithms. OpenAI, arXiv:1707.06347, 2017.
  2. [2]Schulman, Levine, Moritz, Jordan, Abbeel. Trust Region Policy Optimization. arXiv:1502.05477, 2015.
  3. [3]Schulman, Moritz, Levine, Jordan, Abbeel. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv:1506.02438, 2015.
  4. [4]Kingma, Ba. Adam: A Method for Stochastic Optimization. arXiv:1412.6980, 2014.
  5. [5]Williams. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Machine Learning 8, 1992.
  6. [6]Ng, Harada, Russell. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. ICML, 1999.
back to homeaadi2 [at] stanford [dot] edu