Teaching a pixel robot to stack crates
A 3,268-parameter policy, trained with PPO in a NumPy sandbox, running live in your browser at 60 Hz.
by aadi shah·August 2026·9 min read
The live arm needs a wider screen, so it's hidden on mobile. The full write-up is below.
Nothing about the arm above is keyframed. Every frame, a small neural network reads fourteen numbers from the simulator and outputs four commands: drive the base, turn each of the two joints, and open or close the gripper. It learned everything it does over 32 million simulator steps of trial and error, guided by the reward described below.
01Problem formulation
A task like this is usually framed as a Markov decision process. At each step the agent sees a state, a compact summary of what's going on, and picks an action. The environment moves forward and hands back a reward. That reward is just a score for that one step. It doesn't tell the agent what it should have done.
Reinforcement learning is about learning a policy from that feedback by trial and error. What counts is the total reward over the whole episode, so an action can be worth taking even if it pays off little right away.
The Markov part means the current state and action are enough to predict what happens next. The policy doesn't need to know how the arm got to its current pose. If that history mattered, it would have to go into the state.
What we're actually learning is the policy, a function that takes a state and returns a distribution over actions:
is the weights of a small neural network. During training the policy samples from that distribution so it keeps trying new things. In the browser it just takes the mean. Both states and actions are continuous:
02Learning with PPO
So the setup defines what the policy sees, what it can do, and how it's scored. What it never gets is the right answer for any state.
With nothing to imitate, the policy has to act, see what reward comes back, and adjust. That makes training online: the policy generates its own data in the simulator, and every update changes what the next rollout looks like.
Policy gradients:
We want to maximize the expected return . Here is one rollout and discounts rewards that arrive later. The catch is that we can't backprop through the simulator. The policy gradient theorem gets around this by rewriting the gradient in terms of the policy's log-probability, which we can differentiate with respect to :
Value and advantage:
Intuitively, the update makes actions that went better than expected more likely and actions that went worse less likely. To pin down “expected”: the value of a state is the return the policy expects from there on, and the action-value is the same thing if you commit to one action first. The gap between them is the advantage:
means the action beat the policy's own average from that state, and means it fell short. We don't know or , so a second network, the critic , learns , and is estimated from the critic plus the rewards we actually saw. Subtracting as a baseline doesn't bias the gradient. It just cuts the variance.
In practice you estimate the expectation by averaging over a batch of steps. That's REINFORCE [5] with a baseline. It's simple and unbiased, but noisy enough that big update steps aren't safe.
Advantage estimation:
We never observe the advantage directly, so every step needs an estimate built from the rewards that followed it and the critic's predictions. The simplest version looks one step ahead, and it's called the TD residual:
That's low-variance, but it's only as good as . At the other extreme, Monte Carlo sums the actual rewards to the end of the episode. That's unbiased, but much noisier. GAE [3] blends the two using :
recovers the one-step estimate, the Monte Carlo return minus the baseline, and anything in between is a geometric blend. I used , which keeps variance manageable over 600-step episodes without leaning too hard on a critic that starts out pretty bad.
Clipped objective:
With an advantage for every step, the remaining question is how big a step to take. There's a catch: means the rollouts came from the current policy, so the gradient estimate only holds for the parameters that collected them. That's what makes the method on-policy. After one update the batch describes an older policy, and reusing it without correction makes the estimate worse. So a big update can hurt twice: it can make the policy worse, and it can leave you with data that no longer describes it.
TRPO [2] handles this with a hard limit on the KL divergence between the old and new policy (a measure of how different two distributions are). It enforces that limit with second-order methods, which use curvature as well as slope and are expensive. PPO [1] replaces this with a clipped first-order objective. You can take several gradient steps on the same batch, but there's no extra reward for pushing any one action's probability too far in a single update.
To do that, PPO tracks how much each sampled action's probability has changed:
When the data is collected, . During the update it drifts: above 1 means the action got more likely, and below 1 means less likely. PPO clips the ratio once it moves past in the direction the advantage is pushing it:
How the clip plays out depends on the sign of the advantage.
Positive advantage (). Making the action more likely helps, but only up to . Past that, the clipped term is smaller, the picks it, and since it's constant in , that sample stops contributing gradient.
Negative advantage (). Now the objective wants the action to become less likely. Below the term goes flat, so one bad sample can't push its own probability to zero.
The still picks the unclipped term whenever it's smaller, so if an update moves the ratio the wrong way, there's still a gradient pulling it back. It isn't a hard constraint either: nothing stops from leaving the interval, and other samples keep moving the parameters. Clipping just removes the incentive to keep going in the direction that's already helping.
Total loss:
Putting it together, the loss has three parts:
The value term trains to predict the returns we actually observed, which is what feeds . The entropy term keeps the action variance from collapsing too early, so the policy keeps exploring. I annealed its coefficient from to over training.
In the browser the policy just outputs its mean action and never samples. The Gaussian only matters during training.
03Reward
PPO will optimize whatever reward you write down, loopholes included. My first few reward functions let the policy rack up points without ever stacking a crate, so I kept revising it based on what went wrong.
Sparse and shaped rewards:
My first attempt gave one point for placing a crate and zero otherwise. Nothing happened. A random policy would have to drive over, lower onto a crate, close, drive back, and open at the right height, all by accident, before it ever saw a nonzero reward. Until then there's nothing to learn from.
So I swapped it for a dense reward that pays out for getting closer to the current target. Here are the main terms in the final version:
is the distance from the gripper to its current target: the crate when it's empty, the drop slot when it's carrying one. This term does most of the work.
Failures and fixes:
Every other term is there because the policy did something I didn't want. The left column is what it did, and the right column is the fix.
04Training
Learning the whole task from scratch didn't work within my training budget, so I split it into four stages, each starting from the previous stage's checkpoint:
Curriculum stages:
Configuration:
05Further reading
If you want to go deeper, here's what I'd recommend, roughly in the order I'd read them:
- Reinforcement Learning: An Introduction (Sutton & Barto) for the foundations: MDPs, value functions, policy gradients.
- CS224R (Chelsea Finn, Stanford) for modern deep RL, with a more practical focus.
- Deep RL Course (Hugging Face) for guided exercises that train several kinds of agents.
- Spinning Up (OpenAI) as a reference: short derivations next to readable single-file implementations.
06References
- [1]Schulman, Wolski, Dhariwal, Radford, Klimov. Proximal Policy Optimization Algorithms. OpenAI, arXiv:1707.06347, 2017.
- [2]Schulman, Levine, Moritz, Jordan, Abbeel. Trust Region Policy Optimization. arXiv:1502.05477, 2015.
- [3]Schulman, Moritz, Levine, Jordan, Abbeel. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv:1506.02438, 2015.
- [4]Kingma, Ba. Adam: A Method for Stochastic Optimization. arXiv:1412.6980, 2014.
- [5]Williams. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Machine Learning 8, 1992.
- [6]Ng, Harada, Russell. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. ICML, 1999.