I came to robotics late, and from the wrong direction. I'd spent years on machine learning: sequence models, speech, language models. Robotics looked like a different field, full of motors and torque and matrices with dots on them. It took me a while to notice that it wasn't a different field. It was the same field with a harder output.
A language model reads text and writes text. A robot reads sensors and writes forces. That's the whole difference. Everything else in robotics, all the mechanics that made it look intimidating, is there because forces are harder to write than words. Words don't have mass. They don't fall over or crush the cup.
1. Physics
You can't skip physics. I tried. The trouble is that every robot learning paper assumes you know what a pose is, what a Jacobian is, why torque matters. There are only about five ideas, though, and they're all simple once someone explains them without the notation.
1.1 Pose and SE(3): Where is the object?
To pick up a cup, a robot needs to know two things: where the cup is, and which way it's facing. Together these are the cup's pose.
Position is three numbers, (x, y, z). Orientation is three more, (roll, pitch, yaw). So a rigid object has six degrees of freedom, and SE(3) is just the name for the math that keeps track of all six at once. That's all it is. When you see SE(3) in a paper, read "where it is and which way it's pointing."
The reason this matters is that "there is a cup" isn't enough to act on. The robot needs the cup's pose, and its own hand's pose, and it needs both in the same coordinates.
Pose is where plus which way.
1.2 Jacobian: How does joint movement become hand movement?
A robot arm is a chain of joints: shoulder, elbow, wrist. The motors live at the joints. But nobody cares about the joints. What you care about is the thing at the end, the gripper, which roboticists call the end effector.
The Jacobian is the translation between the two:
ẋ = J(q) q̇
Here q is where the joints are, q̇ is how fast they're turning, and ẋ is how fast the hand moves as a result. You don't need to remember the equation. What you need to remember is that when you say "move the hand ten centimeters left," several joints have to turn at once by different amounts, and the Jacobian is what works out which.
The Jacobian turns joint motion into hand motion.
1.3 Singularity: When a robot gets into a bad pose
Sometimes an arm gets into a configuration where some direction is just hard to move in. Straighten your own arm completely and try to move your hand a little further away from you. You can't. You've run out of arm. That's a singularity.
What surprised me is that the problem isn't the target. The target might be fine. The problem is the shape the robot is currently in. A lot of motion planning is about staying away from these shapes.
A singularity is a pose where the robot loses easy control in some direction.
1.4 Torque: How hard should the motor push?
A motor doesn't get to say "go to thirty degrees." It has to actually produce enough twisting force to get there, against gravity and against whatever it's holding. That twisting force is torque.
τ = r F sin θ
The formula is less important than the intuition, which you already have from using a wrench: the farther from the pivot you push, the more turning you get.
Torque is where the stack bottoms out. Every command, no matter how high-level, ends up as command → controller → motor → torque → movement. There is no other way to move something.
Torque is how hard a motor is trying to rotate a joint.
1.5 Dynamics: How much force is needed to create motion?
Dynamics is the study of how force turns into motion. For a robot it's usually written as
M(q) q̈ + C(q, q̇) q̇ + g(q) = τ
which looks worse than it is. Read it left to right and it says the torque you need is the sum of three things: how hard the robot is to accelerate, the extra forces it makes by already moving, and gravity. Motor torque = inertia + motion effects + gravity. The next three sections are just those three terms.
One piece of notation first. q is where the joints are. q̇ is how fast they're moving. q̈ is how fast that speed is changing. Position, velocity, acceleration. Same as high school physics, just with more joints.
Position says where the robot is. Velocity says how it's moving. Acceleration says how that's changing.
1.6 Inertia: M(q) q̈
This is F = ma wearing a costume. Heavier things are harder to get moving. A robot is more complicated than a block because it's a chain of parts that rotate, and how hard it is to accelerate depends on the shape it's in. An arm stretched out is harder to swing than an arm tucked in, for the same reason a figure skater spins faster with their arms pulled in. M(q) is the term that keeps track of this.
Inertia is how hard this body is to speed up or slow down.
1.7 Motion effects: C(q, q̇) q̇
Joints affect each other. Swing the shoulder fast and the elbow gets dragged along, and the wrist after it. So a robot that's already moving is generating forces just from moving. This term collects those. Coriolis and centrifugal effects, if you want the names. You mostly don't.
A moving robot makes extra forces just by moving.
1.8 Gravity: g(q)
Gravity never stops. Even a robot that's perfectly still has to push against it, or the arm falls. g(q) is how much torque it takes just to hold the current pose. This was a small revelation to me: standing still costs energy.
Standing still can still require force.
1.9 Dynamics vs. Taylor expansion
Both of these use position, velocity, and acceleration, and it took me a while to see that they answer opposite questions.
Dynamics asks: if I push this hard, how fast will things accelerate?
Taylor expansion asks: given how things are moving now, where will they be a moment from now?
q(t + Δt) ≈ q(t) + q̇(t) Δt + ½ q̈(t) Δt²
Put them together and you get the mental model I actually use: τ → q̈ → q̇ → q → new state. Torque changes acceleration, acceleration changes velocity, velocity changes position. Every simulator in the world is running this loop.
Torque changes acceleration. Acceleration changes velocity. Velocity changes position.
1.10 Contact: What happens when the robot touches something?
Moving through empty air is the easy part. The hard part is touching things. When a robot touches something it has to deal with force, friction, slipping, stiffness, how much the thing deforms, and exactly where the contact is. Grip a cup too gently and it slips. Too hard and it cracks.
This is why manipulation is so much harder than it looks. Getting to the right position is a solved problem. Doing something useful once you're there isn't.
Robotics gets much harder the moment the robot touches the world.
2. Robot systems
That's the body. Now the loop that runs on it. Every robot, from a Roomba to a humanoid, runs some version of perception → state estimation → planning → control → actuation. If you've built an ML system, most of this will look familiar, just with a physical world on both ends.
2.1 Perception: What do I see?
Robots have sensors: cameras, depth cameras, LiDAR, microphones, IMUs, force sensors, tactile sensors, joint encoders. Sensors give you raw numbers. Perception is turning those numbers into something useful: pixels → object → position → pose.
The important thing, and the thing that separates people who've shipped robots from people who haven't, is that a camera doesn't give you the truth. It gives you an observation. Observations are wrong all the time.
Perception turns sensor data into something the robot can act on.
2.2 State estimation: What is probably true?
Because sensors lie, a robot shouldn't believe any single reading. They're noisy, they drop frames, they lag, things get occluded, calibration drifts. So instead the robot keeps a belief and updates it: what it thought before, what it sees now, and what it just did.
Here's the example that made this click for me. The robot sees a cup. Then its own arm passes in front of the camera. A naive system concludes the cup is gone. A robot with state estimation concludes the cup is probably still there. That's the whole idea.
State estimation is the robot's best guess about reality.
2.3 Planning: What should I do?
Planning takes the current state and a goal and produces a sequence of actions: approach → grasp → lift → move → place. It answers "what should happen next?" and nothing else.
Planning picks the sequence of actions.
2.4 Control: How do I make it actually happen?
A plan says "move the hand here." A motor has no idea what that means. Control is the layer in between. It turns the plan into positions, velocities, torques, currents, and then, crucially, it checks. It measures error = desired − actual and corrects, hundreds of times a second, forever.
I think of it this way. Planning is deciding what should happen. Control is making it happen, and noticing when it doesn't.
Control is a feedback loop: observe, correct, repeat.
2.5 Actuator: Where software becomes physical
An actuator is the thing that actually moves: an electric motor, a hydraulic cylinder, a pneumatic piston. It's the last link in the chain, bits → electrical signal → actuator → force → movement, and the only one that isn't software.
Actuators are where software becomes physical.
3. Robot learning
Classical robotics wrote the rules by hand. Robot learning tries to learn them, from demonstrations, from data, from trial and error. This is the part of the field where I felt at home, because it is, in a very literal sense, machine learning. The four ideas you'll run into are reinforcement learning, imitation learning, diffusion policies, and vision-language-action models.
3.1 Reinforcement learning: Learn from consequences
In RL nobody shows the robot what to do. It tries something, sees what happens, gets a reward, and adjusts: act → see result → get reward → improve → try again.
It's slow, because the robot has to discover everything for itself, and it's dangerous, because a robot that's still discovering things is a robot that knocks over furniture. But when it works, it finds behaviors nobody would have thought to demonstrate.
RL learns through trial, feedback, and reward.
3.2 Imitation learning: Learn by copying
Watch someone do the task, then do what they did. The training data is (observation, action) pairs, and you fit a policy π(a | s): given state s, which action a? The simplest version is called behavioral cloning, and it's exactly supervised learning. That's not an analogy. It's the same code.
Some people call this IRL, inverse reinforcement learning. Strictly, that's the version where instead of copying the actions, you infer the reward the demonstrator was optimizing and then learn from that. Either way the signal comes from someone else's behavior. That's the difference from RL, which learns from the consequences of its own.
Imitation learning is learning from examples of correct behavior.
3.3 Diffusion policy: Generate a whole movement
A simple policy predicts one action at a time. But movements are sequences, and predicting them one step at a time produces jerky, indecisive robots. A diffusion policy generates a whole short trajectory at once, by the same trick image models use: start with noise and refine it, noisy actions → refine → refine → smooth trajectory.
The reason this works well is that most tasks have several equally good solutions. You can reach for the cup from the left or the right. A model that averages them reaches for the middle and knocks it over. A diffusion model can hold both in its head and commit to one.
Diffusion policy generates coordinated action sequences instead of one action at a time.
3.4 VLA: Vision-language-action
A vision-language-action model takes what the robot sees and what a person said, and produces actions: (vision, language) → action. You say "put the red cup in the sink" and the model has to connect the words to the pixels to the motion. If you've followed language models at all, this is the obvious next thing to build, and lots of people are building it.
VLA models try to turn instructions directly into behavior.
4. World models
A policy asks "what should I do?" A world model asks "what will happen if I do it?" These sound similar and are completely different, and the difference is most of what's interesting in robotics right now.
4.1 One-step world model
Take the current state sₜ and an action aₜ. A world model predicts p(sₜ₊₁ | sₜ, aₜ): the next state. That's it. It's learning the transition function of the world.
World models predict the consequences of actions.
4.2 Physics and world models are closely related
Here's the thing I wish someone had told me at the start. Physics is a world model. The dynamics equations from section 1 predict what happens next: τ → q̈ → q̇ → q. A learned world model predicts the same thing from data: (sₜ, aₜ) → sₜ₊₁. The only difference is where the transition rule comes from. Physics derives it. Learning fits it. They're answering the same question.
4.3 Long-horizon world models
One step isn't much. What you want is to predict a whole future, p(sₜ₊₁ … sₜ₊ₕ | sₜ, aₜ … aₜ₊ₕ₋₁): if I do this sequence of things, what happens? Then the robot can:
- imagine an action,
- predict the result,
- imagine a different action,
- compare,
- pick one,
- do it.
Which is to say, it can think before it moves. That sounds like a small thing. It's the thing.
A world model lets the robot imagine before acting.
5. Robotics abstraction
The full robot loop
perceive → estimate → predict → plan → act → observe
Each step is one question:
- Perception — what do I see?
- State estimation — what's probably true?
- World model — what happens if I act?
- Planning — what should I do?
- Control — how do I actually do it?
- Learning — how do I do better next time?
The core idea
Robotics starts with physics, force → motion. Learning adds behavior, observation → action. World models add prediction, (state, action) → future state. The whole thing, seen from far enough away, is human intent → physical intelligence → real-world action.
What I eventually realised is that none of the steps were foreign. Perception is representation learning. State estimation is inference under noise. World models are sequence prediction. Planning and control are what search and optimization look like when the output has mass. The physics isn't a different subject. It's the constraint the output has to satisfy.
You don't get to ignore physics. But you don't have to expose it either. The goal is to understand it well enough to hide it behind good software, which is what we've done with every other hard thing computers do.
That's what it means to turn robotics into software.