Nicole Junkermann on Patience, Practice and Reinforcement Learning
Machines that learn by trying arrive the way anybody does: many attempts, honest feedback and more patience than the story usually admits.

Reinforcement learning is a plain idea wearing a technical name: try something, find out how it went, try again.
Nicole Junkermann has a weakness for ideas that turn out to be simpler than their vocabulary, and this is one of them. A system is set a task, it makes an attempt, and it receives a signal telling it whether that attempt went better or worse than the one before. Then it tries again, and again, until something that works has been found.
Trying, and finding out
Nobody writes the solution down in advance. Somebody writes down what counts as an improvement, and the system searches for its own route towards it across an enormous number of attempts. The route it settles on is frequently not the one a person would have chosen, which is part of what makes the process interesting and part of why it needs watching.
A family likeness
Anyone who has learned anything slowly will recognise the shape of this. The resemblance to ordinary practice is less a metaphor than a family likeness, and it turns up in the same few places.
- A single attempt says almost nothing. The information is in the hundredth one.
- The feedback has to measure what was actually meant, or the wrong thing quietly improves.
- Early results look discouraging, and looking discouraging is not the same as going wrong.
- Progress arrives unevenly, in long flat stretches with sudden steps in them.
- Someone still decides what counts as good, and that decision comes before everything else.
The middle stretch
Accounts of this kind of learning tend to skip to the result, which makes it sound like a conjuring trick. The interesting part is the middle: the long run of attempts that are not yet any good, continued because the direction is right even while the evidence is thin. That is not a technical quality at all. It is the same patience a quiet hour asks for, and the same one that fills a notebook long before it fills a page worth keeping. The carnets are largely a record of that middle stretch.
Keeping the claim small
It is worth being careful about what is being described. A system that learns this way does not understand the task, hold an opinion about it or want anything from it. It finds behaviour that fits a reward somebody chose, inside a narrow frame, and outside that frame it has nothing at all. Which hands the interesting questions straight back to people: what is worth asking for, and what should count as a good answer.
Why it belongs in a journal about looking
Practice is the common ground. Both kinds of learning depend on repetition that looks unproductive from outside, on feedback that is honest rather than flattering, and on somebody deciding in advance what would count as better. What these tools are teaching us about looking takes the same question from the other side, and a short introduction sits here for anyone who would like the context first.
The interesting part is not the result. It is the long run of attempts that are not yet any good.
Frequently asked questions
What is reinforcement learning, in one sentence?
A system attempts a task, receives a signal saying whether the attempt went better or worse than the last one, and adjusts the next attempt accordingly, over a very large number of rounds.
Why does patience come into it?
Because the early rounds look discouraging and the progress is uneven. The long middle stretch of attempts that are not yet any good is where the result is actually made, and it is usually the part left out of the telling.
Does a system like this understand what it is doing?
No. It finds behaviour that fits a reward somebody chose, inside a narrow frame. Deciding what task is worth setting, and what should count as a good answer, stays with people.



