Cobra Effect · Decisions and strategy
Explore and exploit
When to try something new, and when to stick with what works.
7 cards, read aloud in 2:39, with a test and sources.
It is Friday and you are hungry. The good place, or the new place?
You know the good place. It has never let you down. The new place might be better. It might be terrible. This is a small question with a whole branch of mathematics behind it.
Imagine a row of machines that pay out at different rates.
You do not know which one is best. You have a fixed number of pulls. Every pull spent finding out is a pull not spent on the best one. Every pull on your favourite is a pull that might have found a better one. Mathematicians call this the multi armed bandit.
It is an old problem, and a serious one.
William Thompson wrote about it in 1933, working on which medical treatment to give. Herbert Robbins set it out properly in 1952. It was never really about gambling. It is about how long to keep learning.
The answer turns on one thing. How much time is left.
In your first week in a new city, try everywhere. A bad meal costs one evening and buys you a map. In your last week, go to the best place you found. There is no future left to spend the information in. The same person, the same restaurants, opposite advice.
Which explains rather a lot about people at different ages.
A teenager tries everything. A grandparent has a favourite table and a favourite chair. Neither one is being foolish. They are solving the same problem with different amounts of time left. Brian Christian and Tom Griffiths make exactly that argument in Algorithms to Live By.
The good strategies all do one thing. They explore less as they learn more.
A simple one spends a small fixed share of your choices on something new and the rest on your current best. A better one tries whichever option could still turn out highest, so promising unknowns keep getting a look. None of them ever stops exploring completely.
So the question is not which option is best.
It is how many more choices like this one you are going to get. Long horizon, take the risk. The information is worth more than the meal. Short horizon, go to the good place. And if you have stopped trying anything new, your horizon just got shorter than it really is.
Sources
- Multi-armed bandit, Wikipedia. The problem, the classic strategies from Thompson in 1933 and Robbins in 1952 onwards, and where it gets used now in trials and in testing.
- Algorithms to Live By, Brian Christian and Tom Griffiths, 2016. The opening chapters are the best plain English account of when to keep looking and when to settle, with the restaurant version and the hiring version.
- Gittins index, Wikipedia. John Gittins showed in the 1970s that the problem has an exact answer under certain conditions. Heavy going, and worth knowing that it exists.
Nearby ideas
- Satisficing. Good enough, chosen quickly, often beats the endless hunt for best.
- Choice overload. Too many options that are hard to compare can stop people choosing.
- The Eisenhower matrix. Urgent things crowd out important ones unless you protect them.
- First principles. Rebuild a problem from what must be true, not from what everyone does.
- Opportunity cost. The real price of anything is the best thing you gave up for it.
- Chesterton’s fence. Before removing a rule, find out why it was put there.
- The sunk cost fallacy. Money already spent is gone, so it should not steer the next choice.
- Inversion. To find the path to success, first ask what would guarantee failure.