Why might injury-minimising programming give people a reason to buy a less safe car?
Because the rule targets the vehicle that will survive the crash better — so a rational buyer might reduce their chance of being chosen at all by driving something that scores badly in crash tests.
* Any optimisation applied to agents who can see it changes the behaviour it was measured against *
The chain of reasoning is uncomfortable but straightforward:
- The disadvantages of being in a particularly safe vehicle would be entirely foreseeable, and a rational actor would let them influence their choice of car.
- They would have to weigh minimising the risk of being in a crash at all against maximising protection if one happens — the two now pull in opposite directions.
- Depending on how strongly the choice of vehicle affects the probability of being involved in a crash, it could suddenly become irrational to buy the safest car available.
- To protect themselves and their family as well as possible, one might be better advised to choose a car that does relatively badly in crash tests, because the risk of being involved in a collision at all could be considerably lower.
Perverse incentives of this kind can undermine the effort to avoid casualties in road traffic — and, depending on how strongly people let themselves be influenced, could drive the goal of injury minimisation ad absurdum: a rule adopted to reduce injuries ends up filling the roads with vehicles that injure their occupants more.
Two points are worth adding, both of which the paper concedes. There is no moral duty to sacrifice oneself for others — praiseworthy as it would be. And one could invoke a special duty of care towards one's own family. So the person gaming the rule is not obviously behaving badly; they are responding sensibly to a badly designed rule. That is what makes it a design problem rather than a moral failing.
Tip: This is the general failure mode of any optimisation applied to agents who can see it. The rule changes what people do, so a rule evaluated against today's behaviour will be evaluated wrongly.
Go deeper:
Wikipedia: Perverse incentive — the general failure mode, with the cobra-effect examples that make it memorable.
Wikipedia: Goodhart's law — »when a measure becomes a target it ceases to be a good measure« — the same problem in one sentence.