Part 13 of 16History

When History Ran the Experiment

Twelve parts have ended at the same wall: you cannot assign women to different rules and see what happens. But history has assigned them, repeatedly, and sometimes almost cleanly. This part goes and looks at the results.

Before We Begin — Where We Left Off

Part Twelve was about enforcement in a world where the community stopped having edges. Reputation, it turned out, was never a feeling — it was a working technology with three properties that followed from physics rather than from anyone’s choice: a bounded audience, a fading memory, and an enforcer with a return address. Those three were also a family’s three remedies. You could wait, you could move, or you could negotiate. All three broke at once and nothing replaced them.

The finding I did not expect came from Chapter Six, and criminals produced it rather than researchers. Sextortion has split into two patterns. Girls are extorted for further images and for compliance; boys are extorted for money, fast. The reason is that a threat against a boy loses its value within days while a threat against a girl in a strict community never expires. Two sets of extortionists, working independently and with no interest in the question, have therefore priced the difference between what a community will do to a son and what it will do to a daughter. Their estimate is the most precise measurement in this entire series.

And Part Twelve ended where every part has ended: with a mechanism identified, a remedy that reaches the wrong layer, and no way to test the counterfactual.

That is the wall. Thirteen parts of argument and I have said some version of the same sentence in every one of them — the experiment cannot be run. You cannot assign a thousand women to strict rules and a thousand to loose ones and come back in thirty years. No review board would approve it and no researcher would propose it.

But history has done it. Repeatedly. Sometimes almost cleanly.

Countries have reversed their rules overnight. Populations have been cut in half by a line on a map and governed differently for forty years. Constraints that stood for centuries have been lifted in a single announcement. And in one case — the best evidence in this entire series, and it is Indian — a government actually randomised.

This part is the payoff. It is also the part where I have to be most careful, because a natural experiment is the easiest kind of evidence to abuse, and Chapter One is about how.

How this document is built

Six kinds of box, each doing one job, each looking different so you can see at a glance what you are about to read.

The first explains a hard word the moment it appears.

Word Box

Counterfactual: what would have happened otherwise. The thing you are always comparing against and can never observe.

If a country changes its rules and something happens afterwards, the question is not what happened. It is what would have happened without the change. That world does not exist, so every claim in this part is an argument about an imaginary place, and the whole craft consists of finding the least imaginary substitute.

The second takes a number too big or too abstract to picture and gives it a body.

In Real Terms

“A country’s economic output fell by half a per cent.” That sounds like a rounding error in a budget document.

For Afghanistan it means that the goods and services the country produces in about two days of every year have simply stopped existing, permanently, every year, and the reason is that girls are not allowed past the sixth grade.

The third shows the actual evidence and then says what it cannot show. In this part it does more work than anywhere else in the series.

How We Actually Know This

There is a hierarchy of evidence in this part and it is worth stating up front, because the cases are not equally good.

Strongest: a genuine randomisation, where who got the change was decided by something with no connection to the outcome. There is exactly one of these in this series and it is in Chapter Two.

Strong: a sharp change in one place with a genuinely similar place alongside it that did not change.

Moderate: a sharp before-and-after in one place, with the trend beforehand visible so you can see whether the line bent.

Weak, and very common: two countries that differ, compared after the fact, with a story attached.

Almost everything quoted in public argument about women is the fourth kind. This part is organised so you can always tell which kind you are reading.

The fourth is for places where informed people genuinely disagree. Each side gets its best case.

The Argument — a worked example of the box

A question, then the sides, then a verdict that does not claim more certainty than exists.

One side

Its strongest case, put the way its best advocate would put it, with the evidence it actually has.

The other side

The same, with equal care. If a position sounds foolish here, I have failed to understand it, not proved it wrong.

Where things stand: what is genuinely agreed, and where the disagreement really sits.

What would settle it: the evidence that would decide it. Sometimes the honest answer is that nothing available would.

The fifth is the signature of this series. It digs out what both sides are assuming without noticing.

The Hidden Assumption

Not caveats. Unexamined premises sitting underneath an argument everybody is having. There are five in this part, and the last one is turned on the way I built this part.

The sixth closes every chapter in the plainest words available.

Remember This

Six boxes. Word explains. In Real Terms converts. How We Actually Know This shows the evidence. The Argument gives every side its best shot. The Hidden Assumption digs underneath. Remember This closes the chapter.

Simple words, serious content, nothing left out.

Three warnings specific to this part

The first is about how easy this material is to abuse. Case studies are the most persuasive and least reliable form of evidence there is. A vivid country comparison will convince a reader of almost anything, and both sides of this argument have a favourite one they deploy as though it settled the matter. I have tried to include the cases that cut against me as carefully as the ones that do not, and Chapter Eight exists specifically to count the places where the traditionalist forecast was right. If that chapter reads as grudging, I have failed at it.

The second is about live events. Chapter Five is about Afghanistan, where the experiment is running now. Almost everything published about its consequences is a projection — a model of what will happen, not a measurement of what has. I have marked every one of those, because the difference between “maternal deaths rose by half” and “maternal deaths are projected to rise by half” is the difference between evidence and a forecast, and the reporting collapses it constantly.

The third is about me, and it belongs at the front rather than buried in Chapter Nine. Nobody handed me a list of natural experiments. I chose them. That is a serious problem for a part whose first chapter warns about choosing cases, and I have not solved it — I have only made it visible. The last Hidden Assumption box in this part is about exactly that, and it is the most uncomfortable thing I have written in thirteen parts.

1What a Natural Experiment Is

The whole value of this part depends on knowing which comparisons are worth something and which are stories. The rule turns out to be simple, and it disqualifies almost everything you have ever seen quoted.

1.1 — Why a coin toss is worth more than a country

Suppose you want to know whether a medicine works. You could give it to a hundred people and see how they do. But the people who took it chose to take it, or a doctor chose for them, and those choices track everything else about them — how ill they were, how much money they had, whether they follow instructions.

So instead you toss a coin. Heads gets the medicine, tails does not. Now the two groups differ in exactly one respect and every other difference has been scrambled by the coin, including the differences nobody thought to measure.

That last clause is the entire magic of it. Statistical control can only remove factors you know about and measured. Randomisation removes the ones you never thought of, for free.

Word Box

Natural experiment: a situation in the real world where something outside the researcher’s control assigned people to different conditions in a way that had nothing to do with the outcome being studied.

The word “natural” means nobody arranged it as research. A war, a border, a law passed on a particular date, a lottery for a permit. The value comes from the assignment being unrelated to the thing you want to measure, not from it being accidental.

Why it matters: a good natural experiment gets you close to the coin toss without the coin. A bad one is two countries that differ, with a story attached.

1.2 — The checklist

Six questions separate the good ones from the stories. I will apply them to every case in this part, and you should apply them to every case anybody quotes at you afterwards.

One. Was the assignment unrelated to the outcome? If a country liberalised its laws because women were already gaining ground, then the law did not cause what followed — the same underlying change caused both.

Two. Is there a comparison group? Before-and-after in one place tells you what happened next, which is not the same as what the change did.

Three. Were the two sides on the same path beforehand? If the lines were already diverging before the change, the change is not what separated them.

Four. Was the change sharp in time? A reform that arrived over twenty years cannot be distinguished from anything else that happened over those twenty years.

Five. Did the change come alone? A revolution alters a thousand things at once. Whatever follows, you cannot attribute it to any one of them.

Six. Is the outcome measured the same way on both sides? Two countries counting “employment” or “violence” differently will show a difference that exists only in the definitions.

In Real Terms

Run those six on the comparison people reach for most often — “look at Sweden and look at Afghanistan.”

Was the assignment unrelated to the outcome? No: Sweden is not governed differently by accident. Same path beforehand? They were never on the same path. Sharp in time? Neither change was. Did it come alone? Everything else about the two countries also differs — income, climate, history, war, oil, geography, religion, colonialism, literacy in 1900.

That comparison fails five of the six tests and it is the single most-quoted piece of evidence in this entire argument, by both sides. It establishes that two very different countries are very different, which we knew.

1.3 — What history can supply

Having disqualified the easy comparisons, what is left is genuinely useful, and there is more of it than people expect. Four kinds turn up repeatedly.

A line drawn through one people. Germany divided, Korea divided. Same language, same ancestry, same starting point, different rules, forty years. This is as close as political history comes to a laboratory, and Chapter Three is about how close that actually is.

An abrupt reversal. A society that was moving one way is turned round in a year by a change of regime. Iran is the cleanest example on Earth and Chapter Four is entirely about it.

A constraint lifted. One specific thing that was forbidden becomes permitted, on a known date, with everything else roughly unchanged. Saudi Arabia’s driving ban is nearly ideal on the checklist.

A rule applied by lottery. Rare, precious, and India has one.

The Argument — can history substitute for an experiment at all?

Some people think this whole enterprise is a category error. That is a serious position and the case for it is strong.

It cannot, and the pretence is worse than nothing

Societies are not laboratories and the differences between them are not noise to be scrubbed out — they are the substance. Every “natural experiment” involving whole countries fails at least one of the six tests, usually the fifth: revolutions, wars and reforms change everything simultaneously and there is no honest way to isolate one strand. Worse, the method launders an argument. A researcher picks two places, finds a difference, and the machinery of comparison gives a hunch the appearance of a finding. If you want to know about human societies, read history properly rather than dressing it up as data.

It is imperfect and it is what we have, and imperfect beats nothing

The alternative to flawed evidence is not perfect evidence; it is opinion, tradition and confident assertion — which is exactly what filled this space for three thousand years. The checklist above is not a rubber stamp; it disqualifies most of what gets quoted, including things people on my side of the argument like. And some cases genuinely do pass: a lottery is a lottery, a border is a border, and a driving ban lifted on a Sunday is a sharp change in one variable. Rejecting the method wholesale means having no way to be surprised, and being surprised is the only thing that distinguishes inquiry from advocacy.

Where things stand: the honest position is graded rather than binary. The randomised case in Chapter Two is real evidence by any standard, and nobody serious disputes it. The border comparisons are strong but contaminated. The revolutions are vivid and weak. What is not defensible is treating all four as the same kind of thing, which is what public argument does constantly and in both directions.

What would settle it: nothing, because this is a dispute about method rather than fact. But it can be made honest: state which tier of evidence each case belongs to before drawing anything from it, which is what the front matter of this part does.

Why people care so much: because a vivid country comparison is the most persuasive rhetorical object available and everyone has one. Grading the evidence takes almost everybody’s favourite argument away.

There is a deeper problem than any of the six tests, and it applies to this document rather than to the cases in it.

The Hidden Assumption — that a difference we noticed is an experiment

Every case in this part, and every case anybody cites in this argument, arrived the same way: somebody noticed it. Iran is famous because the reversal was dramatic. Germany is famous because the division was total. Saudi Arabia is famous because the change was sudden and recent.

Now ask what kind of case never becomes famous. The one where a society changed its rules and nothing measurable happened. There is no book about that. No researcher builds a career on it, no newspaper reports it, and it does not become an example anybody quotes — because a null result is not a story.

So the set of natural experiments available to any writer is not a sample of what happens when rules change. It is a sample of the occasions on which something visible happened. That is selection on the outcome, and it is the oldest error in empirical work.

The consequence is specific and it cuts against both sides equally. If the rules about women mostly do not matter much, we would still expect a handful of dramatic cases where they appeared to, and those cases would be the famous ones, and everybody would argue using them. The world in which rules matter enormously and the world in which they mostly do not look identical from inside a list of famous examples.

The general form: sampling on the outcome and calling it evidence. It runs through everything built on cases. Successful founders studied for their habits, with the identical habits of the failures unrecorded. Cures credited to a remedy, with the people it did not cure not writing in. Historical lessons drawn from the wars that were decisive.

There is a partial defence and it is worth naming because it is what this part relies on. Where a case includes a proper comparison group — a border, a lottery, a neighbouring district — the null results are inside the study rather than missing from it. Those cases survive this problem. The vivid single-country reversals do not, and they are the ones everybody quotes.

With that grading in hand, the rest of this part can be read for what each case is worth rather than for how vivid it is. We start with the only one that needs no grading at all.

Remember This

The value of a coin toss is not fairness. It is that randomising removes the differences nobody thought to measure. Statistical control can only remove the ones you knew about.

A natural experiment is a real-world situation where something outside anyone’s control assigned people to different conditions for reasons unrelated to the outcome.

Six tests: was assignment unrelated to the outcome, is there a comparison group, were the two on the same path before, was the change sharp, did it come alone, and is the outcome measured the same way on both sides.

“Look at Sweden, look at Afghanistan” fails five of the six, and it is the most-quoted evidence in this argument by both sides.

And the deepest problem is that these cases became famous because something visible happened. A world where the rules matter enormously and a world where they mostly do not look identical from inside a list of famous examples.

2The One Time Somebody Randomised

The best evidence in this entire series comes from Indian village councils, and it exists because a constitutional amendment accidentally built a lottery. It is the only case here that passes all six tests.

2.1 — The accident in the amendment

In 1993 India amended its constitution to devolve power to village councils, and wrote in a requirement that one-third of council seats — and one-third of the council leader positions — be reserved for women.

The reservation itself is well known. What matters here is the mechanism used to decide which villages got a reserved leadership in any given election. Rather than letting states choose, the rule rotated the reservation through villages by serial number.

A village’s serial number has nothing to do with how progressive it is, how rich it is, how its women live, or anything else. So the question of whether a village would be led by a woman was decided by something completely unrelated to the outcomes anyone wanted to study.

That is a lottery. Nobody intended it as one. It is the closest thing to a coin toss that exists anywhere in this argument.

How We Actually Know This

Run the six tests from Chapter One on this case.

Assignment unrelated to outcome? Yes — serial number rotation. Comparison group? Yes — the two-thirds of villages not reserved that cycle, in the same district, under the same state government, at the same moment. Same path beforehand? Yes, and it can be checked directly. Sharp in time? Yes, one election. Change came alone? Yes — this is the crucial one, and it is why this case is worth more than every revolution in this part. Nothing else about those villages changed. Measured the same way? Yes; the same survey teams visited both.

Six out of six. I am not aware of another case in this entire subject that manages it.

What it still cannot do: it tells you about villages where a woman held a specific office for one or two terms. It does not tell you about a country, a religion, or a norm.

2.2 — What changed

Three findings came out of this, from a series of studies through the 2000s and 2010s, and they get progressively more interesting.

First: women in charge governed differently. Researchers compared what village councils actually spent money on. In reserved villages, investment shifted towards the things women in those villages complained about most in surveys — drinking water in particular, and in some states roads. Not towards what women in general are assumed to want. Towards what the women there had actually said.

That is a smaller and better finding than it sounds. It does not show that women are better leaders. It shows that leaders respond to the constituency they belong to, which is a claim about representation rather than about women.

Second: exposure changed what villagers believed. This is the one that surprised the researchers. Villagers in reserved villages were tested on their evaluations of leaders, including hypothetical ones. Before exposure, a speech attributed to a woman was rated worse than the identical speech attributed to a man. After two rounds of having a female leader, that gap shrank substantially — and women became more likely to stand for, and win, seats in later elections that were not reserved.

Two things there are worth separating. Bias against women leaders was measurable and real. And it was not fixed: it responded to experience within a few years.

Word Box

Implicit bias: an automatic association that shapes judgement without the person intending it or noticing it. Measured here by giving people identical material attributed to a man or a woman and comparing the ratings.

Why the design matters: nobody is asked what they think about women leaders, which would produce the answer they think is expected. They are asked to rate a speech. The difference between the two ratings is the bias, and the person rating it never knows they are being measured on that.

Third, and this is the finding that belongs to this series: the gap between what parents wanted for their sons and what they wanted for their daughters narrowed. So did the gap between what adolescent boys and adolescent girls wanted for themselves. The gender gap in school attendance in those villages closed, and girls spent measurably less time on household chores.

The reported sizes were substantial: roughly a quarter of the parental aspiration gap and about a third of the adolescents’ own aspiration gap, closed within about seven years, by nothing more than the presence of a woman running the village council.

In Real Terms

Think about what that mechanism actually is. No law changed for those girls. No school was built for them specifically. Nobody gave their families money or lectured them about equality.

A woman was in the chair at the village meeting. Their fathers saw her there. They saw her there.

And within seven years, a third of the distance between what a boy in that village expected from his life and what a girl expected from hers had closed.

That is the cheapest intervention in this entire series, it was produced by an administrative rotation rule, and nobody designed it to do this.

2.3 — What it does not show

Now the honest limits, because this is the evidence I most want to be true and that is exactly when to be careful.

It is about political office. Whether the same mechanism operates on marriage, sexual conduct, dress or mobility — the domains this series is actually about — is not established by any of it. Those norms may be far stickier, and Parts Three and Eight gave reasons to expect they are.

The attitude effects mostly required two rounds. One term of a female leader did much less. Whatever is happening needs repetition, which means short interventions may show nothing and be wrongly judged failures.

Some of the findings are uncomfortable in their own right. Villagers in reserved villages often rated their female leader’s actual performance worse even where objective measures were equal or better. Exposure reduced bias about women in the abstract while the specific woman in front of them was still marked down.

And the effects were measured over roughly a decade. Whether they persist across a generation is not known, because not enough time has passed.

The Argument — does this generalise?

Everybody accepts the findings. The dispute is entirely about what they license.

It generalises, and it is the strongest thing anybody has

The mechanism identified is not about panchayats. It is about what happens when people repeatedly observe a woman doing something they believed women do not do. That mechanism is general by construction, and it is the same one invoked in every argument about representation in every field. Crucially, it settles a question this series has been stuck on since Part One: whether the beliefs underneath these rules are fixed features of a society or responsive to evidence. They are responsive, measurably, within seven years, in rural north India — which is about the least promising place anyone could have picked to test it.

It generalises much less far than people want

Holding a village office is a low-cost belief to revise. It costs a father nothing to accept that a woman can run a council. It costs him a great deal to accept that his daughter may choose her own husband, because that touches property, lineage, caste and his family’s standing — which is exactly the machinery Part Eight described. Norms are not one substance with one stickiness; the ones this series is about are the ones defended hardest. And a study of leadership in a state-created institution says nothing about a domain governed by families rather than by law.

Where things stand: the finding is solid and its reach is genuinely uncertain. What it establishes beyond dispute is that bias is not fixed — that a belief about what women can do responded to observation, in a conservative setting, within a decade. What it does not establish is that beliefs about what women may do respond the same way. The second is the harder case and it has not been tested, because nobody has randomised it and nobody will.

What would settle it: a comparable randomisation in a domain closer to this series — reserved positions, or a similar rotation, in an institution touching marriage or property, with aspiration and attitude measured the same way. India’s own reservation system generates natural variation of this kind that has never been fully exploited.

Why people care so much: because if beliefs are responsive, then the whole argument is about how to change them and the answer is exposure. If they are not, the argument is about coexistence with something that will not move. Almost everybody’s position on the rest of this series depends on which one they assume.

Whichever way that generalises, one thing about this case is settled: it is the only piece of evidence in thirteen parts that came from a lottery. Everything after this chapter is weaker, and the next case is the strongest of what remains.

Remember This

India’s 1993 constitutional amendment reserved a third of village council leaderships for women and rotated them by village serial number — which is unrelated to anything about the village. That is a lottery nobody intended.

It is the only case in this series that passes all six tests, and it passes the hardest one: nothing else about those villages changed.

Three findings. Women in charge spent differently, on what the women there had actually asked for. Bias fell after two rounds of exposure, and more women then won unreserved seats. And the gap between what parents wanted for sons and for daughters narrowed by about a quarter, with the adolescents’ own gap narrowing by about a third.

No law changed for those girls. A woman was in the chair at the village meeting, and their fathers saw her there.

What this establishes beyond dispute is that bias is not fixed. What it does not establish is that beliefs about what women may do move as easily as beliefs about what women can do — and the second is the one this series is about.

3Two Halves of One People

One nation, one language, one ancestry, cut in half and governed on opposite principles for forty years. It is the closest thing political history offers to a laboratory, and it fails one of the six tests in a way that matters.

3.1 — Two answers to the same question

Between 1949 and 1990 the same people, in the same country, were run according to two opposite theories about what women are for. That has never happened anywhere else at that scale with that much record-keeping.

In the East, women worked. Not as an aspiration — as a near-universal fact. By the 1980s the overwhelming majority of working-age East German women were in employment or training. The state supplied childcare from infancy, kept it open through the working day, and treated a mother’s employment as normal rather than as a problem to be managed.

In the West, the arrangement was the opposite and it was written into law. Until a reform in 1977, a West German husband could legally object to his wife taking paid work. The tax system rewarded households with a single earner. And the school day ended around lunchtime, which by itself made full-time work by a mother close to impossible without private money.

Word Box

The male-breadwinner model: a set of arrangements built on the assumption that a household has one earner, who is the husband, and one person doing the unpaid work, who is the wife.

Why it matters that it is a model rather than a preference: it is assembled from separate pieces — tax rules, school hours, childcare availability, pension credits, shop opening times — none of which mentions women. Each looks like an administrative detail. Together they make one arrangement easy and the other one exhausting, and nobody has to forbid anything.

3.2 — What the comparison showed

The differences were large and they were exactly what you would predict. East German women were employed at far higher rates, worked full-time far more often, had children younger, and had them while working rather than instead of working. The pay gap between men and women was narrower in the East.

Then in 1990 the border came down, and the East was absorbed into West German institutions almost overnight — West German law, West German tax rules, West German childcare provision, West German school hours.

What happened next is the most useful thing in the case.

In Real Terms

East German fertility collapsed. The rate fell from roughly 1.5 to 1.6 children per woman before reunification to below 0.8 within about four years.

To put that in the terms of Part Nine: a fall of that size, that fast, in a large population at peace, has essentially no precedent. It is steeper than anything Korea has done, and Korea took thirty years to get where East Germany went in four.

Then it recovered, and by the late 2000s eastern and western German fertility had converged.

Nothing happened to East German women’s bodies in 1990. Something happened to the arrangements around them, and their children stopped being born.

Reading a change of that size correctly requires one technique, and it is the standard tool for every case in this part.

Word Box

Difference-in-differences: the standard method for reading a case like this. You do not compare East and West at one moment. You compare how much each changed, and take the difference between the two changes.

Why it is better than a before-and-after: if something happened to all of Germany in 1990 — and a great deal did — it affects both sides, and subtracting one from the other removes it.

Why it is not magic: it only works if the two sides would have carried on in parallel without the change. That assumption cannot be tested directly, only made plausible by showing the lines were parallel beforehand.

3.3 — What lasted

The other finding is about persistence, and it is the one that speaks to this series.

Decades after reunification — long after the East German state, its childcare system and its ideology had all ceased to exist — women who had grown up in the East were still markedly more likely to work full-time, more likely to use childcare for very young children, and more likely to hold the view that a mother’s employment does not damage her children. Those differences narrowed with time but did not disappear.

So a set of arrangements that lasted forty years produced beliefs that outlived the arrangements by another thirty. That is a real measurement of how sticky this material is, and it cuts in both directions: it means norms can be changed by institutions, and it means the change is not quick to reverse in either direction.

The Argument — what actually caused the fertility collapse?

The collapse is not disputed. Its cause is, and the answer matters because it is the cleanest test anyone has of whether family arrangements are about preferences or about infrastructure.

The economic shock did it

Reunification devastated the eastern economy. Unemployment went from essentially nil to enormous within two years, whole industries closed, and hundreds of thousands of young people moved west. People do not have children during a collapse of that kind, and every other case of sudden economic catastrophe produces a fertility drop. The recovery afterwards is the proof: as the eastern economy stabilised, births returned. Nothing about childcare policy is needed to explain any of it.

The institutions did it

The economic explanation cannot account for the size. Ordinary recessions move fertility by a tenth of a child, not by half. What changed for an East German woman in 1990 was not only her job prospects: it was that the crèche closed, the school day ended at noon, the tax system started penalising her earnings, and having a child stopped being compatible with working — all at once, imposed from outside, with no adjustment period. The postponement was rational because the entire arrangement that had made early childbearing workable was removed in a single year.

Where things stand: both, and demographers have largely converged on a mixed account — a severe economic shock arriving simultaneously with an institutional one, which is precisely why the case fails the fifth test from Chapter One. The change did not come alone. What can be said with confidence is that the collapse was far too large for the economic shock alone, and that the recovery tracked both stabilisation and the gradual rebuilding of childcare in the east.

What would settle it: a case where one arrived without the other — an institutional change of that magnitude without an economic shock. Nothing comparable exists, which is why this remains open thirty-five years later.

Why people care so much: because if arrangements determine fertility to that degree, then Part Nine’s conclusion — that no country has policied its way back to replacement — becomes a statement about what has been tried rather than about what is possible.

3.4 — The other divided nation

Korea is the obvious companion case and I am going to be brief about it, for a reason that is itself informative.

The two Koreas separated in 1945 with a shared language, culture and Confucian family structure, and have been governed on opposite principles ever since. It should be as valuable as Germany. It is not usable, because there is no reliable data on the northern side — no credible census, no independent survey, no verifiable statistics on employment, education or fertility.

So the best available comparison on that peninsula is not between two Koreas. It is between South Korea and its own past, which is a before-and-after in one place, which fails the second test.

That is worth stating rather than skipping, because it illustrates something about this whole field: the availability of evidence is not random either. We know about divided Germany because both halves kept records and both halves let people count. The cases where a society is most tightly controlled are the cases where least is known, and those are exactly the cases this series most needs.

The Hidden Assumption — that the control group stayed still

Everybody who uses this comparison — and every comparison like it — talks as though one side were the treatment and the other the baseline. East Germany did something to women; West Germany is what would have happened otherwise.

West Germany was not a baseline. It was also a treatment, and an unusually strong one. A legal power for a husband to forbid his wife’s employment until 1977 is not the absence of a policy about women. A tax system that penalises a second earner is not neutrality. A school day ending at noon is a decision. The West ran an active programme to keep married women at home, and it ran it through administrative details rather than announcements, which is why it reads as the natural state of things and the Eastern arrangement reads as an intervention.

This matters for reading the numbers. When the two sides converged after 1990, that was widely described as the East becoming normal. Both were moving. Western German women’s employment rose enormously across the same decades, driven by changes with nothing to do with the East.

The general form: mistaking a comparison for a counterfactual. It runs through every A-versus-B argument. Comparing a reformed country to an unreformed one and calling the second the baseline, when the second is running its own vigorous programme. Comparing a treated patient to an untreated one who is also doing something.

And notice which side gets called the intervention. It is always the one that departs from what the observer regards as ordinary — which means the label is doing the work of an argument before any evidence arrives. In this series that matters more than usual, because the whole question is whether the traditional arrangement is a policy or the absence of one.

Which is the right frame for the next case, where a state did announce what it was doing, loudly, and got the opposite of what it announced.

Remember This

Divided Germany is the closest thing to a laboratory this subject has: one people, two opposite theories, forty years, good records on both sides.

East German women worked at near-universal rates with state childcare from infancy. West Germany ran a male-breadwinner model assembled from tax rules, school hours and — until 1977 — a husband’s legal power to forbid his wife’s job.

After reunification, East German fertility fell from about 1.5 to below 0.8 in four years, one of the sharpest peacetime collapses ever recorded, then recovered. The cause was both an economic shock and the removal of every arrangement that made children compatible with work — which is why the case fails the “change came alone” test.

The attitudes outlived the state that produced them by thirty years. Institutions can change norms, and the change is slow to reverse in either direction.

And West Germany was never a control. It ran its own vigorous programme through administrative details rather than announcements — which is exactly why it reads as the natural state of things.

4The Reversal

One country reversed its rules about women almost completely, in a year, in living memory. The traditionalists won the rules and lost the outcomes, and the reason why is the most interesting thing in this part.

4.1 — What was reversed

Before 1979, Iran was on the standard modernising path. A family law passed in the 1960s and strengthened in the 1970s had raised the age of marriage, restricted polygamy, given women grounds for divorce and moved family disputes into courts. Women sat as judges. Unveiling had been enforced, briefly and brutally, decades earlier.

After the revolution, that was undone, quickly and comprehensively. The family law was suspended. The minimum age of marriage was lowered dramatically. Women were removed from the judiciary. Veiling became compulsory and was enforced by law. Public space and education were segregated by sex.

If the rules about women determine the outcomes for women, this is the case where you would expect to see it. A society moving one way was turned round completely, at a known moment, and held there for four decades. So what happened?

4.2 — What happened

Female education exploded. Female literacy in Iran was around a third in the mid-1970s. By the 2000s it was above eighty per cent, and among younger women it approached universal. Girls’ primary enrolment became near-total. And by the turn of the millennium women were a majority of those entering Iranian universities — a threshold most Western countries crossed at around the same time or later.

Fertility collapsed. Iran went from roughly six and a half children per woman in the early 1980s to about two by the year 2000, and is now somewhere around 1.6. Part Nine called this the fastest sustained fertility decline ever recorded for a large country, and it happened under a theocracy that had begun by encouraging large families.

And women’s paid employment did not move. Female labour force participation in Iran has stayed low — in the mid-teens as a percentage — through the whole period. A country produced a generation of highly educated women and did not employ them.

In Real Terms

Take an Iranian family in a village in 1975. The mother cannot read. There are seven children. The daughters will marry at fifteen or sixteen and will not go to secondary school.

Take her granddaughter in 2015. She has one or two siblings. She finished school and probably went to university. She is very likely unemployed, and she wears a headscarf she is legally required to wear.

Every rule about her conduct is stricter than her grandmother’s. Almost every fact about her life is the one the reformers wanted and the traditionalists warned against.

4.3 — Why the rules did not deliver

Three explanations, and they are not rivals — all three appear to have operated.

The regime built the machinery of the thing it feared. The Islamic Republic ran an enormous rural development programme: roads, electricity, clinics, and a national literacy campaign. Part Three of this series identified exactly those as the drivers of fertility decline — children stop dying, schooling spreads, cities grow. The government supplied every one of them for reasons of its own, and got the demographic consequence attached.

It reversed course on family planning, hard. After the war with Iraq ended, the same clerical establishment endorsed one of the most effective family planning programmes ever run — free contraception, religious endorsement, and compulsory pre-marriage counselling. Fertility fell off a cliff. When the leadership reversed again after 2012 and began urging larger families, fertility carried on falling — which is Part Nine’s asymmetry, observed inside a single country under a single government.

And the veil functioned as a permission slip. This is the claim that deserves the most care, because it is the most interesting and the most contested.

Word Box

Enabling constraint: a restriction that makes an activity possible by making it acceptable.

The argument in this case: a conservative rural father in 1975 would not send his daughter to a mixed school taught by unveiled strangers in a state he regarded as irreligious. After 1979 the school was segregated, the teachers were veiled, the curriculum was religiously framed, and the state was one he trusted. The same act — sending a daughter out of the house to be educated — went from unthinkable to permissible without his beliefs changing at all.

Why it matters: the constraint and the opportunity arrived in the same package. On this account the compulsory veil is not something that happened despite the education boom. It is part of what produced it.

Before weighing that claim, it is worth scoring the case itself, because Iran is quoted far more often than its evidential quality warrants.

How We Actually Know This

Score Iran on the six tests from Chapter One, because it is the case most often quoted and it does not do well.

Assignment unrelated to outcome? No. A revolution is not assigned; it happens for reasons connected to everything else about a society. Comparison group? None. Neighbouring countries differ in too many ways to serve. Same path beforehand? This it does have — the pre-1979 trends are documented, so you can see whether the line bent. Sharp in time? Yes, unusually so. Change came alone? Emphatically not: a revolution, a war with a million casualties, sanctions, oil price swings and a development programme all at once. Measured consistently? Broadly yes; Iranian censuses and surveys are reasonably good.

So Iran passes three tests and fails three, including the most important one. It is a strong before-and-after and a weak experiment.

What that means practically: Iran cannot tell you what the rules caused. It can tell you what the rules failed to prevent, and that is a genuine finding — a negative one, and negative findings from a single case are the most reliable kind, because you only need one counterexample to refute a universal claim.

4.4 — What it refutes

That last point is worth drawing out, because it is how a weak case can still do real work.

A claim of the form “you cannot have strict rules about women’s dress and conduct together with mass female education” is refuted by Iran on its own. One counterexample is enough. Similarly, “high fertility requires only that a state wants it” is refuted: Iran’s state wanted it and did not get it.

What Iran cannot do is establish that the rules caused the education boom, because too much else was happening. The enabling-constraint account is plausible and it is not proved.

The Argument — did the rules cause the education boom, or merely fail to prevent it?

This is one of the genuinely open questions in this part, and the two answers imply opposite things about every conservative society on Earth.

They caused it, and the mechanism is real

The timing is not a coincidence. Rural girls’ enrolment did not creep up after 1979; it jumped, and it jumped fastest in the most conservative provinces — precisely where the pre-revolutionary state had been least able to reach. Segregated schools, veiled teachers and religious framing removed the objection that had kept those girls at home, and the state that made the offer was one their fathers trusted. The same logic explains parallel cases elsewhere: participation by conservative women in public life often rises when the setting is made religiously acceptable rather than when it is made secular. On this reading, insisting on unveiling in 1936 kept girls out of school, and requiring the veil in 1983 put them in it.

They failed to prevent it, and the credit belongs elsewhere

Female education was already rising before 1979 and was rising across the whole region in the same decades, including in countries with no revolution. What Iran did was pour oil money into rural roads, clinics, electricity and a literacy campaign, which is what every country that expanded schooling has done. Attributing the result to the veil rather than to the infrastructure is reading a slogan off a machine. And note what the same regime did to the women it produced: barred from the judiciary, largely excluded from paid work, and policed in the street. A permission slip that admits you to a university and then locks the exit is not obviously a permission slip.

Where things stand: unresolved and probably partly both. The regional trend is real, so some of the rise would have happened anyway. The provincial pattern — the sharpest gains in the most conservative places — is the strongest evidence for the enabling-constraint account and it is not easily explained away. What is agreed is the outcome: whatever caused it, the rules did not stop it.

What would settle it: a comparison of Iranian provinces by how conservative they were before 1979 against how much girls’ enrolment jumped after, with the infrastructure spending held constant. The census data to attempt this exists.

Why people care so much: because if the enabling-constraint mechanism is real, then a great deal of what looks like oppression is also, simultaneously, the thing making participation possible — and neither side of the wider argument has anywhere comfortable to put that.

Iran is a country that reversed its rules and did not get the outcomes. The next case is a country that removed them altogether, and it is running now.

Remember This

Iran reversed its rules about women almost completely in 1979: family law suspended, marriage age lowered, women removed from the judiciary, veiling made compulsory, public space segregated.

Then female literacy went from about a third to over eighty per cent, women became a majority of university entrants, and fertility fell from about six and a half children to around 1.6 — the fastest sustained decline ever recorded for a large country. Paid employment did not move.

Three things did it: the regime built the development machinery that drives fertility decline; it ran one of the most effective family planning programmes in history and then could not reverse it; and the veil may have worked as a permission slip that let conservative fathers send daughters out of the house.

On the six tests Iran passes three and fails three. It is a strong before-and-after and a weak experiment.

But a single case can refute a universal claim, and this one refutes two: that strict rules and mass female education are incompatible, and that a state which wants high fertility can have it.

5The Removal

The most complete exclusion of women from public life in modern times is running now. Almost everything published about its consequences is a forecast rather than a finding — and one strand of it is neither, because it is mechanical.

5.1 — What was removed

Since September 2021 no girl in Afghanistan has attended school beyond the sixth grade. In December 2022 universities were closed to women. Employment with aid organisations and the United Nations was barred. In late 2024 the restriction reached training in midwifery and nursing, which had been one of the last routes remaining.

By the most recent school year, roughly 2.2 million girls were out of secondary education, with several hundred thousand more added each year as new cohorts reach the cut-off. On present policy the figure passes two million deprived of any education beyond primary by 2030.

This is not a partial restriction of the kind this series has been describing elsewhere. It is a near-total removal, imposed at a known date, on a national population, with a documented twenty-year period of the opposite policy immediately beforehand.

5.2 — Measured, modelled, and the difference

Here is where I have to slow down, because the reporting on this does something that matters.

Word Box

Measurement: somebody counted the thing. Births attended, girls enrolled, deaths recorded.

Projection: somebody built a model of how the thing usually responds and calculated what should follow. No counting of the outcome has taken place.

Why the distinction is critical here: projections in this case are made by serious institutions using well-established relationships, and they are the best available guide to what is coming. They are still not evidence that it has happened. A sentence like “maternal deaths rose by half” and a sentence like “maternal deaths are projected to rise by half” describe completely different epistemic situations, and the second routinely becomes the first in the retelling.

So it is worth going through the Afghan figures one at a time and putting each in the right column.

How We Actually Know This

Sorting the Afghan figures into the two categories.

Measured. The bans themselves and their dates. Enrolment: roughly 2.2 million girls out of secondary school. Child marriage from a household survey covering 2022–23: about 28.7 per cent of girls married before eighteen and about 9.6 per cent before fifteen. Female illiteracy around 84 per cent. Fertility around 5.3 to 5.4 children per woman with modern contraceptive use under a quarter. Maternal mortality around 638 deaths per 100,000 live births, already among the highest in the world.

Modelled. The widely quoted consequences: a 25 per cent rise in child marriage, a 45 per cent rise in adolescent childbearing and at least a 50 per cent rise in maternal mortality by 2026; an additional 1,600 maternal deaths and over 3,500 infant deaths; an estimated 0.5 per cent reduction in national output already, from an assessment published in April 2026.

What the modelled figures are good at: they use relationships that are extremely well established across many countries — mothers’ education and child survival is one of the most replicated findings in development economics.

What they cannot do: confirm that it happened. Afghanistan’s statistical capacity has degraded severely since 2021, which means the country least able to be counted is the one generating the most quoted numbers about itself.

5.3 — The one chain that is not a model

There is a strand of this that does not depend on statistics at all, and it is the strongest thing in the chapter.

Afghanistan enforces strict segregation in medical care: a woman is, in practice, treated by a female clinician or by nobody. That is not a projection; it is the operating rule.

Now add the December 2024 restriction on women training in midwifery and nursing. The existing female health workers age, retire, leave or die, and no replacements are being produced. Within a decade there will be a shortage of the only people permitted to attend Afghan women in childbirth.

This does not require a regression. It is arithmetic on a closed system. If only women may treat women, and no new women may be trained, the number of people able to treat women falls to zero on a schedule that can be calculated.

In Real Terms

Afghanistan’s maternal mortality is roughly 638 deaths per 100,000 live births. That means for every thousand Afghan women who give birth, between six and seven die — before any of this began.

Put the chain end to end. A rule about girls attending school past the sixth grade becomes, twelve years later, a rule about who is qualified to deliver a baby, which becomes a rule about which women survive childbirth.

Nobody who wrote the education rule was making a decision about obstetrics. That is what a rule about women’s schooling turns into in a society that also forbids a male doctor from touching a female patient — and the two rules were made by the same people, for the same reasons, without anyone connecting them.

5.4 — What the case can and cannot establish

The Argument — what does Afghanistan actually prove?

It is the most emotionally decisive case in this part and it is not the most evidentially decisive, and those two facts are worth separating carefully.

It is the definitive demonstration

Here is total exclusion of half a population, applied at a known date, in a country with two decades of contrary policy immediately before it for comparison. The outcomes are catastrophic and running in exactly the predicted direction across every domain measured — education, marriage age, health, output. If this does not count as evidence that the rules matter, nothing ever will, and the demand for a cleaner case is a way of refusing to look at the clearest one available.

It fails the isolation test badly

Everything happened at once. The government changed by military collapse; international aid, which had funded a large share of the state, stopped; foreign reserves were frozen; the banking system seized; a drought was under way. Afghanistan’s economy would have contracted severely with no change to women’s rules at all. So the aggregate outcomes cannot be attributed to the exclusion of women, and citing them that way is exactly the reasoning Chapter One warned about. It demonstrates what happens when a country collapses.

Where things stand: both, and the resolution is to split the claim. The aggregate figures — output, poverty, humanitarian need — are hopelessly confounded and should not be used to prove anything about women’s rules. The specific mechanical chain in the section above is not confounded at all, because it does not run through the economy: it runs through a rule about who may treat whom and a rule about who may be trained. Aid could be restored tomorrow and that chain would be unaffected.

What would settle it: restoration of statistical capacity and a proper survey round, which would convert the projections into measurements. Whether that happens depends on decisions nobody reading this controls.

Why people care so much: because Afghanistan is deployed as a closing argument by everyone who wants the rules-matter conclusion, and treating a confounded case as decisive is how a strong argument acquires a weak foundation. The mechanical chain is a better argument than the aggregate one, and almost nobody uses it.

5.5 — The repeat

One structural feature of this case is rare enough to note. The same policy was applied to the same country twice — from 1996 to 2001, and again from 2021 — with two decades of the opposite policy in between.

That is close to a reversal design: treatment, removal, treatment. It is uncommon in social science and it lets you ask whether the second application produced the same effects as the first.

The honest answer is that we do not yet know, because the first period was even less well documented than the current one, and because the country in 2021 was not the country of 1996 — twenty years of schooling had produced a generation of literate women who are now barred from teaching. That difference cuts an unexpected way: the second application began from a much higher base, so the same rules remove much more.

Remember This

Since September 2021 no Afghan girl has attended school past the sixth grade; universities closed to women in December 2022, and midwifery and nursing training in late 2024. Roughly 2.2 million girls are out of secondary school.

Separate the measured from the modelled. Enrolment, child marriage rates, literacy, fertility and maternal mortality are counted. The headline consequences — a quarter more child marriage, half again as many maternal deaths, half a per cent of national output — are projections, and the retelling constantly converts them into findings.

One chain is neither a model nor a projection. If only women may treat women, and no women may be trained, the number of people able to attend an Afghan birth falls to zero on a calculable schedule. That is arithmetic on a closed system.

The case fails the isolation test badly — a military collapse, an aid cutoff, frozen reserves and a drought arrived together — so the aggregate figures prove nothing about women’s rules specifically.

Which means the mechanical chain is a stronger argument than the catastrophe, and almost nobody uses it.

6The Lifting

Two cases where something that had been forbidden or impossible became available, on a known date, in a society that did not otherwise transform. They are the closest thing this argument has to a test of whether the old position was a preference or a wall.

6.1 — Saudi Arabia, 2018 and 2019

In June 2018 Saudi Arabia lifted its ban on women driving. In August 2019 it changed the guardianship rules so that women over twenty-one could obtain a passport and travel without a male relative’s permission, and could register births, marriages and divorces in their own right.

These are sharp, dated, specific changes to constraints that had held for decades. On the Chapter One checklist that is unusually good. What followed is one of the largest measured changes in women’s economic position anywhere, ever.

Saudi female labour force participation stood at roughly 17 per cent in 2016. By 2025 the national statistics authority put it at around 36 per cent. That is a rise of about nineteen percentage points in under a decade, and it represents on the order of 1.3 million additional Saudi women in paid work.

In Real Terms

Nineteen points in nine years. For comparison, the equivalent shift in the United States — from roughly a third of women in the workforce to around half — took about thirty years, from the 1950s to the 1980s, and was regarded at the time as a social transformation.

Saudi Arabia did a comparable distance in less than a third of the time, starting from a much lower base, in a society that had not otherwise been reorganised.

Nothing changed about Saudi women between 2016 and 2025. Something changed about what they were allowed to do, and about a third of a million a year began doing it.

A number that large, moving that fast, deserves scrutiny of where it comes from before anyone builds on it.

How We Actually Know This

The figures come from Saudi Arabia’s own statistics authority, which has run quarterly labour force surveys since 2016 and supplements them with administrative employment records. That is good coverage and consistent method — but three cautions belong on the record.

First, the government set a public target of 30 per cent as part of its national programme and then exceeded it. Any statistic that is also a performance indicator for the people producing it deserves an extra look, and this one is both.

Second, the definitions and survey methods changed during the period, which is normal and which makes the earliest and latest figures not perfectly comparable.

Third, the employment is geographically concentrated in the largest cities. The national figure hides a country in which the change has been far smaller in smaller towns and rural areas.

None of this touches the direction or the rough magnitude. Independent international estimates show the same shape. It does mean the precise number should be read as approximate.

The honest complication is that the lifting did not arrive alone. It came inside a large economic programme that also included employment quotas for citizens, transport and childcare subsidies aimed at working women, and a background of falling oil revenue that made a second household income more necessary. So the case fails the fifth test: what was removed was a legal constraint, and what was added at the same time was a set of inducements.

6.2 — Bangladesh, and a better-identified case

The second case is less famous and evidentially stronger.

From the 1980s, Bangladesh built an export garment industry which by the 2010s employed several million people, the large majority of them women. Factories opened in some places and not others, and the pattern of where they opened had more to do with roads, ports and land than with how progressive the surrounding villages were.

That gives researchers something close to a comparison group: villages near a new factory against otherwise similar villages further away. When that comparison was done properly, the results were consistent and large. Girls growing up near garment factories married later, had their first child later, and stayed in school substantially longer — and the schooling effect was larger than that produced by a well-known government scheme that paid families directly to keep girls in school.

The mechanism is worth spelling out because it is not the one people assume. The factories did not employ schoolgirls. They employed young women, and the jobs required basic literacy and numeracy. So the return on educating a daughter became visible, local and specific: families could see what school was for, in the form of a wage a neighbour’s daughter was bringing home.

Word Box

Revealed preference: the idea that what people actually do tells you more about what they want than what they say.

Why it matters in these two cases: for decades, low female employment in both countries was explained as a preference — women, or their families, did not want it. Revealed preference is the standard tool for testing that, and here the test was run by events rather than by researchers.

The catch, which matters just as much: a preference expressed under constraint is not a free preference either. What people do when an option is forbidden tells you almost nothing, and what they do when it is newly permitted and newly rewarded tells you something mixed.

And the cost belongs in the record, because a chapter that reported only the gains would be doing exactly what Chapter One warns about. The Bangladeshi garment industry has killed workers in large numbers, most notoriously in the collapse of a factory building in 2013 in which more than eleven hundred people died. The same industry that delayed marriages and filled schools also produced some of the worst industrial safety failures of the century, and the people who bore both were the same women.

The Argument — does lifting a constraint reveal what women wanted?

This is the question these two cases were supposed to answer, and they answer part of it.

Yes, and the numbers settle a long argument

For decades the low participation of women in Saudi Arabia and rural Bangladesh was explained as cultural preference — this is simply not what women there want, and outsiders should stop projecting. Then the constraint moved and behaviour moved with it, immediately and at enormous scale. Preferences held by whole populations do not shift nineteen points in nine years. Walls come down in an afternoon. Whatever the previous level was measuring, it was not what women wanted, and the people who insisted it was should say so.

It shows less than that, in both directions

Neither case is a clean lift. Saudi women got permission alongside quotas, subsidies and household economic pressure; Bangladeshi families got a factory wage, which is an inducement rather than a permission. So what these cases measure is the response to a package, not the revelation of a latent want. And there is a harder version of the objection: a woman raised inside these rules does not have a preference sitting underneath them waiting to be uncovered. Her preferences were formed by the same arrangements. There is no free-standing answer to what she “really” wants, so no lifting can reveal it.

Where things stand: the cases establish a negative claim decisively and a positive claim not at all. Decisively: the previous level was not a ceiling set by preference. Something that responds that fast to a change in permissions and incentives was being held down, and the cultural-preference explanation for the old level is dead. What they cannot establish is what the level would be under no constraint, because that condition has never existed anywhere and the second side is right that it may not be a coherent idea.

What would settle it: a case where a permission was lifted with no accompanying inducement and no economic pressure. Rare, because governments that lift restrictions usually want the resulting participation and pay for it.

Why people care so much: because “this is what they want” has been the single most durable defence of every arrangement in this series, and these two cases are the strongest evidence anyone has ever assembled against it.

That is the record of reforms that worked. The next chapter is the record of reforms that did the opposite of what they were for, and it is the more useful of the two.

Remember This

Saudi Arabia lifted its driving ban in June 2018 and relaxed guardianship rules in August 2019. Female labour force participation went from about 17 per cent in 2016 to around 36 per cent by 2025 — roughly nineteen points in nine years, a shift that took the United States about thirty.

The lifting did not come alone: quotas, transport and childcare subsidies and falling oil revenue arrived with it. The case fails the isolation test.

Bangladesh is the better-identified case. Garment factories opened where the roads and ports were, not where attitudes were. Girls near them married later, had children later and stayed in school longer — by more than a government scheme that paid families directly. Because a factory wage made visible what school was for.

The cost belongs on the record too. The same industry produced the factory collapse of 2013 that killed more than eleven hundred people, and the women who gained were the women who died.

Together they kill one claim decisively: whatever the old level of women’s work was measuring, it was not preference. Something that moves nineteen points in nine years was being held down.

7When the Reform Backfired

The most useful chapter here for anyone who actually wants to change something. Four Indian reforms that produced the opposite of what they intended, and one clean rule that explains all four.

7.1 — The law that caused what it banned

In 1929 the colonial government passed an act raising the minimum age of marriage for girls. It was passed in the autumn and took effect the following spring, and in the months between, families across India rushed to marry their daughters before the law could bite.

How We Actually Know This

This one is unusually well documented, because a census fell immediately afterwards. The 1931 census recorded a marked surge in the marriage of very young girls in the window between the act’s passage and its commencement — a spike visible in the age-at-marriage distributions and remarked on at the time.

Why this evidence is strong: the spike is confined to the gap between passage and commencement, which is exactly the window the mechanism predicts and no other explanation produces. It is a sharp change with a known date and a documented before-and-after in the same population.

What it cannot tell you: the long-run effect of the act itself, which is a different and much harder question involving enforcement, registration and everything that happened over the following ninety years.

The announcement of a prohibition is itself an event. It tells everybody that a window is closing, and the rational response to a closing window is to move before it shuts. Reform is not a switch that flips a society from one state to another; it is a signal that people act on, including in the interval before it applies.

7.2 — Dowry: prohibited, and then it grew

India prohibited dowry in 1961. In the decades that followed the practice did not shrink. It spread — into communities that had not previously practised it, including some where the payment had historically run the other way — and the amounts rose.

The mechanism is not mysterious once you look at what prohibition actually did to the transaction.

Before, a dowry was negotiated openly between families, was often at least partly in goods, and had a degree of social visibility that acted as a constraint — everyone knew what had been asked, and asking too much carried a reputational cost. Prohibition did not remove the demand. It moved the transaction out of the visible, negotiable, semi-documented space and into undocumented cash.

What was lost in that move was not the payment. It was every check on the payment: the bride’s family’s ability to complain publicly, the community’s knowledge of what was normal, and any legal claim on what had been handed over. A woman whose in-laws demanded more after the wedding now had no record of what had already been given.

Word Box

Substitution effect: when blocking one route does not remove the underlying behaviour but redirects it into another, usually worse, channel.

Perverse incentive: when a rule creates a reason to do more of the thing it was meant to stop.

Why both matter here: neither is an argument against regulating anything. They are an argument for asking one specific question before you legislate — where does this behaviour go if I block this route? Every failure in this chapter is a case where nobody asked it.

7.3 — The machine that was meant to save mothers

The third case is the most serious, because the technology involved was introduced entirely for good reasons and its misuse has cost more lives than anything else in this series.

Ultrasound scanning arrived in India as maternal health equipment. It made pregnancies safer. It also made the sex of a foetus knowable, cheaply, from mid-pregnancy, in a society with a strong preference for sons.

India’s child sex ratio worsened steadily across the decades that followed. In 1981 there were around 962 girls per 1,000 boys aged nought to six. By 2001 it was about 927, and by 2011 about 918. Millions of girls who should exist do not.

In Real Terms

Now the detail that ought to be famous and is not. Sex selection in India was not concentrated among the poorest and least educated. It was strongest among families that were richer, more urban and better educated.

The reason is arithmetic. A family that wants six children and prefers sons will probably get sons without doing anything. A family that wants two children and prefers sons has to make it happen. Add a cheap scan and a clinic, and it happens.

Falling fertility and son preference combine to kill girls, and it is the modernising families who do it. Every assumption that education and prosperity would dissolve this preference was wrong in the specific direction that mattered most.

India legislated against it in 1994 and strengthened the law in 2003, prohibiting the use of diagnostic techniques to determine sex. Enforcement has been patchy and prosecutions rare, but the ratio has improved somewhat in recent survey rounds. The point for this chapter is not that the law failed. It is that the harm was created by a maternal health technology, and that the people best placed to use it were the ones the reformers expected to be furthest along.

7.4 — Protection that excluded

The fourth case is the one that recurs in every country. Laws written to protect women workers — bans on night work, exclusions from certain industries, restrictions on hours — were passed by people who genuinely intended protection, in response to real hazards.

The effect was to make women more expensive and less flexible to employ, and therefore to exclude them from exactly the shift-based manufacturing work that has been the main route out of agricultural poverty for women everywhere, including in the Bangladesh case in Chapter Six. In India, restrictions on women’s night work in factories stood for decades and were only unwound state by state from the 2000s onwards.

What makes this case instructive is that the hazard was real. Night work in an unsafe factory in an unsafe city genuinely was dangerous for women. The reform addressed the danger by removing the women rather than the danger, which is the same structure as every restriction in this series — and this time it was written by reformers.

The Argument — do these cases argue against reform, or against bad reform?

This chapter is quoted by two very different sorts of people and they draw opposite conclusions from it.

They argue for caution about reform in general

Four serious attempts by intelligent people to improve women’s position produced more child marriage, more dowry, fewer girls alive and fewer women employed. That is not a run of bad luck; it is what happens when you intervene in a system whose workings you do not understand. Social arrangements that have persisted for a long time usually persist because they are solving something, and removing a piece without knowing what it was doing produces exactly this. Humility is not conservatism — it is the appropriate response to a record like this one.

They argue for better reform, and the pattern is diagnosable

Look at what the four failures have in common: every one of them attacked a symptom while leaving the underlying demand untouched. Dowry was prohibited while the marriage market that generated it was left alone. Sex determination was banned while son preference was left alone. Night work was banned while the danger was left alone. Child marriage was outlawed while nothing changed about why families marry daughters young. That is not an argument against intervening. It is a specific, learnable diagnosis — and the reforms in this series that did work, like the reserved seats in Chapter Two, worked because they changed the underlying structure rather than forbidding an output of it.

Where things stand: the second reading fits the cases better, and it is not a comfortable finding for either camp. The failures share a signature: prohibit the visible symptom, leave the demand in place, and watch it re-emerge somewhere less visible and less regulated. That predicts failure in advance rather than diagnosing it afterwards, which is what makes it a usable rule rather than hindsight.

What would settle it: a systematic accounting of reforms in this area by whether they targeted the symptom or the structure, with outcomes measured the same way. Nobody has assembled one, and this chapter is a sample of four that I chose, which is precisely the problem Chapter One described.

Why people care so much: because “reform backfires” is one of the strongest arguments the traditionalist case has, and it is strong. What it does not support is the conclusion usually drawn from it, which is that nothing should be attempted.

Both sides of that exchange share a way of describing what a law is, and it is worth pulling out, because it explains why this particular argument never gets anywhere.

The Hidden Assumption — that the reform is the intervention

Everybody in that argument evaluates these laws by what they say. Dowry prohibition is assessed as a prohibition on dowry. The sex determination act is assessed as a ban on sex determination.

But a law is not what is delivered. What a population actually receives is the text plus whether anyone enforces it, plus how it is locally understood, plus what happens to the behaviour it displaces. Those four things together are the intervention. The text is only the first.

Run it on dowry. The text said: this payment is illegal. What was delivered was: this payment is now undocumented, unnegotiable in public, and unenforceable if your in-laws want more later. Nobody wrote that. It is what the text became after passing through an unenforcing state and a marriage market that still needed the money to move.

This is why arguments about these reforms are so unproductive. One side says the law was a good idea and the other says it failed, and they are discussing different objects — a text and a delivery. Both can be right at once, permanently, without either learning anything.

The general form: evaluating the instruction rather than what was received. It is everywhere. A curriculum reform judged by the syllabus rather than by what was taught. A safety rule judged by the regulation rather than by what the inspector actually does. A target judged by the number rather than by what people did to hit it.

And there is a sharper version specific to this series, which Part Eleven found from the other direction. A rule about women lands where it can be enforced. So the delivered version of any reform is shaped by exactly the same thing that shapes the delivered version of any restriction: not what was intended, but who has the power to make it real in the room where it applies. In an Indian household, that is very rarely the legislature.

So much for the failures. The next chapter does the counting that a series like this owes its readers: which side actually forecast events correctly.

Remember This

The 1929 marriage age act was passed in autumn to take effect in spring, and the census recorded a surge of child marriages in the gap. An announced prohibition is itself an event, and a closing window makes people move.

Dowry was prohibited in 1961 and then spread and grew. Prohibition did not remove the demand; it moved the payment out of visible, negotiable goods into undocumented cash — removing every check on it and leaving the bride with no record of what had been given.

Ultrasound arrived as maternal health equipment and produced millions of missing girls. India’s child sex ratio fell from about 962 girls per 1,000 boys in 1981 to about 918 in 2011 — and sex selection was strongest among richer, more educated, urban families, because small families plus son preference require action.

Protective labour laws addressed a real hazard by removing the women rather than the danger — the same structure as every restriction in this series, written this time by reformers.

All four share one signature: prohibit the visible symptom, leave the demand in place, and watch it re-emerge somewhere less visible and less regulated. That predicts failure in advance, which makes it a usable rule rather than hindsight.

8The Predictions That Came True

Thirteen parts of testing produce a scorecard. Both sides made confident forecasts and both got badly wrong the thing they were most certain about. This chapter counts it, including the entries I would rather not have.

8.1 — What the traditionalist got right

Four predictions have been vindicated, and I want to state them without hedging, because a series that only found the other side wrong would not be worth the paper.

Fertility would collapse and would not come back. Part Nine established this comprehensively. Every rich country has fallen below replacement, no country has ever policied its way back, and the money spent trying runs to hundreds of billions of dollars. This prediction was correct, and it was correct more strongly than most people on the other side expected.

Marriage would decline sharply. Later, less common, and in East Asia collapsing outright. Correct.

The care of the old would become an unsolved public problem. Part Nine’s Chapter Seven: ten working-age people per pensioner becoming four, and a bill nobody has costed. Correct.

Family stability would fall. Divorce rose steeply across the rich world through the second half of the twentieth century, and single-parent households multiplied. The honest footnote is that divorce rates have fallen in several countries since their peaks around 1980 to 2000 — but the prediction was about direction over the period, and over the period it was right.

8.2 — What the traditionalist got half right

“She will be less happy.” This deserves a careful look because it is the traditionalist case’s best empirical card and it is usually either overstated or dismissed.

How We Actually Know This

Across the United States and much of the developed world, surveys of subjective wellbeing since the 1970s show women’s reported happiness declining relative to men’s. Women started the period reporting higher wellbeing than men and ended it reporting the same or lower. This is a real, replicated finding across multiple national datasets.

What it establishes: the relative position moved, over exactly the decades when women’s legal, educational and economic position improved enormously. That is a genuine puzzle and pretending otherwise is not honest.

What it does not establish, and this is where most retellings go wrong. It is a relative measure — much of the movement is men’s reported wellbeing rising. It is a survey of stated feeling, and what people are willing to state changes: a woman in 1972 asked whether she was satisfied with her life was answering inside a norm that made complaint costly. And a wider range of options can lower satisfaction while raising almost everything else, because satisfaction is measured against what you think you could have had.

So: the finding is real, the interpretation is genuinely open, and both sides quote it as though it were settled.

There is a specific reason happiness comparisons across decades are so hard to read, and it has a name.

Word Box

Reference-group effect: people rate their lives against what they believe is available to them, not against an absolute scale.

Why it complicates every happiness comparison: if a woman in 1960 compared herself only to other housewives, and a woman in 2010 compares herself to everyone including her male colleagues, the second can have a far better life and report a worse score. The ruler changed at the same time as the thing being measured.

This does not dismiss the finding. It means a happiness survey cannot distinguish between “life got worse” and “expectations rose faster than life improved”, and those are very different worlds.

“She will regret it.” Part Ten found a real, robust, moderate asymmetry in sexual regret that survived testing in the most egalitarian society available. Half right — because what predicted an individual woman’s regret was not her sex but whether she was worried, pressured or disappointed, all of which are conditions rather than dispositions.

8.3 — What the traditionalist got wrong

“Women do not want to work.” Killed by Chapter Six of this part. Nineteen percentage points in nine years.

“Educating women will destroy the family and then the society.” No society that educated its women has collapsed. Every one of them got richer, healthier and longer-lived, and their children stopped dying.

“Bias against women in authority is natural and fixed.” Refuted by Chapter Two, in rural north India, within seven years, by rotation.

“Strict rules and mass female education are incompatible.” Refuted by Iran — and note that this claim was made by both sides, which is why Chapter Four is uncomfortable for everybody.

8.4 — And the reformer’s failures

The other column, and it is not shorter.

“Development will dissolve son preference.” Wrong, badly. Chapter Seven: sex selection was strongest among the richer, more educated, more urban families.

“Prosperity and education will end dowry.” Wrong. It grew.

“Equality at home will restore fertility.” Wrong. Part Nine: the Nordic countries went furthest and fell anyway.

“Legal reform changes behaviour.” Often wrong, and Chapter Seven shows the pattern — prohibit the symptom, leave the demand, watch it re-emerge somewhere worse.

The predictionWhoseVerdict
Fertility collapses and does not returnTraditionalistRight, decisively
Marriage declines sharplyTraditionalistRight
Old-age care becomes an unsolved problemTraditionalistRight
Family stability fallsTraditionalistRight over the period; partly reversed since
Women end up less happyTraditionalistHalf right; the measure is relative and contested
She will regret itTraditionalistHalf right; predicted by conditions, not by sex
Women do not want to workTraditionalistWrong
Educating women destroys societiesTraditionalistWrong
Bias against women leaders is fixedTraditionalistWrong
Development dissolves son preferenceReformerWrong, and lethally
Prosperity ends dowryReformerWrong
Equality at home restores fertilityReformerWrong
Legal reform changes behaviourReformerOften wrong
The Argument — whose forecast did better?

Read that table and the obvious question is who won. Two serious answers.

The traditionalists forecast better

Count the entries. Their four unambiguous hits are large, civilisational and irreversible — a fertility collapse with no exit, a marriage system dissolving, an ageing bill nobody can pay. The reformers’ failures are on the exact points where they were most confident and most dismissive: they were certain that education and prosperity would dissolve son preference and dowry, and millions of girls are dead because they were wrong. A forecast is judged on the big calls, and the big calls went one way.

The traditionalists forecast the costs and missed the substance

Every traditionalist hit is a prediction about aggregate social arrangements — birth rates, marriage rates, care burdens. Every traditionalist miss is a prediction about women themselves: what they want, what they can do, whether others’ bias about them is fixed. On the question of what a woman is, they were wrong every time it was tested. And the reformers’ misses are misses of optimism about speed, not of direction — they were wrong that prosperity alone would fix son preference, and right that the preference could be reduced, which several places have since done.

Where things stand: the second reading survives contact with the table better, and it is a real pattern rather than a rescue. The traditionalist case has been consistently right about consequences at the level of societies and consistently wrong about capacities at the level of persons. The reformer case has been right about persons and repeatedly naive about systems. Neither of those is a small thing to be wrong about.

What would settle it: nothing, because “whose forecast did better” is not one question. It depends entirely on which predictions go on the list — and that is the subject of the box below.

Why people care so much: because both sides have spent decades claiming vindication by events, and this is the first table in this series that puts the entries side by side. Neither side comes out of it looking the way its advocates say it does.

Before leaving that table, there is a question about it that nobody in the argument above asked.

The Hidden Assumption — that the scoreboard is neutral

Look again at that table and ask a question nobody asks: who chose the rows?

I did. And notice what determines the result. Fertility, marriage rates, family stability, old-age care — put those four on a scoreboard and the traditionalist wins. Literacy, maternal mortality, income, life expectancy, political representation, violence — put those on instead and it is not close in the other direction.

Both lists are real. Neither side is cheating. Each is measuring the things it thinks a society is for, and the disagreement about what to count is not a disagreement about evidence at all — it is the actual disagreement, arriving disguised as a methodology choice.

This is why the argument never ends. Two people can agree on every number in this series and still disagree completely, because one of them is counting continuity between generations and the other is counting how a life goes for the person living it. No amount of further research adjudicates that, and further research is what both sides keep demanding.

The general form: the winner is decided when the scoreboard is chosen, not when the evidence is gathered. It runs through every contested field. A school judged on exam results or on what its pupils become. An economy judged on output or on how the median person lives. A hospital judged on mortality or on suffering. In every case the metric selection looks technical and does all the moral work.

And it is worth naming what has never happened. In a hundred years of this argument, nobody has ever proposed a shared scoreboard — a list both sides would agree in advance to be judged by. Not once. That is not an oversight. Agreeing the metrics in advance is the one move that would risk losing.

With that in mind, the last thing this part owes you is a single accounting: what all nine cases, graded honestly, actually establish.

Remember This

The traditionalist got four things right: fertility collapses and does not return, marriage declines sharply, old-age care becomes unpayable, family stability falls. These are large and they are not in dispute.

Two half right: women’s relative reported happiness did fall, though the measure is relative and the ruler changed; and the regret asymmetry is real but predicted by conditions rather than by sex.

Several wrong: that women do not want to work, that educating them destroys societies, that bias against them is fixed, that strict rules and female education cannot coexist.

The reformer’s failures are not shorter, and one was lethal: development did not dissolve son preference, prosperity did not end dowry, equality at home did not restore fertility, and legal reform frequently did not change behaviour.

The pattern: the traditionalist has been right about consequences at the level of societies and wrong every time about capacities at the level of persons. And whichever side you think won depends entirely on which rows go on the scoreboard — which nobody has ever agreed in advance, because agreeing it is the one move that risks losing.

9What All of It Establishes

Nine cases, graded by how much weight each can bear. Six findings survive the grading. And one problem with the whole exercise that I can make visible but cannot fix.

9.1 — The cases, graded

Chapter One set out four tiers of evidence and six tests. Here is every case in this part scored against them, so you can see what each is actually worth before anything is drawn from it.

CaseTierWhat it can support
India: reserved council seatsRandomisedBias about women’s capacity is not fixed. Representation changes spending and changes what girls expect of their lives.
Bangladesh: garment factoriesQuasi-randomVisible returns to schooling change family behaviour faster than paying families does.
Divided GermanyBorder, both sides movingInstitutions set employment and fertility levels. Norms outlive the institutions that made them by about a generation.
Saudi Arabia since 2018Sharp, bundledThe previous level of women’s work was not a ceiling set by preference.
Iran since 1979Sharp, heavily bundledRefutes two universal claims. Cannot establish what the rules caused.
Afghanistan since 2021Confounded; mostly modelledThe mechanical chain about who may treat whom. Not the aggregate figures.
Indian reform backfiresCases, chosenProhibiting a symptom while leaving the demand relocates the behaviour somewhere worse.

9.2 — What survives

Six findings hold across more than one tier, which is the minimum I am willing to build on.

One. Beliefs about what women can do respond to observation. The strongest single result in this series, from the only randomisation in it, in rural north India, within seven years. Not beliefs about what women may do — that is untested, and it is the harder case.

Two. The observed level of women’s work is set by arrangements, not by preference. Germany and Saudi Arabia say this from opposite directions: impose the male-breadwinner model and employment falls; lift a constraint and it rises nineteen points in nine years. Whatever the old numbers measured, it was not what women wanted.

Three. Norms outlive the institutions that produced them by roughly a generation. East German women’s attitudes and employment patterns persisted for thirty years after the state that formed them ceased to exist. This cuts both ways and both are important: institutions can change norms, and neither direction is quick to undo.

Four. Rules land where they can be enforced, and so do reforms. Part Eleven found this about restrictions. Chapter Seven found the same shape about reforms: what a population receives is the text plus enforcement plus local interpretation plus whatever the displaced behaviour does next.

Five. Prohibiting a symptom while leaving the demand in place makes things worse. Four Indian reforms, one signature. This one predicts in advance rather than explaining afterwards, which is what makes it usable.

Six. Force works downwards and not upwards. Coercion has repeatedly and rapidly reduced births, and has never raised them. Iran ran both directions under one government within thirty years and got the same asymmetry Romania and China got.

In Real Terms

Put findings one and two together and you get something concrete enough to act on.

If bias responds to observation, and if the observed level of women’s participation is set by arrangements rather than preferences, then the two reinforce each other. Change the arrangement, women appear in a role, people observe them there, bias falls, and the next change is easier.

That is what the reserved seats did. A rotation rule put a woman in a chair; her presence did the rest.

The cheapest thing in this entire series was an administrative schedule that nobody designed to change anybody’s mind.

9.3 — What would change my mind

Stating this properly is the only thing that separates a document like this from advocacy, so here it is, specifically.

Finding one falls if a second randomisation, in a domain closer to marriage or property, produces no attitude movement. That study is entirely feasible and has not been run.

Finding three falls if the East German persistence turns out to be an artefact of who stayed — if the women who remained in the east were unrepresentative, the attitude gap could be a migration pattern rather than a norm.

Finding two weakens badly if Saudi participation falls back when the subsidies and quotas end. That is a real possibility and the next decade will show it.

The whole Afghanistan chapter weakens if outcomes over the coming years come in far below the projections — which would not mean the exclusion was harmless, but would mean the models overstated it and that everyone quoting them, including me, was quoting a forecast.

And Part Nine’s central claim — that no country has restored replacement fertility — falls the moment one does. Korea’s small recent rise is the live test and I do not know how it will resolve.

The Argument — does the whole body of evidence favour either side?

Thirteen parts. Here is the closing exchange, put as strongly as I can put both.

It favours the traditionalist

Every well-identified finding here is about capacity — women can lead, can work, can learn, and others’ bias about that shifts. Nobody ever seriously doubted any of it; the arguments in this series were never about whether a woman could run a village council. The traditionalist claim was always about what happens to a society that reorganises around individual choice, and on that the record is one-sided: fertility gone and unrecoverable, marriage dissolving, an ageing bill nobody can pay. The reformers won every measurable point and lost the thing the argument was about.

It favours the reformer

Look at what the traditionalist case has actually had to retreat to. It began as a claim about women — their nature, their wants, their capacities, what would happen to them personally. Every one of those has been tested and lost. What remains is an argument about aggregate demographic consequences, in which no individual woman is doing anything wrong and the cost is simply spread across a society. That is a completely different argument from the one being made to a girl in a village, and the fact that it is the only ground left is itself the finding.

Where things stand: both descriptions of the retreat are accurate, which is the honest and unsatisfying result of thirteen parts. The claims about women have been tested and have not survived. The claims about societies have been tested and largely have. Those are different claims and the same person usually makes both, which is why the argument has never resolved — winning one has never cost anybody the other.

What would settle it: nothing, and Chapter Eight explained why. This is a disagreement about what a society is for, arriving dressed as a disagreement about evidence.

Why people care so much: because both sides believe the other is not merely mistaken but is producing suffering — and on the evidence in this part, both are partly right about that too.

One more thing before the closing list, and it is about this part rather than about anybody in it.

The Hidden Assumption — the one in how I built this part

Chapter One warned that the natural experiments available to any writer are not a sample of what happens when rules change. They are a sample of the occasions when something visible happened, because a null result never becomes a famous case.

Then I wrote eight chapters using nine cases that I selected. Nobody handed me a list. There is no register of natural experiments in this field.

So ask why these nine. Because they are well documented, because researchers have already worked on them, because the findings are clean enough to explain in simple English, and — the one I least want to write down — because I had heard of them. Every one of those criteria correlates with something visible having happened. I selected on exactly the property Chapter One identified as the fatal error, and then used the resulting sample to draw conclusions.

There is a second turn and it is worse. I also graded the cases. I put the Indian reservation study in the top tier — which is correct on the checklist, and which is also the case whose finding I most wanted to be true. I do not think I bent the grading. I have no way to verify that from inside the document, and neither does anyone who trusts it.

The general form: running the audit with a sample you assembled yourself. It is the structural flaw in every literature review, every list of lessons from history, and every argument that begins “look at what happened in”. The method looks rigorous precisely because the selection happened before the method started.

What I can offer is partial and I will not dress it up. The checklist was written before the cases were graded and it is printed in Chapter One, so anyone can regrade them and see whether I was consistent. I included the chapter where reforms backfired and the chapter where the traditionalist forecast was right, which are the cases that cut against me — though I chose those too. And I have told you the direction the error runs.

That is not a solution. It is the difference between a flaw that is visible and one that is not, and after thirteen parts it is the most I have found a way to give you.

Which brings us, as every part does, to the accounting of what is known and what is not.

Remember This

Nine cases, graded. Only one is a genuine randomisation. Two are quasi-random or border comparisons. The famous ones — Iran, Afghanistan — are the weakest, and they are the ones everyone quotes.

Six findings survive: bias about women’s capacity is not fixed; the observed level of women’s work is set by arrangements rather than preference; norms outlive their institutions by about a generation; rules and reforms alike land where they can be enforced; prohibiting a symptom while leaving the demand makes things worse; and force works downwards but never upwards.

Each of those has a stated way it could fail, and Korea’s small recent fertility rise is the live test of the largest claim in the series.

The traditionalist case has retreated from claims about women — which have been tested and lost — to claims about societies, which have largely held. Those are different arguments, and the same people make both.

And I chose these nine cases myself, using criteria that all correlate with something visible having happened, which is the exact error Chapter One warned about. I can make that flaw visible. I have not found a way to remove it.

10An Honest List of What We Do Not Know

Two lists, as in every part. What is genuinely unknown, with the reason. And what is solid enough to build on. After thirteen parts, this is the chapter that matters most.

10.1 — Genuinely unknown

Whether beliefs about what women may do move like beliefs about what women can do

The randomised evidence shows bias about women’s capacity falling within a decade. Nothing tests whether beliefs about women’s permitted conduct — marriage, sexual behaviour, mobility, dress — respond the same way. Those are the beliefs this series is actually about. The reason we do not know: nobody has randomised anything in that domain, and it is hard to imagine an institution that could.

What the null cases look like

Every case in this part is one where something happened. The occasions on which a society changed its rules and nothing measurable followed are absent from the record entirely, because nobody writes them up. The reason we do not know: null results in case-based history are not published, not remembered, and not collected anywhere.

Whether the Iranian education boom was caused or merely not prevented

The enabling-constraint account — that segregated, religiously framed schooling let conservative fathers permit what they would otherwise have refused — is plausible and unproven. The reason we do not know: nobody has done the provincial analysis comparing pre-revolutionary conservatism against the size of the post-1979 enrolment jump, holding infrastructure spending constant. The census data to attempt it exists.

What Afghanistan’s outcomes actually are

The consequences everybody quotes are projections. The measurements have not been made, because the country’s statistical capacity collapsed alongside everything else. The reason we do not know: the state least able to be counted is generating the most quoted numbers about itself.

Whether the Saudi change holds

Nineteen points in nine years, arriving alongside quotas, subsidies and economic pressure. Whether participation stays if the inducements are withdrawn is the test of whether a constraint was lifted or an incentive was applied. The reason we do not know: not enough time has passed, and this one will resolve itself.

Whether East German persistence is a norm or a migration pattern

Attitudes and employment patterns persisted for thirty years after the institutions vanished. If the women who stayed in the east were systematically different from those who left, some of that gap is selection rather than culture. The reason we do not know: the analysis requires linking migration histories to attitude data across three decades, which is possible in principle and has not been done at the necessary scale.

Which reforms would have worked

Chapter Seven diagnosed four failures with one signature. It does not follow that structural reforms succeed, because the successful cases are as selected as the failures. The reason we do not know: there is no systematic accounting of reforms in this area scored by whether they targeted a symptom or a structure, and assembling one would be a large piece of work that nobody has funded.

How We Actually Know This

One item on that list deserves to be treated as a finding rather than a gap, because it recurs in every part of this series and it always points the same way.

Three of the seven unknowns above could be closed by analyses that require no new data collection at all. The Iranian provincial comparison uses existing censuses. The East German migration question uses existing panel data. The reform accounting uses published evaluations.

These are not expensive. They are not politically dangerous in most of the places that could run them. They are simply not the studies that get commissioned.

Set that alongside what earlier parts found: no Indian survey asks whether a woman has private use of her phone; no national survey asks whether her private messages have been used against her; the honour-killing count is produced by a state that counts fraud to the rupee. The pattern across thirteen parts is not that this subject is hard to study. It is that the specific studies which would settle it keep not being done.

10.2 — Solid

Bias about women’s capacity is not fixed

From the only randomisation in this series. Rotation by village serial number, rural north India, two electoral cycles: evaluations of female leaders improved, more women won unreserved seats afterwards, and the gap between what parents wanted for sons and daughters narrowed by about a quarter. This passes all six tests.

The observed level of women’s paid work is set by arrangements, not preference

Germany imposed the male-breadwinner model on the east and employment fell. Saudi Arabia lifted constraints and participation went from about 17 per cent to about 36 per cent in under a decade. Two countries, opposite directions, same conclusion.

Visible returns change family behaviour faster than payments do

Bangladeshi girls near garment factories married later, had children later and stayed in school longer — by more than a scheme that paid families directly to keep them there.

Norms outlive the institutions that made them

East German women’s employment patterns and attitudes persisted for roughly thirty years after the state that formed them ceased to exist.

Institutional collapse produces fertility collapse

East German fertility fell from about 1.5 to below 0.8 in four years when the arrangements supporting early childbearing were removed — one of the sharpest peacetime declines ever recorded, and it recovered as the arrangements were rebuilt.

Strict rules and mass female education are compatible

Refuted as a universal claim by Iran alone: female literacy from about a third to above eighty per cent, and women a majority of university entrants, under compulsory veiling and segregation.

Prohibiting a symptom relocates the behaviour

Dowry prohibited in 1961 and it grew, moving from visible goods into undocumented cash. A marriage-age law announced in 1929 produced a documented surge of child marriages before it took effect. Protective labour laws removed the women rather than the danger.

Falling fertility plus son preference kills girls, and prosperity does not prevent it

India’s child sex ratio fell from about 962 girls per 1,000 boys in 1981 to about 918 in 2011, and sex selection was strongest among richer, more educated, urban families — because small families plus son preference require action, and those families had access to the scan.

Force works downwards and not upwards

Coercion has repeatedly and rapidly reduced births and has never raised them. Iran ran both directions under one government within thirty years and got the same asymmetry as Romania and China.

Remember This

Unknown: whether beliefs about what women may do move like beliefs about what women can do; what the null cases look like, since nobody records them; whether Iran’s education boom was caused or merely not prevented; what Afghanistan’s outcomes actually are rather than are projected to be; whether the Saudi change holds without its subsidies; whether East German persistence is culture or migration; and which reforms would have worked.

Solid: bias about women’s capacity is not fixed; the level of women’s paid work is set by arrangements rather than preference; visible returns beat payments; norms outlive their institutions by a generation; removing supporting arrangements collapses fertility; strict rules and mass female education are compatible; prohibiting a symptom relocates the behaviour; small families plus son preference kill girls and prosperity does not prevent it; and force works only downwards.

Three of the seven unknowns need no new data — only analyses of records that already exist, which are cheap, safe, and consistently not commissioned.

After thirteen parts the pattern is unmistakable: this subject is not hard to study. The specific studies that would settle it keep not being done.

Sources & further reading — Part 13

Timeline

A century of experiments nobody designed as experiments.

YearWhat happened
1926Turkey adopts a civil code abolishing polygamy and religious marriage — the sharpest legal break of its kind anywhere, and the model reformers across the region argued about for the rest of the century.
1929India’s act raising the minimum marriage age is passed in the autumn to take effect the following spring. The census two years later records a surge of child marriages in the gap.
1945–49Korea and Germany are divided. Two populations with shared language and ancestry begin four decades under opposite theories of what women are for.
1961India prohibits dowry. Over the following decades the practice spreads into communities that had not had it, and the amounts rise.
1966Romania issues Decree 770. Births nearly double within a year, then decay back over a decade, at a cost of thousands of women’s lives.
1967–75Iran’s Family Protection Law raises the marriage age, restricts polygamy and gives women grounds for divorce.
1977West Germany removes the provision allowing a husband to forbid his wife’s employment. East German women had been in near-universal employment for two decades.
1979The Iranian revolution reverses all of it within a year. Female literacy then rises from about a third towards universal, and fertility begins the fastest collapse ever recorded.
1980sBangladesh’s export garment industry spreads, placing factories by roads and ports rather than by local attitudes — creating, accidentally, a comparison group.
1990German reunification. East German fertility falls from about 1.5 to below 0.8 in four years as every arrangement supporting early childbearing is removed at once.
1993India’s 73rd constitutional amendment reserves a third of village council leaderships for women and rotates them by village serial number — the only genuine lottery in this subject.
1994India legislates against using diagnostic techniques to determine the sex of a foetus, thirteen years after the child sex ratio began falling.
1996–2001The first Afghan exclusion of women from education and public life, poorly documented at the time.
2004The first published analysis of the reservation lottery finds women in charge of village councils spent differently — on what the women there had actually asked for.
2011India’s census records about 918 girls per 1,000 boys aged nought to six, down from about 962 in 1981.
2012Further work on the same lottery finds two rounds of a female village leader narrowed the gap between what parents wanted for sons and daughters, and closed the school attendance gap.
2013A garment factory building collapses in Bangladesh, killing more than eleven hundred people — the women whose employment had delayed marriages and filled schools.
2016Saudi Arabia announces its national economic programme with a target of 30 per cent female labour force participation. The rate at the time is around 17 per cent.
2018–19The Saudi driving ban is lifted in June 2018; guardianship rules on passports and travel are relaxed in August 2019.
2021In September, Afghan girls are barred from school beyond the sixth grade.
2022In December, Afghan universities are closed to women, and employment with aid organisations is barred.
2024Afghan restrictions extend to midwifery and nursing training — closing the last route to producing the only people permitted to attend Afghan women in childbirth.
2025–26Saudi female labour force participation reaches around 36 per cent. An assessment published in April 2026 estimates the Afghan education ban has already cost about half a per cent of national output.
NextPart Fourteen starts here.

Glossary

Every hard word used in this part, in plain English.

TermWhat it means
CounterfactualWhat would have happened otherwise. The thing every claim is compared against and nobody can observe.
Difference-in-differencesComparing how much each of two groups changed, rather than comparing their levels. Removes anything that affected both.
Enabling constraintA restriction that makes an activity possible by making it acceptable. The proposed explanation for why segregated schooling raised girls’ enrolment in Iran.
Implicit biasAn automatic association that shapes judgement without the person intending it. Measured by comparing ratings of identical material attributed to a man or a woman.
Male-breadwinner modelArrangements built on the assumption of one earner and one unpaid worker per household — assembled from tax rules, school hours and childcare, none of which mentions women.
Natural experimentA real-world situation where something outside anyone’s control assigned people to different conditions for reasons unrelated to the outcome being studied.
Perverse incentiveWhen a rule creates a reason to do more of the thing it was meant to stop.
ProjectionA modelled estimate of what should follow, as opposed to a measurement of what did. The distinction collapses constantly in reporting.
Reference-group effectPeople rate their lives against what they believe was available to them. If the comparison group changes, the score changes without the life changing.
Revealed preferenceThe idea that what people do tells you more than what they say. Its limit: a preference expressed under constraint is not a free preference either.
Substitution effectWhen blocking one route redirects a behaviour into another channel rather than removing it.
Child sex ratioThe number of girls per 1,000 boys in a given age group. India’s fell from about 962 in 1981 to about 918 in 2011 among children under six.
ConfoundingWhen something else changed at the same time, so an effect cannot be attributed to the thing being studied. The failure mode of every revolution used as evidence.
The isolation testShorthand in this part for the fifth of the six checks: did the change arrive alone, or did a hundred other things arrive with it?
Quasi-experimentA comparison where assignment was not random but was close enough to unrelated — factory placement decided by roads rather than by local attitudes, for instance.
Selection on the outcomeAssembling your evidence from cases that became famous because something visible happened, then treating the sample as representative.

Download The complete book · 4.0 MB