Part 05 of 16The Paradox

The Paradox of the Free Countries

In the freest countries on earth, men’s and women’s choices diverge more, not less. The finding keeps arriving from unrelated directions. It is also contested — and India is sitting inside it.

Where We Left Off

Before we begin

Part Four measured every difference between men and women that has been measured, and reported each one with its size attached. Most cognitive differences turned out to be close to a coin toss. Physical differences turned out to be enormous and narrow. And the largest psychological difference of all turned out not to be an ability at all but a preference — the Things–People dimension, at roughly 0.93, from over half a million people.

Three times in that part I ran into the same finding and deferred it. Interest differences appeared to be larger in richer and more gender-equal countries. Personality differences did the same. So did measured preferences over risk and patience. Each time I said that Part Five would deal with it.

This is Part Five, and it deals with nothing else.

The reason it needs a part to itself is that it is the single most consequential empirical claim in this series. If it holds, it damages the simplest version of the argument that these differences are produced by constraint — because removing constraint should shrink them, and it appears to do the opposite. And if it does not hold, then a great deal of what has been built on it over the last twenty years has to come down.

I want to be honest about the shape of this part before you start it. It is the most technical part of the series and the least conclusive. There is a genuine, unresolved, professionally conducted dispute at the centre of it, involving competent researchers on both sides, and I am not going to pretend otherwise or pick the side that suits the story. Four chapters of this part are about the finding and four are about why it might not mean what it appears to mean.

What I can promise is that by the end you will know exactly which parts of it are solid, which are contested, and which are being oversold — including by people you agree with.

How to Read the Boxes

The notation, in case this is where you started

Six kinds of box run through the series. One live example of each.

Word Box

The gender-equality paradox: the observed pattern in which differences between men and women — in occupational choice, personality, and stated preferences — appear larger in countries that are richer and score higher on measures of gender equality.

It is called a paradox because it is the opposite of what the most straightforward account predicts. If differences are produced by constraint, removing constraint should shrink them.

Why it matters here: this is the whole subject of this part, and it is the finding that most seriously complicates the argument made in Parts Two and Three.

A Word Box appears the first time a hard word does — never later, never only in a glossary.

In Real Terms

Here is the finding in its rawest form, with no statistics at all.

Look at the share of science, technology, engineering and mathematics graduates who are women, country by country. Near the top of that list you find Algeria, Tunisia, Albania, Oman and the United Arab Emirates.

Near the bottom you find the Netherlands, Switzerland, Belgium, Norway and the United States.

That is not a subtle statistical artefact requiring a specialist to see. It is a list, and the list is upside down relative to every expectation anybody would have brought to it.

Chapters Three to Five are about whether the list means what it looks like it means. Chapter One is about taking it seriously first.

Any number too big to picture gets a body attached.

How We Actually Know This

The evidence in this part comes from three very different kinds of source, and they have different weaknesses, which matters enormously for Chapters Four and Five.

Administrative records — how many women actually graduated in which subject, collected by national education ministries and compiled internationally. Nobody is reporting a feeling. These are counts.

Standardised tests — international assessments of fifteen-year-olds, using identical instruments in dozens of countries. Also not self-report about traits, though the questionnaires attached to them are.

Self-report surveys — asking people to rate their own personality, interests, patience or willingness to take risks. This is where most of the personality evidence comes from and it carries a specific cross-country problem set out in Chapter Two.

Why the distinction matters: the most serious objection to the paradox applies to the third category and not to the first. Any argument that dismisses the whole finding on measurement grounds has to explain the graduation counts, which are not measurements of anybody’s mind.

That split runs through the whole part and is worth carrying with you from here.

The Argument — is the paradox a finding or an artefact?

The dispute at the centre of this part, stated at the front so you know what is coming.

A finding

It appears in graduation statistics, in international assessments, in personality inventories, and in behavioural preference measures, produced by different teams using different instruments in different decades. A pattern that keeps arriving from unrelated directions is not usually an artefact of any one of them.

An artefact

Much of it depends on how you choose to measure “the gap,” and different reasonable choices produce different answers from identical data. Much of the rest rests on self-reports, which are known not to be comparable across countries in the way this argument requires. And the equality indices being used measure something odd.

Where things stand: both are partly right and the honest position is that the pattern is real in some measures and shaky in others. This part sorts them one at a time. Chapter Ten states which is which without hedging.

An Argument box appears where people who have studied something genuinely disagree.

The Hidden Assumption

Both sides above assume that the paradox, if real, would settle something.

One camp treats it as proof that the differences are innate and were merely being suppressed. The other treats it as so damaging that it must be dismantled. Both are behaving as though the result decides the argument.

It does not, and Chapter Eight is about why. A growing difference needs explaining just as much as a shrinking one, and “freedom lets what was always there come out” is only one of several accounts that fit — several of which are entirely social.

The general form: treating a result as decisive because both sides have agreed in advance that it would be.

This is the signature box. There are five in this part.

Remember This

Every chapter closes with one of these, restating it in the plainest words available, key terms in bold.

Read only these and you should still finish holding the whole argument.

A Note on This Part

Read this before Chapter One

This is the least conclusive part of the series and I am not going to disguise that. The previous four parts each ended with something solid. This one ends with a finding that is partly established, partly contested, and entirely unexplained. If that is unsatisfying, it is unsatisfying because the state of the evidence is unsatisfying, and manufacturing a conclusion here would be the worst thing I could do.

The finding is politically loaded in a specific way, and I want to name it in advance. It is quoted constantly by people arguing that efforts to increase women’s participation in technical fields are futile or misguided. It is dismissed just as constantly by people who have not read the criticisms and are relying on somebody else having answered them. I have tried to give the strongest version of each, and the reader who finishes this part with their prior view fully intact has probably read it badly.

And Chapter Nine is where this stops being abstract for an Indian reader. India sits in an unusual position on every measure in this part, and the Indian numbers turn out to say something quite different from what either camp expects. If you read one chapter, read that one.

1The Result

Algeria produces a higher share of female science graduates than Norway does. That sentence is true, it is not disputed by anybody, and everything in this part follows from working out what it means.

1.1 — The list

Take the share of graduates in science, technology, engineering and mathematics who are women, and rank the countries of the world by it.

Near the top you find Algeria, Tunisia, Albania, Oman, the United Arab Emirates. Near the bottom you find the Netherlands, Switzerland, Belgium, Norway and the United States.

Nothing about this requires interpretation. These are counts of graduates, compiled by education ministries, collated internationally. Nobody is being asked how they feel about anything.

And the ordering is upside down relative to every expectation anybody would bring to it. The countries with the strongest legal protections for women, the longest histories of feminist organisation, the most funding aimed specifically at getting girls into science, and the highest scores on every index of gender equality — those countries produce proportionally fewer female scientists and engineers than countries where women’s lives are, by almost any measure, considerably more restricted.

1.2 — The study that made it a subject

The list had been visible in international data for some time. In 2018 two researchers made it into an argument, in a paper that has been among the most cited and most fought-over in this field ever since.

How We Actually Know This

The study combined two sources. The first was an international assessment of around 470,000 fifteen-year-olds across 67 countries, testing science, mathematics and reading with the same instruments everywhere. The second was national data on what subjects people actually graduated in.

What they found in the test data: girls matched or outperformed boys in science achievement in the majority of countries. This is not a story about girls being worse at science, and the authors said so explicitly.

What they found in the graduation data: the share of STEM graduates who were women correlated negatively with a widely used index of national gender equality. More equal country, smaller share.

What it cannot show: why. The paper proposed an explanation involving economic security, which is discussed in Chapter Six and is not established by the correlation itself. And the way the graduation gap was calculated became the centre of a serious methodological dispute, which is Chapters Three and Four.

1.3 — The finding inside the finding

There is a second result in that paper which is more interesting than the headline and gets almost no attention. It concerns not how good students are, but what each student is relatively best at.

Consider a girl who scores in the ninetieth percentile in science. That is excellent. Now suppose she also scores in the ninety-seventh percentile in reading.

She is very good at science. Science is also, for her personally, her weaker subject.

Now consider a boy in the eighty-fifth percentile in science and the sixtieth in reading. He is less good at science than she is, in absolute terms. But science is clearly what he is best at.

Ask each of them what they are good at, or ask a careers adviser, or simply let each follow their comparative advantage, and the two will sort in opposite directions — despite the girl being the better scientist of the two.

Word Box

Relative strength (researchers say intra-individual strength): what a person is best at compared with their own other abilities, rather than compared with other people.

Why it matters here: this is a completely different quantity from ability, and it can move in the opposite direction. A population of girls could be better at science than boys in every country on earth, and still contain fewer people for whom science is their personal best subject — because of Part Four’s finding that girls read and write better, everywhere measured.

The consequence is uncomfortable for everybody: the reading advantage from Part Four may be part of what steers girls out of science. Being better at something else is a reason not to do this thing.

That mechanism deserves to be much better known than it is, because it requires no discrimination, no discouragement and no difference in scientific ability to produce a sorting effect. It requires only that people, or the adults advising them, pay attention to what each individual is comparatively best at — which is what everybody thinks good careers advice consists of.

1.4 — It is not only science

Before testing the result, one broadening is needed, because a great deal of argument about the paradox proceeds as though it were a fact about engineering departments. It is not.

Ask a wider question: across a whole economy, how separated are men’s and women’s jobs? Not whether women earn less, and not whether they work — simply whether the sexes are doing the same jobs as each other or different ones.

Measured that way, the Nordic countries are among the most occupationally segregated economies in the developed world. Not the least. Among the most.

This is not a hidden result and it does not depend on any disputed technique. It falls out of ordinary labour statistics. Those countries have very high female employment — among the highest anywhere — and that employment is heavily concentrated in a large public sector delivering health, care and education, while the private technical and industrial sector remains heavily male.

So the shape of the finding is this. In the countries with the strongest equality institutions, more women work than almost anywhere, and they work in different jobs from men more sharply than in many countries with weaker institutions.

Two things follow that are worth holding for the rest of the part.

The gap runs in both directions and only one is treated as a problem. Men are rarer in nursing and primary teaching in rich countries than women are in engineering, by a wide margin in several. That imbalance attracts a fraction of the attention, the funding and the concern. Part Four made this point about interests; here it appears in employment statistics.

And "more women working" and "men and women doing the same things" are separate goals that can move in opposite directions. A country can succeed enormously at the first while going backwards on the second, and several have. Anybody using the word "progress" here has to say which of the two they mean.

1.5 — Is the headline result real?

The Argument — is there a gender-equality paradox in STEM?

Stated at the front, then taken apart across Chapters Three, Four and Five.

Yes, and the raw data shows it

The country list in §1.1 requires no statistical technique at all. Women are around forty per cent or more of STEM graduates in several Arab and North African countries and around a fifth to a third in much of northern and western Europe. That gap is large, consistent across data sources and years, and nobody disputes the underlying counts.

The counts are real; the paradox is a construction

The word “paradox” requires the second variable — gender equality — to be measured in a way that makes the relationship meaningful, and Chapter Five shows the indices used are strange objects. It also depends on choosing a particular way to express the gap, and Chapter Three shows that other reasonable choices weaken or remove the relationship. A robust raw pattern can still support a fragile claim.

The narrow version that survives everything

Strip out the contested measures and this remains: there is no simple positive relationship between a country’s gender equality and women’s share of technical fields, and in the raw counts the relationship runs the wrong way. That is far weaker than “freedom causes divergence” and it is enough to destroy the assumption that removing barriers straightforwardly increases representation.

Where things stand: the third position, and it is where I will leave it until Chapter Ten. The strong version is contested. The narrow version is not, and the narrow version is already inconvenient for a very widely held belief.

What would settle it: tracking within countries over time rather than comparing across them. If a single country becomes more gender-equal and its STEM share falls, that is much harder to explain away than a snapshot comparison between Norway and Algeria. Some of this data exists and Chapter Eight looks at what it shows.

Why people care so much: because the strong version is used as an argument against doing anything, and the people using it that way rarely mention that a narrow version is all that is established.

The narrow version is where I will leave the finding until Chapter Ten. Everything between here and there is an attempt to work out whether anything stronger survives.

Remember This

Rank countries by the share of science and engineering graduates who are women. Near the top: Algeria, Tunisia, Albania, Oman, the UAE. Near the bottom: the Netherlands, Switzerland, Belgium, Norway, the United States. These are graduation counts, not opinions.

The study that made this famous used an international test of 470,000 fifteen-year-olds in 67 countries plus national graduation data. Girls matched or beat boys in science achievement in most countries. This was never a story about girls being worse at science.

The finding inside the finding: what matters for sorting is not ability but relative strength — what a person is best at compared with their own other subjects. A girl in the ninetieth percentile in science and the ninety-seventh in reading is an excellent scientist for whom science is a personal weakness.

That means Part Four’s reading advantage may be part of what steers girls out of science. Being better at something else is a reason not to do this thing — and it needs no discrimination, no discouragement, and no difference in scientific ability to work.

The version that survives every criticism in this part: there is no simple positive relationship between a country’s gender equality and women’s share of technical fields, and in the raw counts it runs the wrong way. That is much weaker than “freedom causes divergence” — and it already destroys a very widely held assumption.

2It Keeps Arriving From Everywhere

If the pattern appeared only in graduation statistics it would be easy to dismiss. It appears in personality inventories, in behavioural preference measures, and in vocational interests — different teams, different instruments, different decades, same direction.

2.1 — Personality

The first place this pattern was noticed was not in education data at all. It was in personality research, and it was noticed by people who were not looking for it and did not expect it.

A study across twenty-six cultures in 2001 found that sex differences in personality traits were larger in European and American samples than in African and Asian ones. The authors described the result as counterintuitive, because the expectation had been the reverse.

A larger study seven years later covered fifty-five nations and found the same thing more clearly: differences on the standard personality dimensions from Part Four were bigger in prosperous, healthy and more egalitarian countries. Further work using different inventories and different country sets has repeatedly reproduced it.

2.2 — Preferences

Word Box

A nationally representative sample is chosen so that its composition matches the country — the same mix of ages, regions, incomes and education levels as the population it is drawn from.

The contrast is with a convenience sample: whoever was available. University students, people who answered an online advertisement, volunteers.

Why it matters here: Part Four noted that most psychology is built on convenience samples of unusual people. The study described below is not, and that is the main reason it carries more weight than most of what is in this part.

The strongest single piece of evidence for the pattern comes from economics rather than psychology, and it is worth setting out carefully because its design answers several objections in advance.

How We Actually Know This

A team surveyed roughly eighty thousand people across seventy-six countries, using nationally representative samples rather than students or volunteers. They measured six economic preferences: patience, willingness to take risks, altruism, trust, and willingness to reward or punish others.

The design detail that matters: the survey questions were not invented and hoped for the best. They were first validated in laboratory experiments in which people made real choices with real money, and the questions retained were the ones that predicted actual behaviour.

What they found: sex differences in these preferences were larger in richer countries and larger in more gender-equal countries — and each of those two predicted the gap independently of the other. That second point matters, because wealth and gender equality travel together and it would be easy for one to be doing all the work.

What it cannot show: causation, like everything else in this part. And although the items were behaviourally validated in a laboratory, the survey itself is still people answering questions about themselves, which leaves it exposed to the problem in the next section.

2.3 — And several more

The pattern is not confined to those two literatures. Once researchers started looking for it, it turned up in a series of places nobody had connected.

Basic values. Large cross-national surveys asking people what they consider important in life — security, achievement, benevolence, self-direction — find sex differences in the ranking, and the differences are generally larger in wealthier and more egalitarian societies.

Self-esteem. Measured across dozens of nations, the sex gap in reported self-esteem is wider in more developed countries.

Mate preferences. What men and women say they want in a partner differs everywhere it has been asked. Several analyses find the differences on some dimensions narrowing with gender equality and others not — which is worth flagging as an exception rather than smoothing over, because it is one of the few places the pattern is genuinely mixed.

Physical traits. Even the sex difference in height is somewhat larger in well-nourished populations, for a reason that has nothing to do with psychology: boys’ growth is more responsive to nutrition, so improving everybody’s diet raises male height more. This is a useful case, because here we know the mechanism, and the mechanism is not that something innate was released. It is that a constraint was lifted unevenly.

Hold that last one. It is a worked example of a growing gap with an entirely environmental explanation, and Chapter Eight will need it.

2.4 — Interests

The Things–People dimension from Part Four has been measured across more than fifty nations. The difference appears everywhere it has been looked for, without exception, which is itself notable. And it tends to be larger in more developed countries.

So we now have the same directional pattern in four separate literatures: educational choice, personality, economic preferences, and vocational interests. Different researchers, different instruments, different sampling methods, different decades.

That convergence is the strongest argument this finding has. A single result can be an artefact of one method. A result that keeps appearing when you change the method is much harder to explain that way.

2.5 — And the objection that applies to most of it

Now the counterweight, and it is a serious one that most people repeating the paradox have never encountered.

Word Box

The reference group effect: when people rate themselves on a scale, they implicitly compare themselves to the people around them rather than to humanity in general.

Ask somebody whether they are “tall” and they answer relative to their neighbours, not relative to the world. Ask whether they are “assertive” and the same thing happens.

Why it matters here: it means self-ratings are not straightforwardly comparable across countries, because the yardstick moves with the country. A person can rate themselves identically in two societies whose actual behaviour differs enormously, or differently in two societies where it does not.

There is a specific version of this that could produce the entire paradox in the self-report data without anything else being true.

Research on social comparison suggests that in more individualistic and more gender-equal societies, people are more likely to compare themselves across gender lines — a woman measuring herself against people generally rather than against other women. In more gender-segregated societies, comparisons happen more within gender.

If that is right, then in an egalitarian country a woman rating her own assertiveness is comparing herself to a mixed population including men, which pushes her rating down; a man doing the same is comparing himself to a mixed population including women, which pushes his up. The gap in the ratings widens with no change whatever in anybody’s actual assertiveness.

That is a real mechanism, it is documented, and it is enough on its own to account for a substantial part of the personality findings.

2.6 — Which is why the graduation counts matter

Here is the analytical point that determines how the rest of this part goes, and it is worth stating flatly.

The reference group objection applies to self-report. It does not apply to counting graduates.

Nobody asks a university to rate its own science department against its neighbours. The number of women who completed an engineering degree in Algeria last year is a count. It carries no yardstick problem, no self-comparison, no translation issue about what “assertive” means in two languages.

So the evidence in this part separates into two piles, and they have to be judged separately.

Pile one — self-reported traits and preferences. Large, convergent, and exposed to a documented measurement objection that could generate the pattern by itself.

Pile two — behavioural and administrative records. Smaller in scope, but immune to that objection. This is where the graduation data sits.

Any argument that dismisses the paradox entirely on measurement grounds has to explain pile two. And any argument that treats the whole thing as established has to reckon with the fact that most of the impressive convergence is in pile one.

Remember This

The pattern appears in four separate literatures: educational choice, personality across fifty-five nations, economic preferences across seventy-six, and vocational interests across more than fifty. Different teams, instruments and decades, same direction.

The preference study is the strongest single piece: eighty thousand people, nationally representative, with survey items first validated against real choices with real money — and wealth and gender equality each predicted larger gaps independently of the other.

Convergence from unrelated instruments is the best argument this finding has. A single result can be an artefact of one method; a result that survives changing the method is much harder to dismiss.

But there is a serious objection. The reference group effect: people rate themselves against those around them, so self-ratings are not comparable across countries. And in more equal societies people compare themselves across gender lines, which would widen the reported gap with no change in anybody’s actual behaviour.

That objection applies to self-report and not to counting graduates. The evidence splits into two piles — convergent self-reports exposed to a real measurement problem, and administrative counts that are immune to it. Anybody dismissing the whole thing must explain the counts; anybody treating it as settled must notice that most of the impressive convergence is in the vulnerable pile.

3How To Measure A Gap

There are at least three reasonable ways to say how big a difference between men and women is, and applied to identical numbers they give different answers — sometimes opposite ones. This is the single most important technical idea in the part.

3.1 — Three questions wearing one word

Suppose you want to know how unequal a country’s engineering intake is. It sounds like one question. It is at least three.

What fraction of engineers are women? This is the representation question. It is what people usually mean and what gets reported.

How much more likely is a man than a woman to be an engineer? This is the ratio question, and it can move in a different direction from the first.

Given that somebody went to university at all, how likely were they to pick engineering? This is the conditional question, and it is the one that produced the dispute at the centre of this part.

These are all legitimate. They are also different, and they answer to different concerns.

Word Box

A correlation is a number saying how reliably two things move together across a set of cases — here, across countries.

It runs from −1 to +1. Zero means no relationship. Positive means they rise together. Negative means one rises as the other falls. The paradox is the observation that a particular correlation is negative where everybody expected positive.

Two things a correlation across countries can never do, both of which matter here. It cannot tell you the direction of cause. And it cannot tell you anything about an individual, because a country is not a person — which is Part Four, Chapter One applied at national scale.

3.2 — Watch them disagree

In Real Terms

Two invented countries, with numbers simple enough to check by hand.

Country A — rich. Two hundred people finish university: a hundred women and a hundred men. Twenty of the women study science; forty of the men do. So sixty science graduates in total.

Country B — poor. A hundred and fifty finish university: fifty women and a hundred men, because fewer women get that far at all. Twenty of the women study science; fifty of the men do. Seventy science graduates in total.

Now ask the questions.

What share of science graduates are women? Country A: 20 of 60, about 33 per cent. Country B: 20 of 70, about 29 per cent. The rich country looks better.

Of the women who went to university, what fraction chose science? Country A: 20 of 100, 20 per cent. Country B: 20 of 50, 40 per cent. The poor country’s women chose science at twice the rate.

Same numbers. Opposite conclusions. Nobody cheated.

The reason is hiding in plain sight: in the rich country twice as many women reached university at all. That larger pool inflates their share of every subject, including science — and conceals the fact that a much smaller proportion of them chose it.

This is not a trick and it is not an obscure edge case. It is the ordinary situation, because the countries that score highest on gender equality are also the countries where women’s overall university participation is highest. The two measures pull apart precisely where the argument is happening.

The Hidden Assumption

Everybody in this argument assumes that a gap has an obvious size — that “how unequal is this” is a fact you read off the data rather than a thing you construct.

It is not. Every way of expressing a gap contains a buried assumption about what equality would look like, and the assumptions are not the same.

Share of graduates assumes the target is a workforce that mirrors the population. It is a question about outcomes.

Conditional propensity assumes the target is that a man and a woman with the same access make the same choice. It is a question about choices, holding access fixed.

Those are different political positions, not different statistical techniques. And the second one does something specific that is rarely noticed: by holding access constant, it removes from view the single biggest inequality in most poor countries — that far fewer women get to university at all. A measure designed to isolate choice necessarily discards the constraint.

The general form: a normative assumption smuggled inside a formula. Whenever somebody reports a gap, the first question is not whether the number is correct. It is which of these questions they answered, and almost no reporting says.

3.3 — Why this became the dispute

The 2018 study used a version of the conditional measure. It asked, in effect, how strongly women lean towards science relative to how strongly men do, given that both are graduating.

That is a defensible choice. It is arguably the right measure if the question you care about is whether men and women with equal access choose differently, which was the question the authors were asking.

It is also a choice that systematically flatters the poorer countries in the comparison — for the reason set out in the box above. And it was not described in the paper in a way that let readers work out precisely what had been computed.

Which is where Chapter Four begins.

Remember This

“How big is the gap” is at least three questions. What share of engineers are women (representation). How much likelier is a man to be one (ratio). Given that somebody went to university, how likely were they to choose it (conditional).

Applied to identical numbers these give different answers and sometimes opposite ones. In the worked example, the rich country has a higher share of female science graduates while the poor country’s women chose science at twice the rate — because twice as many women in the rich country reached university at all.

Every way of expressing a gap contains a buried assumption about what equality would look like. Share of graduates assumes the target is a workforce mirroring the population. Conditional propensity assumes the target is equal choices given equal access.

And the conditional measure does something rarely noticed: by holding access constant it removes from view the biggest inequality in most poor countries — that far fewer women reach university at all. A measure built to isolate choice must discard the constraint.

When somebody reports a gap, the first question is not whether the number is right. It is which question they answered — and almost no reporting says.

4The Critics

Two years after the paper appeared, a group of researchers published a formal challenge. A correction was issued. Both sides then said they had been vindicated. Here is what actually happened, and what each side is entitled to claim.

4.1 — What was challenged

The challenge had three parts, and they are of very different strength. Separating them is the whole job of this chapter.

First, a reproducibility complaint. The critics reported that they could not reconstruct the study’s headline measure from the description given in the paper. Working from the published text, they could not obtain the published numbers.

Second, a sensitivity complaint. They argued that when more conventional measures of the gap were used — of the kind in §3.1 — the relationship with gender equality was weaker, and depending on the specification could largely disappear.

Third, a complaint about the explanation. The original paper proposed that economic security and life satisfaction accounted for the pattern. The critics argued this was not adequately supported by the analysis presented.

4.2 — What happened next

A correction was published. It set out precisely how the measure had been computed, which had not been clear in the original.

The authors maintained that their substantive conclusion was unaffected and replied to the criticisms. The critics maintained that the sensitivity problem stood regardless of the clarification.

Both are still saying so. There is no referee.

4.3 — Scoring it honestly

Here is what I think each side is entitled to, and I want to be explicit that this is my reading of a live dispute rather than a settled verdict.

The critics are entitled to the first complaint entirely. A measure that readers cannot reconstruct from the paper is a reporting failure, and it is a serious one, particularly for a result that became this influential this fast. The correction concedes this by existing.

The critics are substantially entitled to the second complaint. Chapter Three showed why: the choice of measure genuinely changes the answer, and the choice made was one that favours the paper’s conclusion. Saying so is not an accusation of bad faith — the choice is defensible, and Chapter Three explains why somebody would make it. But a result that depends on a contested measurement decision is weaker than one that does not, and it should be described that way.

The critics did not overturn the raw pattern, and this is the part their supporters routinely overstate. The country list in Chapter One does not use the disputed measure. It does not use any measure. It is a ranking of counts, and it still runs the wrong way. Whatever happened to the propensity calculation, Algeria still produces a higher share of female science graduates than Norway.

And the original authors are not entitled to the explanation. The economic security account is a hypothesis that fits, and it was presented with more confidence than the analysis supported. Chapter Six treats it as a hypothesis, which is what it is.

The Argument — did the critique refute the paradox?

Both camps quote this exchange, and almost nobody quoting it has read both sides.

It was refuted

A headline result that cannot be reproduced from its own methods section, that required a published correction, and that weakens or vanishes under standard alternative measures, is not a finding. It is an artefact of a measurement choice made by researchers who had a conclusion in view. That it was quoted worldwide for two years before anybody checked is an indictment of how this literature works.

It was dented, not refuted

The critique addressed one paper. The pattern predates that paper, appears in the raw counts without any derived measure, and appears independently in personality, preferences and interests as Chapter Two showed. Refuting a single analysis does not remove a pattern that shows up in four literatures. Treating this exchange as a general disproof is exactly the error the critics accused the original authors of.

What each actually established

The critics established that this paper’s specific quantitative claim was fragile and badly documented. The original authors established that something in the raw data runs contrary to the obvious expectation. Neither established what causes it, and neither has ever claimed to have tested a causal mechanism.

Where things stand: the third position. The exchange is a good example of scientific criticism working — a fragile claim was probed, a correction issued, and the residue is smaller and better specified than the original. That is a success, not a scandal, and reading it as a scandal in either direction misses what happened.

What would settle it: pre-registered analysis specifying the measure in advance, applied to new data, ideally tracking countries over time rather than comparing them at one moment. This has not been done.

Why people care so much: because the strong version is used to argue that nothing should be attempted, and the refutation is used to argue that nothing needs explaining. Both conclusions are larger than the evidence, and both are more comfortable than the actual state of knowledge, which is that a real oddity in the data has no agreed explanation.

4.4 — How to read a dispute like this

A general lesson, because you will meet this shape again and it is worth having a procedure.

Separate the raw observation from the derived measure. The country list and the propensity calculation are different objects and only one of them was challenged. Most public reporting collapses them.

Ask whether the correction changed the conclusion or the description. Corrections come in two kinds and they are treated identically in headlines. This one clarified how a number was produced. That is a real failing and it is not the same as a retraction.

Notice who is being quoted by whom. Both papers are now cited far more often as ammunition than as evidence. Part One, Chapter Six applies: when a countable question is fought this hard, something other than the count is at stake.

And check whether the critique addresses the convergence. This is the one that matters most here. The critique concerned one study in one domain. Chapter Two’s finding — that the same pattern appears in personality, preferences and interests, measured by unrelated teams — is untouched by it, and is also, as Chapter Two showed, exposed to a completely different objection of its own.

Remember This

The critique had three parts. That the headline measure could not be reconstructed from the paper. That the result weakens or vanishes under conventional alternative measures. And that the proposed explanation was not supported by the analysis.

A correction was published clarifying the computation. The authors maintained their conclusion. The critics maintained the sensitivity problem stood. There is no referee.

My reading: the critics are entitled to the reproducibility complaint entirely, substantially entitled to the sensitivity complaint, and did not overturn the raw pattern — because the country list uses no derived measure at all. Algeria still produces a higher share of female science graduates than Norway.

And the original authors are not entitled to the explanation. Economic security is a hypothesis that fits, presented with more confidence than the analysis supported.

Read this as scientific criticism working rather than as a scandal in either direction. A fragile claim was probed, a correction issued, and what is left is smaller and better specified: a real oddity in the raw data with no agreed explanation. Both camps prefer larger conclusions because both are more comfortable than that.

5What “Gender Equality” Actually Measures

The paradox is a relationship between two numbers. Chapters Three and Four examined the first. This chapter examines the second, and the second turns out to be a much stranger object than almost anybody using it realises.

5.1 — Somebody had to decide

When a study says a country is “more gender-equal,” it means the country scored higher on an index. Somebody built that index. They had to decide what counts.

The one used most often in this literature combines four things: women’s economic participation, educational attainment, health and survival, and political empowerment. Each is scored, the scores are combined, and countries are ranked.

Now the crucial design decision, which is stated openly by the people who built it and is almost never noticed by people citing it.

Word Box

A gap measure records the difference between men and women. A level measure records how well women are actually doing.

The main index used in this literature is a gap measure, deliberately. It asks how far women lag behind men in a country, not how good women’s lives are in absolute terms.

Why it matters here: a country where nobody goes to school scores perfectly on educational equality. A country where women live badly and men live equally badly scores well. The index is designed to measure distance between the sexes, and it succeeds — which means it is not measuring what most readers assume when they see the phrase “more gender-equal.”

And that is only the first of the design decisions. The index is also a composite, which brings a second set of choices that are made by somebody and disclosed to almost nobody.

Word Box

A composite index is a single number built by combining several different measurements. Somebody chooses which measurements go in, how each is scored, and how much weight each carries.

Those choices are not technical details. They determine the ranking. Two teams building an index of the same concept from the same data will produce different orderings of countries, and both will be defensible.

Why it matters here: every sentence in this literature beginning "in more gender-equal countries" is a sentence about a composite index, and therefore about somebody’s weighting decisions. That is not a reason to discard the indices. It is a reason to know which one is being used before treating the sentence as a fact about the world.

5.2 — What this produces

The consequences are not hypothetical and they are visible in the published rankings.

Rwanda has for years placed in the top tier of this index, above most of western Europe. The reason is straightforward: it has the highest share of women in a national parliament anywhere in the world, over sixty per cent, and political representation is one of the four components. Whether a Rwandan woman’s life is freer than a German woman’s is a different question, and the index does not ask it.

Several countries in Central America, southern Africa and south-east Asia routinely outrank wealthier ones for similar reasons. This is not an error in the index. It is what a gap measure does, working correctly.

And there is a second design feature worth knowing. The components are capped at parity — a country gets no additional credit for women exceeding men on a measure. So the fact that girls now outperform boys in education across much of the developed world, which Part Four established as one of the larger differences in the data, contributes nothing to a country’s score.

5.3 — Change the ruler, change the answer

Other indices exist and they measure differently. One widely used alternative incorporates maternal mortality, adolescent birth rates and women’s absolute educational attainment — which makes it far more sensitive to how women are actually living rather than how far behind men they are.

Those indices rank countries substantially differently. And the strength of the paradox varies depending on which one you use.

That is a serious problem for the strong version of the claim, and it is worth being precise about why. If a relationship holds with one index and weakens with another, then the finding is not “gender equality predicts smaller female STEM shares.” It is “this particular composite of four indicators, measured as gaps rather than levels, correlates with this particular measure of graduation” — which is a much narrower and much less interesting sentence.

The Hidden Assumption

Everybody in this argument assumes that gender equality is one thing, which countries have more or less of.

Look at what is being bundled. Whether girls attend school. Whether women earn as much as men. Whether women survive childbirth. Whether women sit in parliament. Whether a woman can travel alone, inherit property, refuse a marriage, or leave one.

These do not move together. A country can have high female parliamentary representation and low female literacy. It can have excellent maternal health and no legal right to divorce. It can have equal pay legislation and a marital rape exemption — as Part Three found in a country that combines the two.

Compressing all of it into a single number and ranking countries on it produces something, and what it produces is not “how free are women here.” It is a weighted average of unrelated things, and the weights were chosen by somebody.

The general form: a composite index treated as though it measured a single underlying quantity. Once you see it, the sentence “in more gender-equal countries” stops being informative and starts being a question: equal in what respect, measured how, and weighted by whom?

None of which means the indices are useless. They were built for a purpose — tracking whether countries are closing measurable gaps — and for that purpose they work. The problem is what happens when a tool built for monitoring is borrowed as a variable in a causal argument it was never designed to support.

Remember This

“More gender-equal” means “scored higher on an index,” and somebody built the index. The main one used here combines economic participation, education, health and political representation.

It is deliberately a gap measure, not a level measure. It records how far women lag behind men, not how well women are actually living. A country where nobody is educated scores perfectly on educational equality.

So Rwanda ranks above most of western Europe, driven largely by having over sixty per cent women in parliament. That is the index working correctly, not an error. And no country gets extra credit for women exceeding men — so girls outperforming boys in education across the developed world contributes nothing.

Other indices measure absolute conditions instead, rank countries differently, and the strength of the paradox changes with the ruler you use.

Gender equality is not one thing. Female parliamentary representation, female literacy, maternal survival, the right to divorce and equal pay do not move together. Compressing them into one number produces a weighted average of unrelated things — and the weights were chosen by somebody. “In more gender-equal countries” is not information. It is a question: equal in what respect, measured how, weighted by whom?

6Necessity

The strongest explanation on offer is also the simplest. Where life is precarious, people choose what pays. Where it is not, they choose what they like. It accounts for a great deal — and there is a specific body of evidence it cannot touch.

6.1 — The account

Imagine two young women deciding what to study.

The first lives in a country where graduate unemployment is high, where the state will not support her if she fails, where her family has invested heavily in her education and expects a return, and where a small number of professions offer reliable, respectable, well-paid work. Engineering is one of them.

She may or may not find engineering interesting. That is not the operative question. The operative question is what happens to her if the degree does not lead anywhere, and the answer is: something bad, with no floor under it.

The second lives in a country with a functioning welfare system, a large service economy, many viable careers, and a real prospect of changing course at thirty if the first choice was wrong. If she picks a subject she loves and it pays less, she will be less wealthy and she will be fine.

The first woman optimises for security. The second can afford to optimise for interest. And Part Four established that when people optimise for interest, men and women diverge — because that is what the Things–People difference means.

That is the necessity account in full. It requires no biological claim at all, and it predicts exactly what is observed.

6.2 — What supports it

Several things, and they are not trivial.

The relationship with wealth is at least as strong as the relationship with gender equality. The preference study in Chapter Two found both predicted independently, but wealth was doing real work. Necessity is a story about wealth.

It explains why the effect appears in subject choice rather than in performance. The 2018 study found girls doing as well as or better than boys in science almost everywhere. Necessity predicts precisely this: the constraint is not on capability but on what a person can afford to do with it.

And the countries at the top of the list fit the mechanism. In several of them, engineering and medicine are heavily subsidised, socially prestigious, and among the few professional routes considered appropriate for a woman. That last clause is important and it is not a story about freedom. A woman may be entering engineering because the alternatives are closed rather than because the field is open.

One country demonstrates that sentence more clearly than any statistic can.

In Real Terms

Iran is the case that makes the mechanism concrete, and it is uncomfortable for everybody.

Iranian women live under extensive legal restriction — on dress, on travel, on marriage, on custody, on testimony. By any level measure of women’s freedom, the country sits near the bottom of the range.

Iranian women are also a very large share of university science students, in some fields a majority, and have been for years.

Both sentences are true at once. Read one way, it demolishes the idea that legal freedom drives educational participation. Read another, it shows exactly what Chapter Six is describing: where a small number of routes are open and respectable, and where failure has no floor beneath it, people take the routes that are open.

A high figure in a column marked "women in science" can mean a society opened a door. It can also mean the society closed most of the others.

6.3 — What it cannot explain

The Argument — is the paradox explained by economic necessity?

This is the leading explanation and it is not sufficient. Where it fails is more informative than where it succeeds.

Necessity explains it

The whole pattern falls out of one mechanism: precarity forces optimisation for security, security permits optimisation for interest. It needs no claim about innate anything, it predicts the observed dissociation between performance and choice, and it fits the countries at both ends of the list. It is also the only explanation that is straightforwardly testable, because economic precarity can be measured directly rather than inferred from an index.

Necessity cannot be the whole story

Three problems. It is a theory about career choice, and the same pattern appears in patience, risk tolerance and personality traits, which are not career choices and have no economic return. Post-communist countries in Europe have relatively high female engineering shares while being neither poor nor precarious, which points at history rather than necessity. And the preference study found gender equality predicting the gap independently of wealth, which necessity does not predict.

A large piece of a bigger thing

Necessity very likely accounts for a substantial share of the educational and occupational pattern, and none of the personality and preference pattern. Which suggests the paradox is not one phenomenon with one cause but at least two phenomena that have been given one name because they point the same way.

Where things stand: the third position, and it is the most useful thing in this chapter. Treating “the paradox” as a single object is probably an error. The graduation pattern and the personality pattern may have entirely different causes and have been fused by the fact that both embarrass the same assumption.

What would settle it: testing whether the educational pattern tracks economic precarity better than it tracks gender equality, holding both. This is doable with existing data and has not been done cleanly.

Why people care so much: because necessity is the explanation that makes the paradox politically harmless. If it is all economics, nothing follows about men and women at all — and that is a conclusion one camp badly wants and the other badly does not, which is a reason to check it rather than to adopt it.

6.4 — The uncomfortable version

There is a reading of the necessity account that neither camp raises and that follows directly from it.

If women in poorer countries are entering engineering because their alternatives are constrained — because the field is one of a small number considered respectable, because the family needs the return, because failure has no floor — then the high female STEM share in those countries is not a success. It is a measurement of constraint, appearing in a column that everybody reads as achievement.

And the corresponding reading at the other end is equally uncomfortable. A low female STEM share in a rich country might indicate that women there are choosing freely, which would make it a success measured as a failure.

Both of those are conclusions people reach for when it suits them. What I want to point out is that they are the same argument, and almost nobody holds both. The person who says the Algerian figure shows what women choose when free of Western feminism cannot also say the Norwegian figure shows what women choose when free. And the person who says the Norwegian figure proves persistent bias cannot also say the Algerian figure proves that barriers are the only obstacle.

Chapter Nine puts India into exactly this vice, and India’s numbers do something neither camp expects.

Remember This

The necessity account: where life is precarious, people choose what pays; where it is not, they can afford to choose what they like. And Part Four established that when people optimise for interest, men and women diverge.

It requires no biological claim, predicts exactly what is observed, and explains the key dissociation — girls perform as well or better in science almost everywhere, and still choose differently. The constraint is not on capability but on what a person can afford to do with it.

It cannot be the whole story. The same pattern appears in patience, risk tolerance and personality, which are not career choices. Post-communist Europe has high female engineering shares while being neither poor nor precarious. And gender equality predicted the gap independently of wealth.

Which suggests “the paradox” is not one thing. The graduation pattern and the personality pattern may have different causes, fused into one name because both embarrass the same assumption.

And the uncomfortable reading: if women enter engineering in poor countries because their alternatives are closed, then a high female STEM share is a measurement of constraint appearing in a column everybody reads as achievement. Both camps reach for that argument when it suits them, and neither notices it is the same argument at both ends of the list.

7Freedom, And What Freedom Contains

The second explanation is that when people are secure, they express themselves — and expression is where men and women diverge. It is a good account. It also rests on an assumption about rich societies that is straightforwardly false.

7.1 — The self-expression account

There is a well-established finding in the study of national values: as countries become richer and more secure, what people say matters to them shifts. Concerns about survival, order and material security give way to concerns about autonomy, self-realisation and personal expression.

Apply that here and the account writes itself. In a society organised around survival, a person’s choices are dominated by external requirements. In one organised around self-expression, choices are dominated by internal preference. If men and women differ in internal preference — and Part Four found the largest measured psychological difference is exactly that — then a shift towards self-expression will widen the observed gap.

Notice how neatly this sits with Chapter Six. Necessity explains why the gap is small in poor countries. Self-expression explains why it is large in rich ones. They are the same account viewed from opposite ends, and together they cover the whole range.

It is a genuinely good explanation. It is also incomplete in a specific way that is almost never raised.

7.2 — The thing everybody assumes

The Hidden Assumption

Every version of this argument, on both sides, assumes that a freer society is a society with less socialisation in it.

The picture is of culture as a weight. Remove the weight and what was underneath springs up. So a rich, liberal, gender-equal country is understood as a place where the pressure has been lifted and people are closer to their unshaped selves.

Consider whether that is true of anything else about rich countries.

A rich country has more advertising, not less. More media, watched for more hours. More consumer categories, more finely segmented. More elaborate identity vocabularies, more products attached to each, more sorting of people into marketable groups. A child in a wealthy country encounters vastly more designed, commercially produced messaging about who they are than a child in a poor one.

And gender is the oldest and most profitable segmentation variable there is.

So the picture is backwards. A rich society is not a society with the cultural pressure removed. It is a society with far more cultural production per person, most of it made by people whose job is to sell things, using categories that make selling easier.

The general form: mistaking a change in the kind of pressure for a reduction in the amount. “Freedom” in these indices means freedom from legal and familial constraint. It does not mean, and has never meant, freedom from being marketed to.

7.3 — The evidence that this is not a debating point

How We Actually Know This

Somebody went and counted. A researcher analysed a century of American mail-order catalogues, coding how toys were marketed — explicitly for boys, explicitly for girls, or without a gender attached.

What she found: explicit gendering of toy marketing was at its lowest in the 1970s. It rose sharply through the following decades. By the mid-1990s, the share of toys marketed without a gender attached had collapsed to close to nothing.

The same period covers the largest expansion of women’s legal rights, workforce participation and educational attainment in that country’s history.

What it shows: that legal and economic liberalisation and the intensification of gendered marketing happened at the same time, in the same country. They are not opposites. One did not lift as the other fell.

What it cannot show: that the marketing caused anything. It is a record of what was sold, not of what children became. And it is one country’s catalogues, which is a narrow window on a broad claim.

A related fact is worth knowing because it is so easily checked. The convention that pink is for girls and blue for boys is not old. In the early twentieth century the association was frequently reversed in Western retail advice, and the modern rule settled into place only in the middle of the century. A convention that half the world now treats as almost natural is younger than the aeroplane.

7.4 — What this does to the explanation

It does not destroy it. Self-expression may still be doing real work, and the necessity account in Chapter Six is largely untouched by any of this.

What it destroys is the inference people draw from it — that a difference which grows as constraints lift must have been there underneath all along.

Because there is a second account that fits the same data exactly. As a society gets richer, it does not stop shaping people. It shapes them differently, more intensively, and along commercially useful lines — of which gender is the most useful there is. On that account, the gap grows in rich countries not because something was released but because something else was applied.

Both accounts predict the identical observation. Which means the observation cannot distinguish them, and anybody claiming it does is not reasoning from the data.

Remember This

The self-expression account: as countries get secure, values shift from survival towards autonomy and self-realisation. If men and women differ in internal preference — and Part Four found that is the largest measured difference — then a shift towards expression widens the gap. Combined with Chapter Six, the two cover the whole range.

But every version assumes a freer society has less socialisation in it — culture as a weight that lifts, letting what was underneath spring up.

That is backwards. A rich country has more advertising, more media, more consumer categories, more elaborate identity vocabularies and more products attached to each. And gender is the oldest and most profitable segmentation variable there is.

Somebody counted. Explicit gendering of toy marketing in American catalogues was at its lowest in the 1970s and rose sharply afterwards, until by the mid-1990s almost nothing was sold gender-neutral — across exactly the decades of the greatest expansion of women’s rights in that country’s history. And the pink-and-blue rule is younger than the aeroplane.

So two accounts fit the same data exactly. A difference grows in rich countries because constraint was released — or because a different and more intensive shaping was applied. The observation cannot tell them apart, and anybody claiming it does is not reasoning from the data.

8What The Paradox Cannot Show

Suppose every criticism in Chapters Three to Five fails and the finding is exactly as reported. What then follows? Considerably less than anybody claims, and this chapter sets out precisely how much.

8.1 — Four accounts, one observation

Assume the pattern is completely real. Here are the explanations that fit it, all four of which have been described in this part.

One: release. The differences were always present and constraint suppressed them. Removing constraint lets them appear. This is the account that gets quoted.

Two: necessity. Precarity forces optimisation for security; security permits optimisation for interest. Chapter Six.

Three: elaboration. Rich societies do not shape people less, they shape them more and along more commercially useful lines. Chapter Seven.

Four: measurement. Gap constructions, index composition and the reference group effect generate part or all of the pattern without anything underlying it. Chapters Two, Three and Five.

All four predict the same observation. That is the whole difficulty, and it is not a difficulty that more of the same data can solve — a hundred more countries measured the same way would fit all four just as well.

8.2 — What would tell them apart

Three things would, and it is worth knowing which have been done.

Following countries over time rather than comparing them at one moment. The four accounts make different predictions here. Release predicts the gap grows and then stabilises once constraint is gone. Elaboration predicts it keeps moving and could reverse if marketing conventions change. Necessity predicts it tracks economic security rather than legal equality.

And there is one case where this has effectively been run. It is the most useful piece of evidence in this part and it almost never appears in arguments about the paradox, on either side.

How We Actually Know This

American universities have recorded the sex of graduates by subject for decades. So it is possible to watch a single country, with one set of laws and one culture, over fifty years.

What happened in most fields is what you would expect. Women’s share of medical degrees rose from under a tenth around 1970 to roughly half by the 2010s. Law did something similar. Biology similar. Steady, large, sustained increases across the period of expanding legal equality.

What happened in computer science is not. Women’s share of computing degrees rose through the 1970s and peaked in the mid-1980s at around thirty-seven per cent. Then it fell. By the late 2000s it was around eighteen per cent — less than half its peak — with a modest recovery since.

What this shows: that a single field can move sharply in the opposite direction to every other field, in the same country, in the same decades, under the same laws, during the same expansion of women’s rights. Whatever caused it was specific to computing and was not a change in constraint, because constraint was falling throughout.

What it cannot show: what did cause it. The timing coincides with home computers arriving in households and being marketed heavily as a boys’ product, and with computing courses beginning to assume prior experience that boys were more likely to have. That is a plausible story and a correlation in time, not a demonstration.

Now score the four accounts against it.

Release does badly. If freedom lets an underlying preference emerge, the emergence should not reverse, and it should not run opposite to every neighbouring field in the same country in the same years.

Necessity does badly. Computing became more lucrative and more secure across exactly the period when women’s share of it collapsed. If security drives choice, this went the wrong way.

Elaboration does well. A field acquires a gendered marketing identity, the identity propagates through households and schools, and participation follows. That is precisely what the account predicts and precisely what the timing shows.

Measurement does not apply. These are degree counts.

One case is one case, and I am not going to build a conclusion on it. But it is the single most informative data point in this part, it points away from the account that gets quoted, and the fact that it is rarely raised by anybody arguing about the paradox is worth noticing on its own terms.

Using measures that are not self-report. Chapter Two’s split applies: administrative counts and behaviourally validated measures are immune to the reference group objection. If the pattern holds in the immune measures and fails in the vulnerable ones, that is highly informative. Some of this has been done and the answer is mixed, which is itself worth knowing.

Following migrants. Part Two used exactly this design for the plough hypothesis, and it worked. If women who move from a low-equality country to a high-equality one shift towards the destination pattern within a generation, that rules out anything wholly fixed. As far as I know this has not been done for the paradox, and it is the single most obvious missing study in this part.

The Hidden Assumption

Underneath the entire dispute sits a premise that neither camp examines: that there is a fact about what women would choose if free.

Both sides talk this way constantly. One says the Norwegian figures show what women choose when unconstrained. The other says the Norwegian figures are contaminated by residual bias and the true preference has not yet been revealed. Both assume there is a true preference underneath, waiting to be uncovered by removing enough interference.

But there is no view from nowhere. Every human being who has ever chosen anything did so inside a society, with a particular set of options, a particular vocabulary for describing what they wanted, and a particular set of people watching. There is no condition called “free of culture” in which a preference could be read off cleanly — not in Oslo, not in Algiers, not anywhere.

Which means the question “what do women really want, absent social influence” is not a hard question. It is a question with no possible answer, because the state it refers to has never existed and could not exist.

The general form: demanding a measurement of something in the absence of conditions, when the thing only occurs inside conditions. It is the same error as asking what a language sounds like before grammar.

8.3 — What this does to the rules

Every part of this series has ended by asking what its findings do to the rules in Parts Two and Three. This one has to as well, and the answer is sharper than it looks.

Suppose the strongest version of the paradox is true. Suppose men and women, given freedom, reliably diverge in what they choose. What follows for a rule requiring a woman to marry inside her caste, or forbidding a widow to remarry, or dictating what a woman may wear at a protest?

Nothing — and worse than nothing, because the finding argues against them.

Follow the logic. The paradox says that when women are free, they choose differently from men. That is a claim about what happens in the absence of compulsion. It is evidence, if anything, that compulsion is unnecessary.

A rule exists to make somebody do what they would not otherwise do. If women reliably choose a particular way when nothing forces them, then a rule forcing them that way is not producing the outcome — the outcome was arriving anyway. The rule is producing something else, and the only candidate for what it is producing is the outcome in the cases where the woman would have chosen otherwise.

Which means the entire practical effect of every rule in this series falls on exactly the women whose preferences the paradox does not describe. Not on the average. On the exceptions.

Part Four, Chapter Nine reached this from a different direction and it is worth stating in one line: a preference argument cannot justify a prohibition, and the more strongly you believe the preference finding, the less work there is for the prohibition to do.

The traditionalist reader who has been enjoying this part should sit with that. The paradox is the best empirical card in the traditionalist hand, and it argues for leaving women alone.

8.4 — What the finding licenses

The Argument — what follows from the paradox?

Whatever its causes, people draw conclusions from it. Here is what each is entitled to.

It shows intervention is futile

Decades of programmes aimed at increasing women’s participation in technical fields have coincided with gaps that are stable or widening in exactly the countries running those programmes. At some point a policy that does not produce its intended effect, in its most favourable environment, over fifty years, has to be assessed on results.

It shows nothing about intervention at all

The finding is a correlation between two national aggregates. It contains no information about whether any particular programme works, because it does not compare countries that ran programmes with countries that did not. And it is entirely compatible with there being large numbers of individual women deterred by treatment they encountered, whose absence is invisible in an aggregate.

What it actually licenses

One conclusion, and it is narrow: the assumption that removing barriers straightforwardly produces proportional representation is false. That assumption underlies a great deal of policy and public argument, and the finding is sufficient to retire it. Nothing further follows — not that the remaining gap is natural, not that it is unjust, and not that any specific intervention does or does not work.

Where things stand: the third position. And the first position has a problem it never acknowledges: it treats an aggregate as evidence about individuals, which Part Four, Chapter One disposed of at length. Even under the most favourable reading of the paradox, an individual woman turned away from a laboratory is turned away, and no national correlation speaks to her case.

What would settle it: evaluations of specific programmes against specific outcomes, which exist, are mixed, and are almost never cited by anybody quoting the paradox in either direction.

Why people care so much: because this is the most quotable finding in the whole field, and quotable findings do political work regardless of what they establish. Chapter Ten states plainly how much weight it can bear, which is less than the weight I have been putting on it.

That last sentence is not modesty. Chapter Nine puts India into this framework and finds that the framework is looking in the wrong place, which is a harder criticism of this part than anything in Chapters Three to Five.

Remember This

Assume the pattern is entirely real. Four accounts fit it equally well: release (it was always there, constraint hid it), necessity, elaboration (rich societies shape people more, not less), and measurement artefact. More data of the same kind cannot separate them.

Three things would. Following countries over time — the accounts predict different trajectories. Using non-self-report measures. And following migrants, which worked for the plough hypothesis in Part Two and is the most obvious missing study in this part.

Both camps assume there is a fact about what women would choose if free. There is not. Every person who ever chose anything did so inside a society, with a particular option set and a particular vocabulary for wanting. There is no condition called “free of culture” — not in Oslo, not in Algiers. The question has no possible answer, like asking what a language sounded like before grammar.

What the finding licenses is one narrow conclusion: the assumption that removing barriers straightforwardly produces proportional representation is false. That assumption underlies a great deal of policy and the finding retires it. Nothing further follows — not that the remaining gap is natural, not that it is unjust, and not that any particular programme works or fails.

9Where India Sits

India produces a higher share of female science graduates than Britain, Germany, France or the United States. It also has among the lowest female workforce participation in the world. Both are true, and the space between them is the most important number in this part.

9.1 — The figure nobody expects

India’s national higher education survey records that women make up somewhere around forty-three per cent of graduates in science, technology, engineering and mathematics.

Set that against the comparison countries. The United States is around thirty-four per cent. The United Kingdom around thirty-one. Germany around twenty-seven. France around thirty-two.

So India — which sits in the bottom quarter of most gender-equality rankings, which has the machine described in Part Three still running, which has a marital rape exemption still in force — produces proportionally more female scientists and engineers than any of them.

This is the paradox, in the country this series is about, in its own official statistics. It is not an artefact of anybody’s index. It is a count of graduates.

9.2 — And then they disappear

Now follow those graduates forward.

In Real Terms

Take a hundred Indian women who graduate in a science subject this year, and watch what happens.

Roughly forty-three of every hundred science graduates are women. That is the entry figure and it beats every large Western economy.

Of India’s women STEM graduates, an estimated twenty-seven per cent go on to enter the formal STEM workforce.

Among science faculty in Indian institutions, women were around fourteen per cent.

Among researchers in India generally, the share of women has been reported at under fifteen per cent — one of the lower figures in the world, in a country producing one of the highest shares of female graduates.

Line those numbers up and the shape is unmistakable. India is not failing to educate women in science. It is failing to employ them. The loss does not happen in the classroom. It happens between the graduation ceremony and the first decade of a career, and by the time you look at who is running a laboratory, almost all of it has already happened.

Researchers call this a leaky pipeline, which is a comfortable phrase for something that is not a leak. A leak is gradual and diffuse. This is a fall of roughly sixteen percentage points at one identifiable transition, and Part Three has already explained what happens to Indian women at that stage of life.

9.3 — And inside science, the sorting

There is a second Indian finding, and it is the one that connects this part back to Part Four.

The forty-three per cent is an average across very different subjects. Break it apart and the picture changes completely.

FieldApproximate share of women
Microbiologyaround 67 per cent
Life sciencesaround 56 per cent
All STEM, combinedaround 43 per cent
Undergraduate engineeringaround 29 per cent
Mechanical engineeringunder 7 per cent

Read that table against Part Four, Chapter Five and it is the Things–People dimension, in Indian administrative data, with nobody having asked a single person about their preferences.

Living systems at one end. Machines at the other. Two in three at the top of the table, one in fifteen at the bottom. And a national average of forty-three per cent that conceals both.

So India simultaneously produces the paradox pattern in aggregate and the interest-sorting pattern within. Whatever is going on, it is not that Indian women are indifferent to the distinction that Part Four measured. It is that more of them are getting into the building.

9.4 — The wider number

The STEM figures are a special case of something larger, and the larger thing has to be on the table before the explanation makes sense.

India’s female labour force participation — the share of working-age women in paid work or looking for it — has been among the lowest in the world for a country at its income level, and for a long stretch it was falling while incomes rose. Recent official surveys report a substantial increase, and there is a genuine argument among economists about how much of that increase reflects women entering work and how much reflects a change in how unpaid work on family farms and in family enterprises is counted. I am not going to adjudicate it. Part Twelve takes the measurement dispute properly.

What is not disputed is the shape. Enormous gains in girls’ education across three decades, matched by nothing like a corresponding movement into paid employment.

In Real Terms

Set the two curves beside each other and the mismatch is the whole Indian story.

Girls’ school enrolment, literacy and university entry all rose steeply across the last three decades. On several measures girls now outnumber or outperform boys in Indian higher education.

Over much of that same period, the share of Indian women in paid work went sideways or down.

A country does not usually educate a generation of women and then not employ them. Doing both at once requires an explanation, and the explanation cannot be that the women are unqualified, because the qualifications are the part that worked.

9.5 — What the degree is actually for

Put the two findings together and a specific explanation becomes hard to avoid.

Chapter Six proposed that in precarious economies a technical degree is chosen for security rather than interest. India adds a second function that the necessity account does not cover, and Part Three supplied it.

In a marriage market where families compete for grooms, an educated daughter is a more attractive proposition. A science degree from a recognised institution raises a family’s standing, improves the match, and can reduce what has to be paid. Part Two established that dowry is a bid; a degree is an asset that improves the bid.

On that reading, a substantial share of Indian women’s STEM education is not primarily an input to a career. It is a credential in a different market.

And if that is right, the sixteen-point fall in §9.2 stops being a mystery and stops being a leak. It is what happens when the thing the degree was acquired for has been obtained. Part Three’s finding applies directly: as a family’s income rises, its women leave paid work, because withdrawal is what respectability costs and a household with a surplus can buy it.

I want to be careful here. This is a hypothesis that fits, not a demonstrated mechanism, and it certainly does not describe every Indian woman who studies science. Many intend careers and are prevented — by workplaces, by transport, by safety, by childcare that does not exist, by families that permitted the degree and not the job. Those are different causes producing the same statistic and they are not distinguishable in the aggregate data.

What is established is the shape: education without employment, at a scale unmatched in any comparable country.

It is worth listing the competing mechanisms explicitly, because they call for entirely different responses and are constantly conflated.

If the degree is a marriage-market credential, then the fall after graduation is not a failure of the labour market at all, and every intervention aimed at workplaces will miss. The lever is in the household.

If women want the job and cannot safely reach it — transport, hours, harassment, the cost of a commute in a city with poor public safety — then the lever is infrastructure and policing, and it has nothing to do with attitudes.

If women leave within a few years of starting, at the point of marriage or a first child, then the lever is childcare, parental leave and the distribution of domestic work, and Part Thirteen shows what happens in countries that pulled it.

And if families permit the education and forbid the employment — which Part Three predicts directly, since a working daughter-in-law costs a household respectability that it can afford to buy — then the lever is the one this whole series has been circling, and no policy touches it.

These are not rival theories to be settled by argument. They are four different populations of women, all real, all present, producing one aggregate number. Which is the largest is an empirical question that Indian data could answer and has not.

The Hidden Assumption

Time to turn this on the part, and on the decision to write it.

I deferred to Part Five three separate times in Part Four. I built a whole part around a cross-country correlation. And the Indian numbers in this chapter suggest that the paradox — the thing both camps fight hardest over — is not where the Indian question is.

India’s educational figure is already better than the West’s. If the paradox were the operative issue, India would be a success story. The operative issue is what happens after graduation, and nothing in the international paradox literature addresses it, because that literature measures graduates.

So the assumption underneath this part is: that the question worth arguing about is the one being argued about. Both camps have agreed that cross-country STEM shares are the battleground. I accepted that framing and spent sixty pages on it, and the country I am writing about turns out to have its problem somewhere the framing does not look.

The general form: inheriting a research question along with a research literature. A field that measures graduation will produce arguments about graduation, and everybody downstream will argue about graduation, including people whose actual problem is employment.

I am not going to remove the part. The paradox is quoted at Indians constantly and somebody has to have read the criticisms. But the chapter you have just read is the one that matters here, and it is the ninth of ten rather than the first of sixteen, which is a comment on how I built this.

Remember This

India produces around forty-three per cent female STEM graduates — ahead of the United States (34), Britain (31), France (32) and Germany (27). A country in the bottom quarter of gender-equality rankings, with the machine of Part Three still running. This is the paradox in India’s own official counts.

Then follow them. About twenty-seven per cent of women STEM graduates enter the formal STEM workforce. Women are around fourteen per cent of science faculty and under fifteen per cent of researchers. India is not failing to educate women in science. It is failing to employ them.

And inside science, Part Four’s sorting is fully visible: microbiology around 67 per cent women, mechanical engineering under 7. Living systems at one end, machines at the other, with an average of 43 concealing both.

A hypothesis that fits: in a marriage market where a degree improves the match and reduces the bid, a science education is partly a credential in a different market. Which makes the fall after graduation not a leak but a completion — and Part Three’s finding applies, that women leave work as families can afford the respectability.

And the turn on this part: I deferred to Part Five three times and built sixty pages on a cross-country correlation. India’s educational figure already beats the West. The Indian problem is at employment, and the entire paradox literature measures graduates. I inherited a question along with a literature.

10An Honest List Of What We Do Not Know

This is the least conclusive part of the series, so this chapter matters more here than anywhere else. Two lists, and the first one is long.

10.1 — Genuinely unknown

Almost everything about causation in this part is unknown, and that is not a hedge — it is the state of the field.

Why the pattern exists. Four accounts fit the data equally well and no available evidence separates them. Release, necessity, elaboration and measurement artefact all predict the same observation. Anybody telling you which one is correct is expressing a preference.

How much of the personality and preference pattern is the reference group effect. The mechanism is documented and could generate a substantial share of it. Nobody has measured how much. This is the most important unanswered question about the self-report half of the evidence.

Whether the pattern holds within countries over time. Almost everything here compares countries at one moment. The four accounts predict different trajectories, and outside the computing case this is the cheapest available test that has not been properly run.

What caused the computing reversal. The timing coincides with home computers being marketed as a boys’ product and with courses beginning to assume prior experience. That is a correlation in time and a plausible story. Nobody has established it, and it is the most consequential unexplained event in this literature — because it is the one case where a field went backwards and we have the records to study it.

What happens to migrants. Part Two’s plough study used exactly this design and it was decisive. Nobody has done it here. It is the single most obvious missing study in this part and I do not know why it has not been done.

Whether “the paradox” is one phenomenon. Chapter Six raised the possibility that the graduation pattern and the personality pattern have different causes and were fused because both embarrass the same assumption. If so, arguing about them together has been an error throughout.

What is actually happening to Indian women between graduation and employment. The shape is established. The mechanism is not. Marriage-market credentialism, workplace conditions, safety, transport, absent childcare, and family permission for education but not employment would all produce the same statistic, and the aggregate data cannot separate them.

10.2 — Solid

That the raw country ordering runs against expectation. Several North African and Gulf states produce a higher share of female STEM graduates than the Netherlands, Norway or the United States. These are graduation counts requiring no derived measure, and no participant in the dispute contests them.

That the Nordic countries are among the most occupationally segregated economies in the developed world. Very high female employment, concentrated in a large public care and education sector, alongside a heavily male technical and industrial sector. This falls out of ordinary labour statistics and requires no derived measure.

That women’s share of American computing degrees peaked in the mid-1980s at around thirty-seven per cent and fell to roughly eighteen per cent by the late 2000s, while medicine, law and biology rose steeply across the same decades in the same country. One field moved sharply opposite to its neighbours under identical laws. This is the strongest within-country evidence in the part and it points away from the account that gets quoted.

That the same directional pattern appears in four separate literatures. Education, personality across fifty-five nations, economic preferences across seventy-six, and vocational interests across more than fifty. Different teams, instruments and decades.

That how you express a gap changes the answer. Chapter Three’s worked example is arithmetic. Share of graduates and conditional propensity can point in opposite directions from identical numbers, and they encode different assumptions about what equality means.

That the leading index is a gap measure, not a level measure. Stated openly by its authors. A country where women live badly and men live equally badly scores well on it, and no credit is given for women exceeding men.

That the 2018 study required a correction and that its headline measure is contested. Both are matters of public record. So is the fact that the correction clarified a computation rather than withdrawing a result.

That gendered marketing intensified across exactly the decades of greatest legal liberalisation, at least in the one country where somebody counted. Liberalisation and intensified gendering are not opposites and did not trade off.

That "more women working" and "men and women doing the same jobs" are different goals that can move in opposite directions. Several countries have succeeded enormously at the first while going backwards on the second. Anybody using the word progress here has to say which they mean.

That India produces around forty-three per cent female STEM graduates and employs a fraction of them. Both figures come from Indian official sources and international compilations, and the gap between them is the largest such discrepancy among major economies.

And that the finding retires one assumption. Removing legal and economic barriers does not straightforwardly produce proportional representation. That is established, it is enough to change how a great deal of policy should be argued for, and it is all that is established.

10.3 — Whether this should have been load-bearing

A closing observation, and it is a criticism of my own decision.

This is the most quoted finding in the entire field of sex differences. It is also, as this part has shown, contested in its measurement, ambiguous in its indices, unexplained in its mechanism, and — by Chapter Nine — probably not addressing the question that matters most for the country this series is about.

Those two facts are related. It is quoted so much precisely because it is unresolved. A finding that settled something would be used once and filed. A finding that can be read four ways can be deployed indefinitely by everybody, which is a description of a very useful object and a poor description of knowledge.

Part One, Chapter Six said to notice who is counting and what they get. Applied here: the paradox is valuable to a great many people in its current unresolved state, and the studies that would resolve it — following countries over time, following migrants, separating self-report from administrative measures — are cheap, obvious and have not been done.

I do not think that is a conspiracy. I think it is what happens when a finding is more useful as ammunition than as evidence, and nobody has an incentive to fire the shot that ends the argument.

Remember This

Genuinely unknown: why the pattern exists — four accounts fit equally and nothing separates them. How much of the self-report half is the reference group effect. Whether it holds within countries over time. What happens to migrants. Whether “the paradox” is even one phenomenon. And what is actually happening to Indian women between graduation and employment.

Solid: the raw country ordering runs against expectation, in counts nobody disputes. The same direction appears in four separate literatures. How you express a gap changes the answer. The leading index measures gaps rather than levels. The 2018 study required a correction and its headline measure is contested. Gendered marketing intensified across the decades of greatest liberalisation. India produces around 43 per cent female STEM graduates and employs a fraction of them.

And one assumption is retired: removing barriers does not straightforwardly produce proportional representation. That is enough to change how policy should be argued for, and it is all that is established.

Last observation, against myself. This is the most quoted finding in the field precisely because it is unresolved — a result that settled something would be used once and filed, while a result readable four ways can be deployed indefinitely by everybody. The studies that would settle it are cheap and obvious and nobody has run them.

Four Accounts, One Observation

This part has no chronology to offer either. Its reference table is the state of the argument: the four explanations that fit the data, what each predicts, and what would tell them apart.

AccountThe claimWhat would support itIts main problem
ReleaseDifferences were always present; constraint suppressed them; freedom lets them appear.Gaps grow and then stabilise once legal constraint is gone. Migrants’ children shift to the destination pattern only slowly.It is not the only account that predicts a growing gap, and it is treated as though it were.
NecessityPrecarity forces optimisation for security; security permits optimisation for interest.The gap tracks economic risk more closely than it tracks legal equality. Sudden welfare expansion widens it.It is a theory about careers, and the same pattern appears in patience, risk and personality.
ElaborationRich societies shape people more, not less, along commercially useful lines. Gender is the most useful line there is.Gendered marketing intensity tracks the gap. Conventions change and the gap follows.Documented for marketing content; the link to what children become is not demonstrated.
MeasurementGap constructions, index composition and the reference group effect produce part or all of the pattern.The pattern fails in administrative counts and behavioural measures while holding in self-report.It cannot account for the graduation counts, which carry none of these problems.

Three tests appear in that table more than once, and none of them has been done properly.

Follow countries over time rather than comparing them at one moment. The accounts predict different trajectories and the data largely exists.

Separate self-report from administrative and behavioural measures and check whether the pattern holds in both. Chapter Two showed why this is the sharpest available cut.

Follow migrants. Part Two used exactly this design for the plough hypothesis and it was decisive. Nobody appears to have done it here.

Sources & further reading — Part 5

Glossary

Every hard word used in this part, in plain English. Each was explained where it first appeared.

TermPlain meaning
Conditional propensityOf the people who reached university at all, what fraction chose a given subject. A measure that isolates choice by holding access fixed — and therefore removes access, the biggest inequality in poor countries, from view.
Gap measureAn index recording the distance between men and women rather than how well women are actually living. A country where both sexes fare equally badly scores well.
Gender-equality paradoxThe pattern in which differences between men and women appear larger in richer and more gender-equal countries. Called a paradox because the most straightforward account predicts the opposite.
Level measureAn index recording women’s absolute conditions — maternal survival, literacy, income — rather than the gap between the sexes. Ranks countries very differently from a gap measure.
Reference group effectPeople rating themselves compare themselves to those around them rather than to humanity in general, so self-ratings are not comparable across countries. Applies to self-report and not to counting graduates.
Relative strengthWhat a person is best at compared with their own other abilities, rather than compared with other people. A girl can be an excellent scientist for whom science is a personal weakness.
Self-expression valuesThe finding that as countries become secure, stated priorities shift from survival and order towards autonomy and personal realisation.

Download The complete book · 4.0 MB