Part 04 of 16Biology

The Bodies Themselves

Every measured difference between men and women, with its real size attached. Some are enormous. Most are almost invisible. The largest one is not an ability at all.

Where We Left Off

Before we begin

Three parts of history are now behind us. Part One separated the three claims welded into every defence of these rules — where they came from, what changing them costs, and what you should therefore do — and showed that each needs a different kind of evidence. Part Two found the origin: a field cannot be carried, so it must be inherited, and a man could not verify which children were his. Part Three found the specifically Indian machine, which turned out not to be about paternity at all but about keeping marriage circles closed, and found that a dateable layer of Indian modesty was written in London.

Part Three ended by turning its own argument on itself. Where a rule came from settles nothing about whether it is good. That was the genetic fallacy, named in Part One and then used for four chapters because it wins arguments.

So history has now done what history can do. It has established that these rules are not timeless, that they had functions, that some of the functions are gone and some are not, and that the specific arrangement defended as natural is one option among several that human beings actually built.

What it has not done is answer the question underneath all of it. How different are men and women, actually?

That question has been sitting under every chapter so far. Both camps have an answer they need to be true. One needs the differences to be large, because large differences make the rules look like a reasonable response to reality. The other needs them to be small, because small differences make the rules look like an imposition on people who are basically the same.

This part measures them. Every difference, with its actual size, reported at full strength regardless of which camp it helps. Some of what follows is going to be unwelcome to readers who think of themselves as progressive. Some of it is going to be much more unwelcome to readers who think of themselves as traditional, because the largest measured difference is not an ability and does not support a single one of the rules in Parts Two and Three.

How to Read the Boxes

The notation, in case this is where you started

Six kinds of box run through the series. One live example of each.

Word Box

Distribution: the full spread of a measurement across a group — not just the average, but how many people sit at every value from lowest to highest.

Almost every mistake in this subject comes from talking about averages when the interesting thing is the spread. Two groups can have the same average and produce wildly different numbers of people at the extremes, and Chapter Four is entirely about that.

Why it matters here: whenever you read “men are more X than women,” ask whether the claim is about the middle of the distribution or the edge of it. They are different claims with different consequences.

A Word Box appears the first time a hard word does — never later, never only in a glossary.

In Real Terms

Here is the picture this entire part is built on. Hold it in your head and you can read any claim about sex differences without being fooled.

A hall with one hundred men on one side and one hundred women on the other. Somebody walks in, picks one person at random from each side, and measures whatever is being argued about.

The only question that matters is: how often does the man come out higher?

If the answer is fifty times in a hundred, there is no difference at all. If it is fifty-five, the difference is real and nearly useless for predicting anything about a person. If it is seventy-five, it is large. If it is ninety, it is enormous and visible from the door.

Every number in this part will be translated into that hall.

Any number too big to picture gets a body attached.

How We Actually Know This

This part rests on a different kind of evidence from the previous three: meta-analyses, large national assessments and population databases rather than law reports and genomes.

What that buys: scale. Several findings here rest on hundreds of thousands of people, and the largest on more than half a million. At those sizes, sampling noise stops being the problem.

What it cannot buy: causation, and it does not cure the problems Part One described. A meta-analysis pools existing studies, so it inherits their biases — including publication bias, where results showing a difference are more likely to be written up than results showing none. And a very large sample makes tiny differences statistically detectable, which is why every number in this part is reported as an effect size rather than as whether it reached significance.

The rule I will follow: where a finding rests on one study, I will say so. Where it has been repeated across countries and methods, I will say that too, and I will treat the difference as decisive.

This box names the actual evidence and then says what it cannot show.

The Argument — should differences between men and women be researched at all?

A framing question, and worth settling before ten chapters of measurement.

The case against

Findings in this area are reliably misused. They are quoted out of context to justify exclusion, they reach the public stripped of their effect sizes, and the harm from a misread result vastly exceeds the benefit from a correct one. The field also has a documented history of producing confident conclusions that later collapsed, nearly always in the direction of justifying the arrangements of the day.

The case for

The claims are being made anyway, constantly, by people with no data at all. A vacuum is not neutral — it is filled by assertion, and assertion favours whoever is louder. Refusing to measure also means being unable to answer, and the specific pattern of “men and women are identical” arguments collapsing on contact with evidence has done more damage to the credibility of equality arguments than any finding ever did.

Where things stand: the second position, with the first taken seriously as a warning about method rather than a case for silence. The correct response to a finding being misusable is to report it with its size attached, which is what this part does and what almost no popular coverage does.

What would settle it: nothing empirical. This is a disagreement about risk, and reasonable people land differently.

An Argument box appears where people who have studied something genuinely disagree.

The Hidden Assumption

Both sides above assume that the answer matters for the political question.

One side fears the findings because it believes large differences would justify unequal treatment. The other wants the findings because it believes the same thing in reverse.

Part One, Chapter Four already showed this is a bad bargain: no fact about averages produces a conclusion about anybody’s rights without an added value, and the added value has to be argued separately. Both camps accepted the bargain and are now fighting over the evidence to protect a moral conclusion neither has defended.

The general form: arguing about the premise because you have conceded the inference. Chapter Nine takes this up properly and turns it on this whole part, including on my decision to write it.

This is the signature box. There are five in this part.

Remember This

Every chapter closes with one of these, restating it in the plainest words available, key terms in bold.

Read only these and you should still finish holding the whole argument.

A Note on This Part

Read this before Chapter One

I am not a psychologist, a physiologist or a statistician. I am reading the published literature and reporting it. In previous parts that limitation was manageable because the material was historical. Here it matters more, because the technical questions are genuinely difficult and specialists disagree about several of them. Where they do, I have put the disagreement in an Argument box rather than adjudicating it.

Every number in this part is a group average or a group spread. None of them is a fact about any individual, and Chapter One exists to make that impossible to forget. If at any point a sentence in this part reads as a statement about a person, I have written it badly.

And I have a specific worry about this part that I did not have about the others. The findings here are the ones most commonly quoted, stripped of their sizes, to justify keeping women out of things. I am going to report them anyway, because refusing to look is the moralistic fallacy from Part One and because the claims circulate regardless. But I want to say plainly, before the first number: nothing in this part supports a single rule described in Parts Two and Three. Chapter Nine explains why in detail, and if you only read one chapter, read that one.

1How To Read A Difference

“Men are more X than women” is not a claim. It is the beginning of one. Until you know how big the difference is, the sentence carries no information — and almost everything written about this subject stops before that point.

1.1 — The hall

Put a hundred men on one side of a hall and a hundred women on the other.

Somebody walks in, picks one person at random from each side, and measures whatever is being argued about — height, memory, patience, grip, anything.

The only question worth asking is how often the man comes out higher. That number, and nothing else, tells you what a difference between men and women is worth.

Fifty in a hundred means no difference at all. Fifty-six means a difference that is real and almost useless for predicting anything about a person you have just met. Seventy-one is what researchers call a large difference. Ninety-two is enormous — visible across a room, obvious to a child.

Researchers do not usually report the number that way. They report something called an effect size, and this chapter exists to let you convert one into the other.

Word Box

Effect size: a number expressing how big a difference between two groups is, measured against how spread out those groups are.

The version used almost everywhere in this literature is written as d. A d of zero means the two groups have the same average. A d of 1 means the averages are one full “spread” apart.

The reason spread has to be involved: a difference of five centimetres is enormous for the width of a hand and trivial for the height of a tree. You cannot say whether a gap is big until you know how much things vary anyway.

Why it matters here: d is the single most important number in this part. A headline that reports a difference without it is not reporting a finding; it is reporting that somebody measured something.

1.2 — The conversion table

Here is the table. It is worth returning to, because every number in the next nine chapters will be read through it.

Effect size (d)The man is higher, out of 100 pairsWhat that means in the hall
0.050Nothing. A coin toss.
0.153Detectable with a huge sample. Invisible in a room.
0.256Conventionally “small”. Useless for predicting a person.
0.564Conventionally “medium”. You would notice it in a large group and be wrong about individuals constantly.
0.871Conventionally “large”. Three times in ten the woman still wins.
1.076Very large for a psychological trait. Almost nothing in psychology reaches this.
1.586The territory of physical measurements.
2.092Obvious. Height sits near here.
3.098Nearly categorical. Very few human traits are here.

Look at the row for 0.8 and stay there a moment, because it is the one that does the most damage in public argument.

A “large” difference means that if you pick one person from each group, the woman comes out higher three times in ten. That is not a small residue. In a company of a thousand, it is three hundred people.

So a large sex difference in something does not license a rule about who may do it. Even at the top of this table, using sex to predict an individual means being wrong at a rate no employer, teacher or family would accept in any other context.

The Hidden Assumption

Both camps assume, without ever stating it, that a fact about a group is a fact about its members.

One side hears “men are on average stronger” and treats it as a description of the man in front of them. The other side hears the same sentence, understands that it will be treated that way, and therefore attacks the finding rather than the inference.

Both are working inside the same error. A distribution is not a person. The sentence “men are on average taller” is true and tells you almost nothing about whether the man walking towards you is taller than the woman beside him.

Notice what follows for the argument this series is about. Every rule in Parts Two and Three is applied to individuals — this daughter, this widow, this woman at a protest. A group average cannot justify any of them, not because the average is false but because it is the wrong kind of object. Even a difference at the very bottom of this table would misclassify one woman in ten.

The general form: reasoning from a distribution to an instance. It is the most common error in this subject and it is committed by everybody, including by people who could recite this box.

1.3 — What the headlines do

Three specific moves account for nearly all bad coverage of this subject, and once you can name them you are largely immune.

Reporting a difference without its size. “Study finds men and women differ in X.” True of almost everything, at almost any sample size, because with enough people even a d of 0.03 becomes detectable. The finding is not that a difference exists. The finding is how big it is, and that number is usually in the paper and not in the article.

Converting an average into a category. “Men are the ones who X.” An average difference becomes a statement about kinds, and the enormous overlap disappears. This is the error in the box above, dressed as a headline.

Reporting the exception as the rule, or the rule as the exception. Both directions happen. A single unreplicated study showing a difference gets a headline; so does a single study showing none. Part One’s rule applies to both: one study is not a fact.

Remember This

Put a hundred men and a hundred women in a hall, pick one at random from each side, and ask how often the man comes out higher. That number, and nothing else, tells you what a sex difference is worth.

Effect size (d) is how researchers report it. Convert it: 0.2 means the man wins 56 times in 100. 0.5 means 64. 0.8 — conventionally “large” — means 71, so the woman still wins three times in ten.

Almost nothing in psychology reaches 1.0. Physical measurements live above 1.5. Height sits near 2.0, which is 92 in 100.

A distribution is not a person. The rules in Parts Two and Three are applied to individuals, and a group average is the wrong kind of object to justify any of them — not because it is false, but because even a very large difference misclassifies three women in ten.

“Men and women differ in X” is true of almost everything and means nothing. The finding is always the size, and the size is always in the paper and almost never in the headline.

2The Body

This is where the differences are largest, and there is no honest way to soften them. Physical sex differences are among the biggest measured differences between any two human groups — and the one with the greatest consequence of all runs the other way.

2.1 — Size and composition

Start with the easy one. Men are taller, by something in the region of thirteen or fourteen centimetres on average in most populations. In the language of Chapter One, that is an effect size of roughly two — the man is taller about ninety-two times in a hundred pairs.

Underneath height sit differences that matter more for what a body can do.

Men carry substantially more lean muscle mass and less body fat as a share of weight. The muscle difference is not evenly distributed: it is considerably larger in the upper body than the lower. Men also have larger hearts relative to body size, more circulating haemoglobin — meaning more oxygen delivered per litre of blood — and denser, heavier bones.

Women have, on average, greater joint flexibility, better tolerance of some kinds of sustained low-intensity effort, and different injury patterns rather than uniformly more or fewer injuries.

2.2 — Strength

Strength is where the numbers become extreme, and I am going to report them plainly because rounding them off would be dishonest.

Upper-body strength shows the largest gap. Across studies the male advantage in upper-body strength is commonly in the range of forty to sixty per cent; in the lower body it is smaller, often around twenty-five to thirty per cent, partly because everybody’s legs are trained by walking.

How We Actually Know This

Grip strength is the standard measure, because it is cheap, quick and hard to fake, so it has been recorded on very large numbers of people.

One large study of young adults measured grip in over two thousand people in their early twenties. The grip strength of about ninety per cent of the women fell below that of about ninety-five per cent of the men. The study also included highly trained female athletes; their results overlapped with the untrained men, at the lower end of the male range.

What it shows: a difference so large that the distributions barely overlap. In the language of Chapter One this is above the bottom of the table — it is off it. There is no psychological trait in this entire part that comes close.

What it cannot show: anything about training history, which differs systematically between the sexes and is not fully controlled in most such samples. The athlete comparison partly addresses this and does not eliminate it. And grip is one measure; it is not a summary of physical capability.

2.3 — Sport as an accidental experiment

Competitive sport is the best available natural experiment on physical capacity, because it is the one domain where enormous numbers of people are measured on identical tasks under identical rules with maximum motivation.

The results are consistent. In running, the gap between the best men and the best women is roughly ten to twelve per cent. In swimming it is smaller, around six to ten per cent, because water reduces the advantage of mass. In jumping and throwing it is larger. In events dominated by upper-body power it is larger still.

Those percentages sound modest. At the elite level they are not, because the top of any distribution is crowded.

In Real Terms

In 2017 the United States women’s national football team — then among the strongest sides in the world, holders of a World Cup — played a practice match against the under-fifteen boys’ academy team of a domestic club.

The boys won.

Sit with what that involves. On one side, adult professionals, the best in the world at their sport, with a decade of full-time training each. On the other, boys of fourteen and fifteen who were not among the best of anything and would mostly never play professionally.

A comparable thing happened in tennis in the 1990s, when a man ranked around two hundredth in the world played exhibition sets against each of the two best women in the game and won both comfortably.

This is what a ten per cent difference in the average looks like once you go to the far edge of the distribution. Chapter Four is about why the edge behaves so differently from the middle, and it is the most important statistical idea in this part.

2.4 — Where the gap narrows, and where it closes

Now the part usually left out, and it is not a consolation prize — it is evidence about mechanism.

The gap is not constant across activities. It shrinks as events get longer and as raw power matters less. In ultra-endurance events the difference is markedly smaller than in the marathon, and in some very long swimming events women have won outright against mixed fields. In events requiring extreme cold tolerance, women’s higher body fat is an advantage rather than a disadvantage.

That variation tells you the gap is not a general fact about capability. It is specific to the demands of a task, and it can reverse when the demands change.

Which sets up the correct way to read this chapter. The male physical advantage is real, very large, and narrow. It concerns power, speed and the ability to apply force. That is not a summary of what a human body is for. It is one axis, and it is the axis on which almost all competitive sport happens to be built, which is why sport makes the difference look like the whole story.

2.5 — The difference that kills

There is one more physical difference, it is the largest in consequence of anything in this chapter, and it runs the other way.

Men die younger. In essentially every country on earth, at every income level, under every political system and every religion, female life expectancy exceeds male. The global gap is in the region of five years, and it has been consistently in one direction for as long as reliable records have existed.

It is not a single cause. Male infants die at higher rates than female infants from the first days of life, before any behavioural explanation can apply. Men are more vulnerable to a range of infections. They die more from most major causes of death, at most ages. And they die vastly more from the causes that involve doing something dangerous — which is behaviour, and which Chapter Six takes up.

In Real Terms

Five years is hard to feel as a number, so convert it into a household.

Take a couple married at twenty-five, the same age. On average figures, she will outlive him by around half a decade. Multiply that across a country and you get the composition of every old people’s home, every widowhood statistic, and a substantial part of why elderly poverty is disproportionately female — a woman is much more likely to reach the age at which savings run out, and to reach it alone.

Now set it against the rest of this chapter. The male physical advantages — more muscle, more power, more oxygen-carrying capacity — are real and large. They come attached to a body that fails earlier, in every population ever measured. Whatever the male body is optimised for, it is not lasting.

This belongs in a chapter about physical differences and it is almost never in one, because it does not fit the shape of the argument either camp is having. Part Fifteen counts what it costs in full.

2.6 — And what it does not establish

Two things need saying before this chapter can be used for anything.

Almost no modern work requires the capacities where the gap is largest. The tasks that separate the sexes most sharply are the ones machines took over first. Upper-body power is not what a surgeon, a lawyer, an engineer, a farmer with a tractor or a soldier with modern equipment is primarily selling. Where the requirement is genuine — some emergency and military roles — the honest response is a job-relevant physical standard applied to every applicant, which is a test of the person rather than of the category, and which some women pass and many men fail.

And nothing in this chapter reaches a single rule in Parts Two and Three. Grip strength does not explain why a widow may not remarry, why a woman must marry inside her caste, why she may not go out after dark, or why a photograph of her at a protest was answered with a rumour about her body. The largest measured difference between men and women is irrelevant to every one of them.

That mismatch is worth holding on to for the rest of the part. Whenever somebody reaches for biology to justify one of these rules, ask which measured difference is being invoked and what it has to do with the rule. The answer is usually that the rule was there first and the biology was recruited afterwards.

Remember This

Physical differences are the largest in this part and there is no honest way to soften them. Height sits near an effect size of 2.0 — the man is taller 92 times in 100 pairs.

Strength is more extreme. Upper-body strength differs by roughly forty to sixty per cent. In one large study of young adults, the grip strength of about nine in ten women fell below that of about ninety-five per cent of men — a gap larger than anything psychological in this entire part.

At the elite edge this compounds. A world-champion women’s football side lost a practice match to an under-fifteen boys’ academy team. That is what a ten per cent difference in averages does at the far edge of a distribution.

But the gap is narrow as well as large. It shrinks in ultra-endurance events, closes in some, and reverses where body fat is an advantage. It concerns power and speed — one axis, and the one competitive sport happens to be built on.

And the largest physical difference by consequence runs the other way. Men die around five years earlier, in essentially every country, at every income level — starting with higher infant mortality, before behaviour can explain anything. Whatever the male body is optimised for, it is not lasting.

And it reaches none of the rules. Grip strength does not explain why a widow may not remarry or why a woman at a protest was answered with a rumour about her body. When somebody reaches for biology to justify a rule, ask which difference is being invoked and what it has to do with the rule — usually the rule came first and the biology was recruited afterwards.

3The Mind, Measured

Take every mental ability that has been measured and put each one in the hall from Chapter One. Nearly all of them come out close to a coin toss. The exceptions are real, and one of them runs in a direction almost nobody discusses.

3.1 — The general finding is similarity

In 2005 a researcher reviewed forty-six meta-analyses of psychological sex differences — a review of reviews, covering thousands of studies. She found that around four fifths of the measured differences were small or close to zero.

That is the headline finding of this field and it is almost never reported as one, because a finding of “not much” does not travel.

Word Box

Meta-analysis: a study of studies. Rather than running a new experiment, researchers gather every published study on a question and combine their results into a single estimate.

The strength is scale: a meta-analysis can pool hundreds of thousands of people and drown out the noise that makes any individual study unreliable.

The weakness is that it inherits whatever is wrong with what it pools. Chiefly publication bias — the tendency for studies that found a difference to be written up and printed, while studies that found nothing quietly are not. That bias runs one way, and it means a pooled estimate of a sex difference is more likely to be too large than too small.

Why it matters here: most numbers in this part come from meta-analyses. Read every one of them as a probable slight overstatement rather than a floor.

It should be held firmly, because everything that follows is a list of exceptions, and a list of exceptions read on its own gives a false impression of the whole.

3.2 — The list

Here is what has actually been measured, converted into the hall. The final column is how often a randomly picked man scores higher than a randomly picked woman. Fifty is no difference; below fifty means women score higher.

AbilityEffect sizeMan higher, per 100 pairs
General intelligenceabout 050
Mathematics (school assessments)about 0.0551
Vocabularyabout 050
Verbal ability, generalabout −0.147
Spatial visualisationabout 0.1554
Verbal and episodic memoryabout −0.244
Object location memoryabout −0.2543
Perceptual speedabout −0.2543
Readingabout −0.342
Spatial perceptionabout 0.4562
Writingabout −0.536
Mental rotation of 3D shapes0.5 to 0.964 to 74

Read the middle of that table and the picture is unmistakable. Most of the abilities people argue about most furiously sit within a few points of a coin toss.

3.3 — Mathematics

The maths gap deserves its own section because it is the most confidently asserted difference in this field and the mean difference is, to a very good approximation, nothing.

Large national assessment data covering millions of schoolchildren produces an effect size close to zero. Meta-analyses of the research literature agree. In some countries girls score higher. The average boy and the average girl are, on this measure, the same.

The Argument — is there a sex difference in mathematical ability?

Everybody arguing about this is talking about a different part of the distribution, which is why the argument never ends.

There is essentially none

National assessment data on millions of students shows an average difference indistinguishable from zero. It varies by country, and in several countries it favours girls — which no biological account predicts. Where a gap exists, it tracks measures of the society rather than measures of the brain.

The averages are the wrong place to look

Equal averages are compatible with very unequal extremes, and the extremes are what produce mathematicians. Among American twelve- and thirteen-year-olds scoring at the very top of a difficult mathematics test in the early 1980s, boys outnumbered girls by roughly thirteen to one. That is where the professions come from, not from the middle.

And the extreme ratio moved

The strongest evidence is what happened next. That thirteen-to-one ratio fell over the following two decades to something in the region of four to one, and has been comparatively stable since. Whatever produced the original figure, a large part of it changed within twenty years — which no account resting on fixed biology can accommodate.

Where things stand: the mean difference is settled and is essentially zero. A difference at the very top is real and is discussed in Chapter Four. And the fact that it shrank by roughly two thirds in twenty years is the most informative single number in this section, because it puts a floor under how much of the original was social.

What would settle the remainder: whether the ratio continues to move. It has been roughly stable for some time, which is consistent with either a biological floor or a social one that has stopped shifting. Nobody can currently distinguish those.

Why people care so much: because this specific claim has been used to explain the composition of entire professions, and because a stable ratio is read by one camp as nature and by the other as unfinished business.

3.4 — The gap nobody mentions

Now look at the two rows near the bottom of the table, and notice that they run the other way.

In reading, girls outperform boys. This is not a local finding. In international assessments covering dozens of countries, girls have scored higher in reading in every participating country, in every cycle. There is no country in the dataset where boys read better.

In writing, the gap is larger still, and on some large national datasets it is the biggest sex difference in academic achievement that has been measured — around half an effect size, which puts it in the hall at roughly thirty-six in a hundred.

Put that beside the maths gap of approximately zero and consider how the two are treated. One is a matter of national concern, funded programmes, and decades of public argument. The other is roughly ten times larger, entirely consistent across countries, and most readers of this document are learning about it now.

I want to be careful about what I am claiming. I am not saying the attention to girls in mathematics was wrong; there were real barriers and the gap in participation was never only about scores. I am saying that the pattern of what gets measured, funded and discussed does not track the size of the difference — which is a finding about us rather than about children, and Part One, Chapter Six predicted exactly this.

3.5 — The one large difference

Mental rotation is the exception at the top of the table, and it is worth being precise about what it is.

The task is to look at a three-dimensional shape, then look at another shape, and decide whether the second is the same object rotated or a mirror image. It is a narrow, specific ability, tested in a laboratory, and the male advantage on it is the largest cognitive sex difference reliably found — depending on the version of the test, somewhere between 0.5 and 0.9, which is sixty-four to seventy-four in the hall.

Two honest qualifications sit next to it.

It trains. Performance on mental rotation improves substantially with practice, including with practice on action video games, and the sex gap narrows when both groups are trained. It does not usually vanish, but a difference that moves with a few hours of training is not a fixed property of anything.

And its real-world reach is smaller than it sounds. Mental rotation correlates with some engineering and technical outcomes, modestly. It is not a general spatial faculty, and other spatial abilities in the table show much smaller differences. A single narrow laboratory task is doing a great deal of rhetorical work in public argument that the underlying finding does not support.

Remember This

The headline finding of this whole field is similarity. A review of forty-six meta-analyses found around four fifths of measured psychological sex differences to be small or close to zero. Everything below is a list of exceptions, and exceptions read alone mislead.

Mathematics: the average difference is essentially zero, across millions of students, favouring girls in some countries. A gap at the extreme top was once about thirteen to one — and fell to around four to one within two decades, which no fixed-biology account can accommodate.

Two differences run the other way and are larger. Girls outperform boys in reading in every country measured, every cycle, and the writing gap is around half an effect size — roughly ten times the maths gap and among the largest in academic achievement.

Notice which of those two receives national attention and funding. The pattern of what gets measured and discussed does not track the size of the difference.

The largest cognitive difference is mental rotation — deciding whether a 3D shape is the same object turned around. It is real, it is 64 to 74 in the hall, it improves with practice, and it is one narrow laboratory task carrying far more rhetorical weight in public than the finding supports.

4The Tails

Two groups can have exactly the same average and produce very different numbers of people at the extremes. This is the most consequential idea in the part, it is used badly by everybody, and it cuts in a direction almost nobody expects.

4.1 — Same middle, different edges

Everything so far has been about averages. Averages are the wrong place to look for most of the outcomes people actually argue about.

Word Box

Variability: how spread out a group is. Two groups can share an average while one is tightly bunched around it and the other is stretched out towards both extremes.

The measure used here is the variance ratio — how much wider one group’s spread is than the other’s. A ratio of 1.0 means identical spread. Ratios found for human cognitive traits are typically small, in the region of 1.05 to 1.20, with males the wider group on most measures.

Why it matters here: a spread difference this small sounds negligible and is not. At the far edge of a distribution it multiplies, and the far edge is where professorships, prizes, prisons and diagnoses are.

4.2 — The arithmetic

The best way to see this is to do it, so here it is with numbers.

Take two groups with exactly the same average. Give one group a spread seven per cent wider than the other — a difference far too small to notice in any ordinary setting.

Now count how many people from each group appear above various thresholds.

How far above averageRatio, wider group to narrowerRoughly what this level means
1 spreadabout 1.1 to 1Above average. Nearly balanced.
2 spreadsabout 1.35 to 1Top few per cent. A noticeable tilt.
3 spreadsabout 1.9 to 1Roughly one in a thousand. Nearly two to one.
4 spreadsabout 2.9 to 1Roughly one in thirty thousand. Almost three to one.

Read that again with the first line of §4.2 in mind. The averages are identical. Nothing in this table comes from one group being better at anything. All of it comes from a seven per cent difference in spread.

And if the spread difference is fifteen per cent rather than seven, the ratio at four spreads above average rises to something in the region of eight to one.

This is the single most important statistical fact in the part, and it explains why arguments about averages so consistently fail to explain what people see. You can have complete equality in the middle and a large imbalance at the edge, produced by nothing except a difference in spread.

4.3 — Does it actually happen?

How We Actually Know This

International school assessments test hundreds of thousands of fifteen-year-olds across dozens of countries on identical instruments. They are the best data available for comparing spreads.

What they show: greater male variability on most measures in most countries. In one widely cited analysis of a large international assessment, male variance in mathematics exceeded female variance in the large majority of participating countries, and in reading it did so in all of them.

What it can show: that the pattern is broad, cross-national, and not an artefact of one country’s schooling.

What it cannot show: that it is fixed. The size of the variance ratio differs between countries, sometimes considerably, and has changed over time in some. A biological constant does not vary by nation. Whatever produces this is at least partly responsive to something about societies, and nobody has established how much.

One further caution: these are school populations. Boys are more likely to be excluded, to drop out, or to be absent on test day, which can distort a tail measurement in either direction.

4.4 — The half that is never mentioned

Here is where this chapter turns, and it turns hard.

Greater variability means more people at both extremes. Not one. Both. The same arithmetic that puts more men at the top of a distribution puts more men at the bottom of it, in exactly the same proportion.

And that is what the data shows. Males are substantially over-represented in intellectual disability, in severe learning difficulties, in diagnoses of attention and developmental disorders, in school exclusion, in failure to complete education, and — as Part Fifteen will count in detail — in almost every measure of catastrophic life outcome.

So anybody invoking male variability to explain why the top of a profession is male has, in the same breath and by the same mathematics, explained why the bottom of society is male. The argument does not have a top-only setting.

The Hidden Assumption

Both camps assume that the interesting question is the average.

The entire public argument about sex differences is conducted in the language of averages. Are women as good at X. Are men more Y. Every study reported, every headline written, every dinner-table row.

But almost nothing that people actually fight about is decided at the average. Professorships are not awarded at the fiftieth percentile. Neither are prime ministerships, patents, prison sentences, or diagnoses. The socially consequential outcomes are all at the edges of distributions, and the edges obey different arithmetic from the middle.

This means both camps are frequently arguing correctly about the wrong thing. The reformer who establishes that average maths ability is identical has established something true that does not touch the question of who becomes a mathematician. The traditionalist who points at the top of a field has said nothing about the vast middle where most people live.

The general form: arguing about the middle of a distribution when the decisions happen at its edge. It is why two people can both be right and still disagree completely.

4.5 — What it explains, and what it cannot

The Argument — does male variability explain the composition of elite fields?

This is the most consequential live dispute in the part.

It goes a long way

The arithmetic in §4.2 is not a theory; it is a consequence of the shapes of the distributions. Greater male variability is documented across most countries and most measures. Given that, an imbalance at the extreme top of ability-intensive fields would occur even with perfectly equal averages and perfectly fair treatment. Any explanation that ignores it is incomplete by construction.

It explains far less than claimed

Four problems. The extreme mathematics ratio fell by roughly two thirds in twenty years, and nothing biological moves that fast. Variance ratios differ between countries, so they are responsive to something social. The fields with the largest gaps are not reliably the ones demanding the most extreme ability. And representation gaps exist throughout these professions, not only at the very top — which is the only place this mechanism operates.

Valid mathematics, overloaded

The mechanism is real and contributes something. It cannot account for the size of observed gaps, their timing, or their variation between countries, because it is a constant and those are variables. Anybody presenting it as a sufficient explanation is misusing valid mathematics, and anybody denying it operates at all is denying arithmetic.

Where things stand: the third position. This is a genuine mechanism producing a genuine effect of unknown and probably modest magnitude, routinely inflated into a complete account by people who like the conclusion.

What would settle it: tracking variance ratios and elite representation together across countries and decades. If representation gaps move while variance ratios stay still, the mechanism is not doing the work. Preliminary comparisons point that way and the analysis has not been done properly.

Why people care so much: because it is the last remaining argument that can explain unequal outcomes without anybody having done anything wrong. That is an enormously attractive property, and it is exactly why the claim needs checking rather than adopting.

One practical note before leaving this chapter. Whenever somebody invokes the tails at you, ask which end they are invoking. Almost everybody who reaches for this argument reaches for one end only, and the arithmetic does not have a one-ended setting.

Remember This

Two groups with identical averages can produce very different numbers at the extremes, if one is more spread out. Male variability is slightly greater on most cognitive measures — variance ratios around 1.05 to 1.20.

That sounds negligible and is not. With identical averages and a spread just seven per cent wider, the ratio at the extreme rises to about 1.35 to 1 at two spreads above average, 1.9 at three, and 2.9 at four. At fifteen per cent wider, roughly eight to one.

The pattern is real and cross-national — but variance ratios differ between countries and have changed over time, and a biological constant does not vary by nation.

And it works at both ends. The same arithmetic that puts more men at the top puts more men at the bottom: intellectual disability, learning difficulties, exclusion, dropout, catastrophic outcomes. Anybody using this to explain why the top of a profession is male has, in the same breath, explained why the bottom of society is.

Almost nothing people fight about is decided at the average. Professorships, prizes, prisons and diagnoses are all at the edges — which is why two people can both be right about the middle and still disagree completely.

5What People Want To Do

The largest psychological difference between men and women that has ever been measured is not an ability. It is a preference. It is bigger than every gap in Chapter Three combined, it has been measured on half a million people, and it is bad news for both camps.

5.1 — The finding

In 2009 a group of researchers pooled every study they could find on vocational interests — what work people say they would like to do — and analysed them together. The combined sample was over half a million people.

They found one dimension along which men and women differ more than on anything else in psychology. They called it Things versus People.

Word Box

The Things–People dimension: a measure of whether somebody prefers work oriented around objects, systems and machines, or work oriented around persons.

At the Things end: engineering, mechanics, construction, physics, computing hardware. At the People end: teaching, nursing, counselling, human resources, social work.

It is not a measure of ability, sociability or warmth. It asks what a person would choose to spend their working life doing.

Why it matters here: the sex difference on this one dimension is roughly 0.93 — seventy-four in the hall — which is larger than any cognitive difference in this part and larger than most differences measured between any two human groups on anything psychological.

5.2 — The numbers

Interest areaEffect sizeMan higher, per 100 pairs
Engineeringabout 1.1178
Things versus Peopleabout 0.9374
Realistic — hands-on, mechanicalabout 0.8472
Scienceabout 0.3660
Mathematics as an interestabout 0.3459
Investigative — research, analysisabout 0.2657
Enterprising — business, persuasionabout 0.0451
Conventional — records, administrationabout −0.3341
Artisticabout −0.3540
Social — teaching, helping, caringabout −0.6832

Two things in that table deserve attention beyond the top line.

Interest in mathematics differs by 0.34 while ability differs by roughly zero. That is the whole argument about mathematics and women, relocated. The gap was never in what girls could do. It was, at least substantially, in what they wanted to do with it — which is a completely different problem requiring a completely different response.

And business interest is dead level. An effect size of 0.04 is nothing at all. Whatever explains the composition of company boards, it is not a difference in wanting to be there.

In Real Terms

Take the engineering figure of about 1.11 and put it in a lecture theatre.

Suppose a country had removed every barrier — no discrimination in admissions, no hostility in the workplace, no cultural discouragement, nothing at all except each person choosing freely from an identical menu. If the only remaining input were the measured strength of interest, that engineering class would still not be half women.

Now hold the other end. In a nursing intake, a social work programme, a primary school staffroom, the same arithmetic running the other way produces the imbalance everybody has already seen and nobody calls a crisis.

The uncomfortable observation is that both imbalances have the same cause and only one of them is treated as a problem — which tells you the concern was never really about imbalance.

5.3 — Where does the preference come from?

The obvious reply is that interests are shaped by society, and this is clearly partly true. Children are handed different toys, praised for different things, and shown different futures. Nobody serious disputes that this affects what they later want.

Two findings complicate the simple version.

The first is that the Things–People difference appears in every country where it has been measured, across more than fifty nations with very different economies, religions and gender arrangements.

The second is the one that does real damage to the simple version, and it is the subject of the whole of Part Five: the difference tends to be larger in richer and more gender-equal countries, not smaller. Whatever produces it, the pattern does not shrink as barriers come down.

I am not going to resolve that here. Part Five is entirely about it, including the serious criticisms of the finding. What matters for this chapter is that the straightforward socialisation account predicts the opposite of what is observed, and an account that predicts the opposite of the observation has a problem.

The Hidden Assumption

Every argument about women and work assumes that the question is ability.

Can women do engineering. Are women as good at mathematics. Is there a difference in capability. Both camps have spent decades on this, one trying to establish a difference and the other trying to refute it.

And on the evidence, ability is largely not where the action is. Ability differences in the relevant domains are close to zero. Interest differences are among the largest measured differences in psychology. The variable everybody argues about is the small one.

Now notice why this is worse for both sides, not better.

An ability gap is a solvable problem. Teach, train, fund, remove the barrier, and the gap closes — the mathematics ratio in Chapter Three fell by two thirds in twenty years. Everybody knows how to work on ability.

A preference gap has no such handle. You cannot train somebody into wanting a different life. Addressing it requires either accepting unequal outcomes as legitimate, or telling large numbers of people that what they want is wrong and was installed in them — and that second option is a thing nobody wants to say out loud about adults, in either direction.

The general form: arguing about the tractable variable because the intractable one has no acceptable answer.

I want to be clear that naming the dilemma is not the same as resolving it, and I am not resolving it here. Chapter Nine takes it up properly, once the rest of the measurements are on the table.

Remember This

The largest psychological sex difference ever measured is not an ability. It is a preference. On the Things–People dimension — objects and systems against persons — the effect size is about 0.93, or 74 in the hall, from a pooled sample of over half a million people.

Engineering interest is larger still at about 1.11. Interest in caring and teaching work runs the other way at about −0.68.

Two rows matter most. Interest in mathematics differs by 0.34 while ability differs by about zero — the whole argument relocated. And business interest is 0.04, which is nothing, so whatever explains the composition of boards, it is not a difference in wanting to be there.

The difference appears in every country measured — and tends to be larger in richer and more gender-equal ones, which is the opposite of what a straightforward socialisation account predicts. Part Five is about that.

Both camps argue about ability because ability is fixable. A preference gap has no handle: you cannot train somebody into wanting a different life. Addressing it means either accepting unequal outcomes as legitimate, or telling adults that what they want is wrong and was installed in them.

6Personality, And A Statistical Trap

Measured one trait at a time, personality differences are modest. Measured all at once by one particular method, they look enormous. And one disposition — aggression — shows more clearly than anything else in this part what a moderate average does at the far edge of a distribution.

6.1 — One trait at a time

Personality research mostly organises itself around five broad traits, and the sex differences on them have been measured many times in many countries.

Word Box

The Big Five: the five broad dimensions most personality research uses.

Agreeableness — warmth, trust, cooperation. Neuroticism — how readily a person experiences negative emotion. Extraversion — sociability and assertiveness. Conscientiousness — orderliness and self-discipline. Openness — interest in ideas, art and novelty.

Why it matters here: these are the units in which the argument in this chapter is conducted, and the argument turns on whether you look at them one at a time or all together.

Taken separately, the findings are consistent and moderate. Women score higher on agreeableness by around 0.4 to 0.5 — roughly sixty-two to sixty-four in the hall, running the other way. Women score higher on neuroticism by a similar amount. Differences on conscientiousness and openness at the broad level are small. Extraversion is close to level.

There is one wrinkle worth knowing, because it explains part of the dispute that follows. Each broad trait contains narrower components, and the sexes sometimes differ in opposite directions within the same trait. Within extraversion, for instance, women tend to score higher on warmth and enthusiasm while men score higher on assertiveness. Average those together and they cancel, producing a broad difference near zero that conceals two real ones.

6.2 — The move

Now the trap.

Suppose you have sixteen personality measurements rather than five. On each one, men and women differ modestly. What happens if you ask not “how different are they on this trait” but “how well could you tell a man from a woman using all sixteen at once?”

That is a different question, and it has a different answer. Because the small differences do not all point the same way, combining them can separate the groups far better than any single measure. A team applying this method to a large sample reported a combined distance so large that it implied only around one tenth of the two groups overlapped — a figure vastly beyond anything in Chapter Three.

That claim has been quoted enormously widely, usually without the argument that followed it.

The Argument — are personality differences small or enormous?

Two teams of competent researchers, looking at overlapping data, reporting answers an order of magnitude apart. The disagreement is genuine and it is about method rather than data.

They are large, and single traits are the wrong unit

Nobody meets a person’s agreeableness score. They meet a whole personality, and a whole personality is a combination. Analysing traits one at a time systematically understates how different two groups are, because it throws away the information contained in the pattern. If you can separate two groups almost completely using the full profile, then the groups are almost completely separable, whatever any individual trait shows.

They are modest, and the combined figure is inflated

Three objections. The method corrects for imperfect measurement in a way that substantially enlarges the result. The size of the combined distance depends on how many measures you throw in, so it can be pumped up by adding more. And it answers a classification question, not a difference question — being able to guess somebody’s sex from a profile does not mean any trait differs much, only that many tiny signals add up.

Two questions, one word

Both results are correct answers to different questions. “How different are men and women on trait X?” — modest, and the numbers in §6.1 are right. “How well can you tell them apart using everything at once?” — surprisingly well, and the combined figure is right. The dispute exists because both are reported using the word difference.

Where things stand: the third position, and it is not a dodge — it is the actual structure of the disagreement. A later analysis with a larger sample produced a smaller combined figure than the original, which suggests the first number was somewhat inflated while the phenomenon is real.

What would settle it: agreement on which question is being asked, which is not an empirical matter. For any practical purpose — hiring, teaching, predicting what a person will do — the single-trait numbers are the relevant ones, because you are dealing with a person and not classifying a sample.

Why people care so much: because the combined figure sounds like a licence. “Ninety per cent non-overlapping” reads as two kinds of creature, and it is repeated constantly in exactly that spirit. What it actually describes is a statistical classification exercise, and Chapter One’s rule applies with full force: a classification is not a person.

6.3 — Aggression, and the clearest tail in the book

One disposition deserves separate treatment, because it is among the largest behavioural differences and because it demonstrates Chapter Four’s arithmetic more clearly than any other example available.

Men are more physically aggressive. Meta-analyses put the difference somewhere around 0.5 to 0.6 — in the hall, roughly sixty-four to sixty-six in a hundred. That is a moderate difference, comparable to writing ability running the other way, and if you stopped there you would conclude it was interesting and not dramatic.

Two details make it more than that.

It appears very early. The sex difference in physical aggression is detectable in children by around the age of two or three, before most proposed socialisation could plausibly have done its work, and it is found across cultures.

And indirect aggression runs the other way. Aggression conducted through exclusion, reputation and social manoeuvre — damaging somebody without touching them — shows either no sex difference or one favouring girls and women, depending on age and measure. So the honest statement is not that one sex is more aggressive. It is that the sexes differ in method, and that the male method is the one that leaves marks and enters crime statistics.

In Real Terms

Now watch what a moderate average difference does at the extreme, because this is the cleanest demonstration in the whole part.

Physical aggression differs by roughly 0.6 — about sixty-six in a hundred pairs. Two thirds. Not remotely categorical. Pick a man and a woman at random and the woman is the more physically aggressive of the two a third of the time.

Now go to the far end of that same distribution. Homicide is committed by men in the region of nine cases out of ten, worldwide, in every country that keeps records. Men are also the substantial majority of the victims.

Sixty-six per cent in the middle. Ninety per cent at the edge. Nothing extra was added — no additional difference, no separate explanation. That is Chapter Four’s arithmetic happening to real people, and it is why average differences and extreme outcomes so often look like evidence about different worlds.

The same structure applies to risk. Men take more physical and financial risk, by a modest margin on most measures — the effect varies a great deal by the kind of risk, and is smaller than most people assume. And men die from accidents, drowning, falls and violence at rates far beyond that modest margin, for the same reason.

Hold this beside §2.5. Men die around five years earlier, everywhere. A meaningful share of that gap is not a fact about male bodies at all. It is the tail of a moderate difference in behaviour, arriving as a mortality statistic.

6.4 — What survives

Strip the dispute away and three things are solid.

Women score higher on agreeableness and on neuroticism, by moderate amounts, consistently, across many countries. These are among the more replicable findings in personality research.

Broad traits hide opposite-running components, so a reported difference of zero at the broad level is not evidence of no difference underneath.

And personality differences, like the interest differences in Chapter Five, tend to be larger in richer and more gender-equal countries. The same pattern, in a different literature, measured by different people. Part Five deals with it properly, and the fact that it appears independently in interests and in personality is the reason it cannot be waved away.

Remember This

One trait at a time, personality differences are moderate. Women score higher on agreeableness and on neuroticism by around 0.4 to 0.5. Other broad traits are close to level.

Broad traits hide opposite-running parts. Within extraversion, women score higher on warmth and men on assertiveness; averaged, they cancel into a difference near zero that conceals two real ones.

Combine many measures at once and the picture changes dramatically — one widely quoted analysis reported only about a tenth of the two groups overlapping. That figure has been criticised for correcting measurement error in an inflating way, for growing as you add variables, and for answering a classification question rather than a difference question.

Both answers are correct to different questions. How different are men and women on a given trait — modest. How well can you tell them apart using everything at once — surprisingly well. They collide because both use the word difference.

On aggression: men are more physically aggressive by about 0.6 — sixty-six in a hundred, so a third of the time the woman is the more aggressive of a random pair. Indirect aggression runs the other way. The sexes differ in method, and the male method is the one that leaves marks.

And at the extreme, that sixty-six per cent becomes roughly ninety per cent of homicides, worldwide, with men also the majority of victims. Nothing was added between those two numbers except distance from the average. That is Chapter Four's arithmetic happening to real people.

For anything practical — hiring, teaching, predicting a person — the single-trait numbers are the relevant ones. “Ninety per cent non-overlapping” is a statistical classification exercise, and a classification is not a person.

7What Is Actually In The Brain

More has been written about male and female brains than about any other topic in this part, and a large share of it is unreliable. This chapter separates what survives scrutiny from what does not, and the second pile is much bigger.

7.1 — The one solid number, and why it means less than it looks

Men’s brains are larger, by roughly ten to twelve per cent on average. This is not in dispute.

Nearly all of it is body size. Larger bodies come with larger brains across the animal kingdom, and within each sex the same relationship holds — taller people have bigger brains. Once you account for body size, most of the gap goes.

And brain size is a poor predictor of anything cognitive within the human range. The correlation between brain volume and measured intelligence is real and modest, and it does not do the work that people who cite brain size want it to do — which is why Chapter Three found essentially no difference in general intelligence despite this ten per cent gap.

7.2 — What is reliably found

The best current evidence comes from very large population imaging studies rather than the small laboratory samples that dominated earlier decades.

Word Box

Grey matter is the tissue where the cell bodies of neurons sit — roughly, where processing happens. White matter is the wiring between them, wrapped in a fatty sheath that gives it its colour and speeds signals up.

Cortical thickness is how deep the outer sheet of grey matter is at a given point on the brain’s surface. It is measured in millimetres and varies across the brain in everybody.

Why it matters here: these three are what large imaging studies actually measure, and none of them has a simple meaning. More is not better and thicker is not smarter. A reported difference in any of them is a difference in a physical measurement whose behavioural consequence, in nearly every case, has not been established.

Those studies do find average differences: in the raw volumes of many regions, in the proportions of grey and white matter, and in cortical thickness, where women tend to score higher. The differences are statistically robust at these sample sizes.

They are also, without exception, differences with large overlap. Not one of them separates the sexes. In the language of Chapter One, they sit in the small-to-moderate range, and most shrink further once total brain size is accounted for.

So the honest summary of the reliable brain literature is: measurable average differences in many structures, none of them large enough to characterise an individual, and none of them yet connected to a specific behavioural difference by a demonstrated causal chain.

That last clause is the important one, and it is where most popular writing goes wrong. Finding that a region differs on average is not the same as showing that the difference causes anything.

7.3 — A famous finding that was never there

How We Actually Know This

In 1982 a study reported that a particular bundle of fibres connecting the two halves of the brain was larger in women. It was published in a leading journal and became one of the most widely repeated claims in popular science — the anatomical basis, it was said, for female intuition, for women’s supposedly greater emotional integration, for a dozen confident stories.

The sample was fourteen brains. Nine male, five female, obtained post mortem.

Over the following fifteen years, dozens of studies attempted to reproduce it. In 1997 a meta-analysis pooled forty-nine of them and found no reliable sex difference in the size or shape of that structure once overall brain size was accounted for.

What this shows: not that the original researchers were dishonest. They found a difference in fourteen brains, which is exactly what small samples do — they produce large, striking, unreplicable results. The finding entered public knowledge in the 1980s and is still repeated today, four decades after the evidence stopped supporting it.

What it cannot show: that all such findings are wrong. Some replicate. The lesson is about sample size and about how long a dead claim survives in public circulation, not about the field as a whole.

That case is not an isolated embarrassment. It is representative of a structural problem in this literature, and the scale of the problem is worth stating with a number attached.

In Real Terms

Here is the scale problem in this literature, laid out plainly.

For decades, a typical brain-imaging study comparing men and women used a few dozen people per group. Recent work on how many participants are actually needed to get a reproducible result linking brain measurements to behaviour suggests the answer runs into the thousands.

Put those side by side. A field spent thirty years running studies at a size now understood to be one or two orders of magnitude too small — and publishing the results, which were then read, quoted and turned into books.

There is a second problem sitting on top. In 2016 researchers found that a widely used statistical procedure in brain imaging had been generating false positives at rates far above what anybody assumed, affecting a substantial share of the published literature.

So when you encounter a striking claim about male and female brains, the base rate matters more than the claim. The default assumption should be that it came from a study too small to be trusted, because for most of the history of this field, that is what it was.

7.4 — Two brains or a mosaic?

The Argument — are there male and female brains?

This dispute has the same structure as Chapter Six, which is worth noticing, because a pattern that appears twice in one part is telling you something about how these arguments work.

Brains are mosaics

A large analysis of brain scans looked at whether individuals were internally consistent — whether a person with one “male-typical” feature tended to have the others. Overwhelmingly they did not. Most brains are patchworks, combining features that are more common in men with features more common in women. Brains that are consistently male-typical or female-typical throughout are rare. So there is no such thing as a male brain or a female brain, only a population of mosaics with differing average feature frequencies.

Brains are classifiable

If brains carried no sex information, you could not guess a person’s sex from a scan. You can, well above chance, using standard statistical methods on the full pattern. That accuracy has to come from somewhere. A claim that the categories do not exist is hard to square with the fact that a computer can sort them.

Two questions, one word — again

Both findings are correct and they answer different questions. Is any individual brain reliably identifiable as male or female by inspection? No — the mosaic result stands. Does the overall pattern carry sex-related information detectable in aggregate? Yes — the classification result stands. Exactly as in Chapter Six, the dispute survives because both are reported using the same vocabulary.

Where things stand: the third position, and its recurrence is the finding. Twice in one part, competent researchers have produced apparently opposite conclusions by asking a classification question and a difference question and calling both “difference.” That is not a failure of any particular team; it is a structural weakness in how this whole field reports itself.

What would settle it: nothing empirical, because there is no empirical disagreement. Both results replicate.

Why people care so much: because “male and female brains exist” reads as a licence for everything in Parts Two and Three, and “there is no such thing” reads as a refutation of it. Neither follows. A brain difference would still be a group average, and Chapter One already disposed of what a group average can justify about a person.

So the practical instruction from this chapter is unglamorous. Treat any striking claim about male and female brains as probably wrong until you know the sample size, and treat any claim that connects a brain difference to a behaviour as almost certainly unestablished, because that chain has very rarely been demonstrated.

Remember This

Men’s brains are about ten to twelve per cent larger, almost entirely because of body size, and brain size is a weak predictor of anything cognitive within the human range — which is why Chapter Three found no difference in general intelligence.

Large modern imaging studies do find average differences in regional volumes and cortical thickness. Every one has large overlap, none separates the sexes, and none has been connected to a behavioural difference by a demonstrated causal chain.

A 1982 finding that a fibre bundle was larger in women became one of the most repeated claims in popular science. The sample was fourteen brains. A meta-analysis of forty-nine later studies found no reliable difference. It is still being repeated today.

Typical studies in this field used a few dozen people per group. Reproducible brain-behaviour findings are now understood to need thousands. A widely used statistical procedure was also found to be producing false positives far above assumed rates.

Are there male and female brains? Individual brains are mosaics — internally inconsistent patchworks. And a computer can still sort them above chance. Both are true. It is the same two-questions-one-word collision as Chapter Six, and its appearing twice is itself the finding.

8Where The Differences Come From

Having measured the differences, the obvious next question is whether they are built in or taught. That question has been dead in the research for about fifty years, and understanding why it died is more useful than any answer to it.

8.1 — The strongest evidence available

You cannot run the experiment. Assigning babies to hormone levels or to upbringings is not something any ethics committee will approve, and the one notorious attempt to raise a child as the opposite sex was a single case, ended in tragedy, and is cited with equal confidence by both camps — which is what a single case always permits and never justifies.

So researchers use situations where nature has arranged something close to an experiment.

Word Box

Prenatal androgens: hormones, chiefly testosterone, present in the womb. Male foetuses are exposed to substantially more than female ones during specific developmental windows.

Congenital adrenal hyperplasia: an inherited condition in which the adrenal glands produce excess androgens from before birth. A girl with the condition develops with androgen exposure far above the usual female range.

Why it matters here: this is the closest thing to a controlled experiment that exists in humans. It varies prenatal hormone exposure without anybody choosing to, and it allows a comparison with unaffected sisters raised in the same households.

The comparison with sisters is what makes this worth the space. Two girls in one house, one exposed to high androgens before birth and one not, raised by the same parents with the same expectations. Whatever differs between them cannot easily be explained by upbringing.

How We Actually Know This

Girls with this condition have been studied for decades, compared with their own unaffected female relatives.

The most robust finding is about play. Affected girls show markedly more interest in toys and activities typically preferred by boys, and less in those typically preferred by girls. This has been reproduced across many studies and pooled analyses. Follow-up work finds shifts in later occupational interests in the same direction — which connects directly to the Things–People dimension in Chapter Five.

The strongest detail: the toy-preference difference persists even where parents actively encourage female-typical play. If it were simply a response to how these girls are treated, encouragement in the opposite direction should reduce it, and it does not.

What it cannot show: a great deal. Girls with the condition often have visibly different bodies and know they are different; parents know the diagnosis and cannot be blinded to it; the condition has other physical effects and requires lifelong medication. Sample sizes are usually small because the condition is rare. And the finding is about the direction of interests, not about ability, and not about anything in Parts Two and Three.

8.2 — Monkeys and toys

A second line of evidence removes human socialisation entirely.

Researchers offered toys to monkeys — animals with no exposure to advertising, parental expectation, school, or any idea of what a toy is for. Male monkeys spent more time with wheeled objects and female monkeys more with plush ones. The result has been found in more than one species.

It is striking and it should not be over-read. Critics have pointed out that the objects differed in colour, shape and how they moved, any of which could drive a preference without anything sex-typed being involved. The effects are moderate. And a monkey has no concept of a doll, so what the animals were responding to is genuinely unclear.

What the studies do establish is narrower and still useful: a sex difference in object preference can arise without human culture. That does not tell you how much of the human difference is cultural. It tells you the answer is not all of it.

8.3 — The most misused number in the subject

Before the framing itself, one term has to be dealt with, because it appears in almost every argument about this and almost nobody using it means what it means.

Word Box

Heritability: the share of the variation in a trait, within one population, that is statistically associated with genetic variation in that population.

Read that definition twice, because three words in it are doing work people routinely ignore.

Variation — it is about differences between people, not about how much of any one person’s trait is genetic. Asking how much of your height is genetic is like asking how much of a rectangle’s area is its width.

Within one population — the number is a property of a group in a place at a time. Change the environment and the same trait can have a different heritability.

Associated — not caused by, not fixed by, not unchangeable by.

Why it matters here: “this trait is sixty per cent heritable” is quoted constantly as though it meant “sixty per cent unchangeable.” It does not, and the next section shows why with a case everybody can check.

8.4 — Why the question is broken

The Hidden Assumption

Everything above assumes what everybody assumes: that a trait is either innate or learned, and the job is to work out the proportions.

This framing is dead in developmental biology and has been for decades, and its persistence in public argument is doing serious damage.

Nothing about an organism is specified by genes alone. A gene is an instruction that only means anything in an environment; change the environment and the same gene produces a different outcome. Equally, nothing is learned except through a biological system that determines what can be learned, how fast, and with what effect. There is no stage at which one hands over to the other.

Consider height. It is one of the most heritable traits there is — and average height in many countries rose by ten centimetres or more in a century, which is nutrition, not genetics. Both statements are entirely true. Height is highly heritable and highly responsive to environment, because heritability was never a measure of how fixed something is.

Now the consequence that matters most, and it is routinely got backwards by people quoting these numbers. Heritability within a group tells you nothing about differences between groups. A trait can be eighty per cent heritable inside each of two populations while the entire gap between them is environmental. This is not a technicality; it is the single most misused statistic in the whole subject.

The general form: a false binary imported from an argument that ended before most of the participants were born.

8.4 — The honest scoring

Given all that, here is what can actually be said about causes, scored plainly.

Physical differences (Chapter Two): overwhelmingly biological. Driven by hormones at puberty, visible in every population, and unchanged by any social arrangement anybody has tried. Training moves individuals substantially and does not close the population gap.

Interest differences (Chapter Five): mixed, with a real biological component and a real social one. The prenatal hormone evidence and the animal work establish that some of it is not taught. The size of the effect in humans, and the share of the total it accounts for, are not established. Anybody giving you a percentage is guessing.

Cognitive differences (Chapter Three): mostly too small to have an interesting cause. Where a difference exists and is large enough to study, such as mental rotation, it responds to training — which does not make it purely learned, but does mean it is not fixed.

Personality differences (Chapter Six): unresolved, and complicated by the cross-country pattern. Differences being larger in freer societies is not what either a simple biological or a simple social account predicts. Part Five is about that.

And the tail differences (Chapter Four): unknown. The mathematics ratio moved by two thirds in twenty years, which rules out anything wholly fixed, and it then stopped moving, which rules out nothing else.

Remember This

You cannot run the experiment, so researchers use situations nature arranged. Girls with a condition causing high prenatal androgen exposure show markedly more male-typical play and later interests than their own unaffected sisters — and the difference persists even when parents actively encourage the opposite.

Monkeys offered toys show a similar split, with no advertising, parents or schooling involved. Read narrowly: a sex difference in object preference can arise without human culture. That does not say how much of the human difference is cultural — only that it is not all of it.

The whole “innate or learned” question is dead in the science. Genes only mean anything in an environment; learning only happens through a biological system. Height is among the most heritable traits there is, and average height rose ten centimetres in a century on better food. Both true.

Heritability within a group tells you nothing about differences between groups. A trait can be eighty per cent heritable inside each of two populations while the whole gap between them is environmental. This is the most misused statistic in the subject.

Honest scoring: physical differences, overwhelmingly biological. Interests, genuinely mixed and nobody can give you the proportions. Cognitive differences, mostly too small to have an interesting cause. Personality and the tails, unresolved — and anyone quoting a percentage for any of them is guessing.

9Same, Equal, And Equal Outcomes

Three completely different ideas share one word, and almost every argument about equality is two people using it to mean different things. Separating them dissolves most of the argument — and leaves one genuine dilemma that nobody has an answer to.

9.1 — Three ideas

Take the sentence “men and women are equal.” It can mean three unrelated things.

Same. That men and women have the same distribution of traits. This is an empirical claim. It can be measured, and eight chapters have just measured it. The answer is: identical or nearly so on most cognitive measures, substantially different on some physical ones, and substantially different on interests. So as a general claim it is false, and as a claim about any specific trait it needs checking one trait at a time.

Equal in worth. That men and women are owed the same consideration, the same rights, the same standing. This is a moral claim. No measurement bears on it. You cannot confirm it with a study and you cannot refute it with one.

Equal in outcome. That men and women should end up in the same proportions in every role, at every level, with the same pay and the same representation. This is a political claim about results, and whether it follows from anything depends entirely on arguments that have to be made separately.

Now the observation that matters. Nobody has ever derived the second from the first, and nobody needs to.

Consider how obvious this is elsewhere. A weak man and a strong man are not the same on any measure of strength — the difference is larger than most in this part. Nobody concludes that the weak man has fewer rights. A person with a measured intelligence of 90 and one with 140 differ enormously. Nobody thinks the first should be governed by the second.

Equality of worth was never a claim about measurements. It has never rested on sameness in any other context, and there is no reason it should be made to rest on sameness here.

Which means the entire eight chapters you have just read are irrelevant to it. Every number in this part could double and the moral claim would be untouched.

9.2 — What separate categories actually concede

There is a version of this worth taking seriously because it is what people usually mean when they say the sexes cannot really be equal. It goes: if a man lifts a hundred kilos and a woman lifts fifty, then calling them equal is a fiction — real equality would mean the same number, and since the numbers differ, the equality is invented.

Look at what competitive sport actually does with that problem, because sport has been running the experiment for a century.

It creates separate categories. And a separate category is an explicit, public admission that the sameness claim is false in this domain — nobody sets up a women’s hundred metres if the times are the same.

But notice what it is not conceding. It is not conceding that women’s competition is less worth having, or that female athletes are lesser people, or that the winner deserves less. Weight classes in boxing work identically, and nobody thinks a flyweight champion is a fraud.

So the argument proves less than it seems to. Separate categories are not a fudge covering an embarrassment. They are what you do when a difference is real and the activity is still worth having for both groups. That is a coherent position, held openly, in public, by millions of people every weekend.

There is a limit to it, and the limit is the useful part.

Adjusted standards make sense where the activity is the point. A race exists to find the fastest person in a field. You can define the field however you like, because the race has no purpose outside itself.

Absolute standards make sense where the outcome is the point. A person trapped in a burning building weighs what they weigh. There is no adjusted category of human being to be carried down a ladder. Where a job genuinely requires a physical capacity, the honest arrangement is a standard set by the task and applied to every applicant regardless of sex — which some women pass and many men fail, and which is a test of the person rather than of the category.

Almost every real dispute about physical standards is a dispute about which of those two situations applies, conducted by people who have not noticed there are two.

9.3 — The dilemma nobody wants

Now the genuine problem, which the separation above does not dissolve.

Chapter Five found that men and women differ substantially in what they want to do — around 0.93 on the Things–People dimension, larger than any ability difference in this part.

Follow that through. If a society removed every barrier — no discrimination, no harassment, no discouragement, nothing but people choosing freely — the outcomes would not be equal. Engineering would remain disproportionately male. Nursing and primary teaching would remain disproportionately female. Not because anybody was stopped, but because of what people chose.

Which produces the following, and it is the crux of the entire modern argument:

Equality of opportunity and equality of outcome are not two ways of saying the same thing. Given a real difference in preferences, they are in conflict, and you cannot have both.

Neither camp can accept this cleanly.

For the equality-of-outcome position, it means that any programme aiming at proportional representation must, past some point, be working against what people want — and the honest versions of that programme have to say either that the preferences are illegitimate, or that the outcome matters more than the preference. Both are sayable. Neither is usually said.

For the traditionalist position, it is worse than it looks, and the reason is the pattern in Chapters Five and Six. The interest and personality differences get larger in freer and richer societies. Which means the preferences being appealed to are most visible precisely where women have most choice — so they cannot easily be used to justify restricting choice. You do not need a rule to make people do what they were going to do anyway. Every rule in Parts Two and Three exists to stop somebody doing something. A preference argument cannot justify a prohibition, because a preference needs no enforcement.

And the whole thing sits on an unresolved question. Preferences form inside societies. If they were installed, the argument changes completely. The cross-country pattern is the strongest evidence against the simple installation story, and it is contested, and Part Five is about it.

9.4 — What this part actually established about the rules

Set every finding beside the thing this series is about.

The largest difference measured here is physical, in power and speed. It does not explain why a widow may not remarry.

The next largest is interests. It does not explain why a woman must marry inside her caste.

Cognitive differences are mostly close to zero. They do not explain seclusion, dowry, or a girl married before puberty.

Male variability is real at the extremes. It does not explain why a photograph of a woman at a protest was answered with a rumour about her body.

Not one measured difference in this part maps onto a single rule in Parts Two and Three. The rules are not what you would design if you were responding to these findings; they bear no relationship to them at all. They are what you would design if you were solving an inheritance problem and a boundary problem, which is what Parts Two and Three found.

The Hidden Assumption

And now the one aimed at this part, and at me.

Both camps assume that the size of these differences bears on the political question. Part One named this as a bad bargain and Chapter One of this part repeated the warning. Then I wrote sixty pages of measurement anyway.

Ask why. Not one number in Chapters Two through Eight was needed to establish anything about what anybody is owed. Section 9.1 shows that equality of worth never rested on sameness in any other context, and does not here. The moral question was settled before the first measurement and is unaffected by all of them.

I wrote the part because the argument is being conducted in these terms and a person who refuses to engage with the numbers loses to whoever quotes them loudest. That is a reason to write it. It is not a reason to believe the numbers matter, and there is a real cost: participating in an argument on false terms lends those terms credibility. Every careful, honest, well-sourced page I have just written is also sixty pages of implicitly agreeing that this is what we should be arguing about.

The general form: fighting on ground you have already shown to be the wrong ground, because that is where the other side is standing. I do not have a better answer, and I would rather name it than pretend I do.

What I can do is keep saying it. Every time a number in this part is quoted at you as though it settled something about what a woman may do, the missing step is the same one, and it has never been supplied by anybody.

Remember This

Three ideas share one word. Same is an empirical claim — measurable, and mostly false for physical traits and interests, mostly true for cognitive ones. Equal in worth is a moral claim that no measurement touches. Equal in outcome is a political claim about results.

Nobody ever derived the second from the first, and nobody needs to. A weak man and a strong man differ more than most gaps in this part, and nobody concludes the weak man has fewer rights. Equality of worth never rested on sameness anywhere else.

Separate sporting categories concede that the sameness claim is false in that domain, and concede nothing else. Adjusted standards make sense where the activity is the point; absolute job-relevant standards where the outcome is the point, because a person in a burning building weighs what they weigh.

The genuine dilemma: given a real difference in what people want, equality of opportunity and equality of outcome are in conflict and you cannot have both. But a preference argument cannot justify a prohibition — you do not need a rule to make people do what they were going to do anyway, and every rule in this series exists to stop somebody.

Not one difference measured in this part maps onto a single rule in Parts Two and Three. And I wrote sixty pages measuring them anyway, because that is where the argument is being held — which is also sixty pages of agreeing that this is what we should be arguing about.

10An Honest List Of What We Do Not Know

Two lists, no hedging in either. What is genuinely unknown about differences between men and women, with the reason. And what is solid enough to carry into the rest of the series.

10.1 — Genuinely unknown

This part rests on measurement, which makes its gaps sharper and easier to state than the historical parts.

How much of the interest difference is biological. This is the central unknown and it is the one everybody claims to know. The prenatal hormone work and the animal studies establish that some of it is not taught. Nothing establishes the proportion, and the cross-country pattern points in a direction that neither simple account predicts. Anybody quoting you a percentage is guessing.

Why the differences are larger in freer societies. The pattern appears independently in interests, in personality and in preferences, measured by different teams using different instruments — which is why it cannot be dismissed. Every proposed explanation has serious problems, and the finding itself has been challenged on methodological grounds. Part Five is about this and will not resolve it either.

What produces the difference at the extremes. The mathematics ratio at the very top fell by roughly two thirds in twenty years and then stopped. That the movement happened rules out anything wholly fixed. That it stopped rules out nothing else, and nobody can currently distinguish a biological floor from a social one that has ceased shifting.

Whether variance ratios are stable. They differ between countries and have changed in some. A biological constant does not vary by nation, so something social is involved — how much, nobody has established.

How much of the brain literature is real. Chapter Seven’s honest position is that a large share of published findings came from samples now understood to be far too small, compounded by a statistical problem discovered in 2016. Which specific results survive is not knowable in advance, and the field is only now producing studies at the required scale.

Whether the mental rotation gap means anything outside a laboratory. It is the largest cognitive difference and it is one narrow task. Its connection to real-world outcomes is modest and its relationship to broader spatial ability is unclear.

10.2 — Solid

That physical differences are very large, and narrow. Height near an effect size of two, strength larger still, grip so separated that the distributions barely meet. And confined to power, speed and force application — shrinking in ultra-endurance events and reversing where body composition helps.

That most cognitive differences are close to zero. A review of forty-six meta-analyses found around four fifths small or negligible. This is the headline finding of the field and the least reported.

That average mathematical ability does not differ. Millions of students, multiple countries, favouring girls in several. Settled.

That girls read and write better, everywhere measured. The reading advantage appears in every participating country in every cycle. The writing gap is around half an effect size — roughly ten times the maths gap, and largely absent from public discussion.

That the Things–People interest difference is the largest measured psychological sex difference. About 0.93, from over half a million people, present in every country studied.

That greater male variability is real and operates at both ends. The arithmetic is not a theory. And it puts more men at the bottom in exactly the proportion it puts more at the top.

That equal averages are compatible with very unequal extremes. A seven per cent wider spread produces nearly three to one at four spreads above the mean, with identical means. This is arithmetic and cannot be argued with.

That heritability within a group says nothing about differences between groups. Not contested by anybody who works on this, and misused constantly.

And that none of the above bears on what anybody is owed. Established in Part One, restated in Chapter Nine, and the only conclusion in this part that does not depend on a single measurement.

10.3 — Who this was measured on

One closing observation, following the pattern of the previous parts.

Look at where these findings come from. University students, mostly. School assessments in wealthy countries. Population databases in Britain and Scandinavia. Volunteers who answered online surveys in English.

Part One named this problem: the people most studied by psychology are among the least typical humans available, drawn from countries holding a small minority of the world’s population. Nearly every number in this part carries that limitation.

The specifically Indian version is worse. There is comparatively little of this research on Indian populations, and where cross-national work includes India it is usually as one data point among dozens. So a series about Indian women has just spent sixty pages on measurements taken largely elsewhere.

I have applied them anyway, because the alternative is applying nothing. But the reader is entitled to know that when this part says “men and women differ by 0.93 on interests”, the sentence is built substantially out of people who have never been to the country this series is about — and that the one finding most relevant to India, the pattern in Chapter Five and Six about freer societies, is precisely the one nobody can explain.

Remember This

Genuinely unknown: how much of the interest difference is biological — the central question, and everybody claims to know it. Why differences are larger in freer societies. What produces the gap at the extremes, given that it moved by two thirds and then stopped. Whether variance ratios are stable. How much of the brain literature survives. Whether mental rotation means anything outside a lab.

Solid: physical differences are very large and narrow. Most cognitive differences are near zero — four fifths small or negligible across forty-six meta-analyses. Average maths ability does not differ. Girls read and write better everywhere measured. The Things–People interest gap is the largest measured psychological sex difference. Greater male variability is real and operates at both ends. Equal averages are compatible with very unequal extremes. Heritability within groups says nothing about gaps between them.

And that none of it bears on what anybody is owed — the only conclusion here that rests on no measurement at all.

Note who was measured. University students, school assessments and databases in rich countries, online volunteers answering in English. There is comparatively little of this research on Indian populations. A series about Indian women has just spent sixty pages on numbers built largely from people who have never been there.

Every Difference, In One Table

This part had no chronology to offer, so its reference table is the measurements themselves. Every difference discussed, sorted from largest to smallest, with the hall figure attached. Positive means men score higher.

What is measuredEffect sizeMan higher, per 100
Grip strengthvery large — off the tableabove 95
Upper-body strengthvery largeabove 95
Heightabout 2.092
Interest in engineeringabout 1.1178
Things versus People interestsabout 0.9374
Realistic — hands-on interestsabout 0.8472
Physical aggression0.5 to 0.664 to 66
Mental rotation of 3D shapes0.5 to 0.964 to 74
Spatial perceptionabout 0.4562
Interest in scienceabout 0.3660
Interest in mathematicsabout 0.3459
Investigative interestsabout 0.2657
Spatial visualisationabout 0.1554
Mathematical abilityabout 0.0551
Enterprising — business interestsabout 0.0451
General intelligenceabout 050
Vocabularyabout 050
Verbal ability, generalabout −0.147
Verbal and episodic memoryabout −0.244
Object location memoryabout −0.2543
Perceptual speedabout −0.2543
Readingabout −0.342
Conventional — administrative interestsabout −0.3341
Artistic interestsabout −0.3540
Neuroticismabout −0.439
Writingabout −0.536
Agreeablenessabout −0.536
Social — teaching and caring interestsabout −0.6832

Three things are visible in that table that are not visible in any single chapter.

The top of it is entirely physical and interests. Not one ability appears above 0.5 except a single narrow laboratory task.

The middle is a coin toss. Everything from mental rotation down to writing sits between 40 and 62 in a hall of a hundred — which means that for any of these, sex is a poor guide to a person.

And the largest non-physical differences at both ends are about what people want, not what they can do. Engineering at one end, caring work at the other. The abilities are in the middle where nothing much happens.

Sources & further reading — Part 4

Glossary

Every hard word used in this part, in plain English. Each was explained where it first appeared.

TermPlain meaning
Big FiveThe five broad dimensions most personality research uses: agreeableness, neuroticism, extraversion, conscientiousness, openness.
Congenital adrenal hyperplasiaAn inherited condition producing excess androgens from before birth. Provides the closest thing to a controlled experiment on prenatal hormones in humans.
DistributionThe full spread of a measurement across a group — not just the average but how many people sit at every value. Most mistakes in this subject come from discussing averages when the action is in the spread.
Effect size (d)How big a difference between two groups is, measured against how spread out they are. Convert it using the hall: 0.2 is 56 in 100, 0.5 is 64, 0.8 is 71, 2.0 is 92.
HeritabilityHow much of the variation in a trait within a population is associated with genetic variation. It is not a measure of how fixed something is, and it says nothing about differences between groups.
Mental rotationDeciding whether a three-dimensional shape is the same object turned around or a mirror image. The largest reliably found cognitive sex difference, and one narrow laboratory task.
Meta-analysisA study that pools many earlier studies to produce a combined estimate. Buys scale; inherits the biases of everything it pools, including the tendency for findings of “no difference” not to be published.
Prenatal androgensHormones, chiefly testosterone, present in the womb. Male foetuses are exposed to substantially more during specific developmental windows.
Publication biasThe tendency for results showing a difference to be written up and printed more often than results showing none, which inflates the apparent size of differences across a literature.
Things–People dimensionWhether somebody prefers work oriented around objects and systems or around persons. The largest measured psychological sex difference, at about 0.93.
Variability / variance ratioHow spread out a group is, and how much wider one group’s spread is than another’s. Ratios for cognitive traits are typically 1.05 to 1.20, with males wider — small in the middle and consequential at the extremes.

Download The complete book · 4.0 MB