Part 15 of 16The Rebuttal

The Case Against This Series

Fourteen parts have audited everybody else. This one puts the series in the dock. The strongest argument that the questions were wrong, the evidence was not neutral, the scoreboard was rigged by its own columns, and the even-handedness cost more than it bought.

Where We Left Off

Part 14 asked the only question that matters if any of this is more than talk. If the rules about women are a problem, what would actually shift them? The answer came out as a taxonomy of four shapes. You can ban the output. You can change the price. You can change who is in the room. You can supply the substitute. There is a fifth thing people do, which is tell each other to stop, and it is not an intervention at all.

The scoring was unkind. Every case in the series where a state banned an output failed, and failed in the same way: the demand stayed, so the behaviour moved. Cash payments moved real numbers by small amounts and stopped moving them the moment the money stopped. Only one shape produced a large, durable, well-identified effect — putting women where they are seen holding authority — and the active ingredient turned out to be the seeing, not the office. Supplying a substitute for what a rule was doing worked best where anyone tried it, and almost nobody tries it.

Then Part 14 turned the knife on itself. Its last hidden assumption said that fourteen parts of argument had been written by a person whose central finding is that argument works least, using the only tool he has. The honest description of the series, it said, is a specification for an intervention rather than an intervention.

That admission is where this part starts. If the series is willing to say that about its method, it should be willing to face the harder version. Not “my tool is weak” — that is a modest thing to confess and it costs nothing. The harder version is: the questions were wrong, the evidence would not bear the weight, the scoring was decided by the choice of columns, and the fairness I am proud of has a price I have never put on the page. Fourteen parts have given every side its strongest case. This part gives that treatment to the series itself.

How to read this document

The Six Boxes

Every part of this series uses the same six boxes. They are the notation. Here is one live example of each, so you know what you are looking at before you need to.

Word Box

Audit: a check on something done by going back to the original records, rather than by asking the person who did it whether they did it well. An accountant audits a company by looking at the receipts, not by reading the company’s own report of itself.

Why it matters here: this part calls itself an audit, and the word carries a promise. An audit that is run by the person being audited is a weaker instrument than a real one, and Chapter One says exactly how much weaker.

A Word Box explains a hard word the moment it first appears. Never later, never in a footnote. If a term is used, it is explained on the page where it is used.

In Real Terms

This series now runs to roughly 300,000 words. An ordinary adult reads non-fiction at about 200 to 250 words a minute. So reading the whole thing, once, at speed, without stopping to think, takes somewhere around twenty to twenty-five hours.

That is three full working days. It is more time than most people spend on any single voluntary thing in a year that is not a television series or a wedding.

An In Real Terms box takes a number you cannot picture and turns it into something with a body — a walk, a wage, a week, a room. A number nobody can picture is decoration.

How We Actually Know This

Where the evidence for a claim comes from, and what it cannot show. For example: this part says the series changed its author’s mind in specific places. The evidence is that the parts are published, dated and public on misterlove.in, and I cannot go back and quietly edit what Part 1 predicted before Part 10 tested it.

What it cannot show: whether I changed my mind because of the evidence or because of the writing. Nobody has access to that, including me.

This is the box that builds the most trust and the one most writers skip. It names the actual thing the claim rests on — the register, the transcript, the survey, the court record — and then says what that thing is bad at.

The Argument — should a writer prosecute his own work?

The question of this whole part, put once at the front.

Yes — he knows where the bodies are

Nobody else has read every draft. Nobody else knows which chapter was thin and got padded with a better example, or which claim survived only because the counter-evidence was hard to find in English. An outside critic attacks what is on the page. The author can attack what is not.

No — it is the strongest defence available

A book that criticises itself has bought immunity cheaply. The reader relaxes. Every later objection now sounds like something the author already thought of and priced in. Writing your own prosecution is the most effective possible way of preventing a real one.

Where things stand: both are true at once, and that is not a fudge. This chapter is genuinely better informed than an outside critic could be, and it genuinely functions as armour.

What would settle it: nothing I can do inside this document. Only an outside reader who reads this part and then still finds something it missed — and publishes it.

An Argument box is used where serious people disagree. Each side gets its strongest case, not a version that is easy to knock down. Then a verdict says where things stand and what would settle it. Often the honest verdict is that nothing available would settle it.

The Hidden Assumption

The signature of this series. Not a caveat and not a disclaimer. It digs out a premise that both sides of an argument accept without noticing, because that shared premise is usually where the real trouble is.

Example, and it is the assumption running under this entire part: that being wrong is a property of answers. Everyone arguing about a book argues about its conclusions. But a document can have every conclusion defensible and still be badly wrong, because the damage was done when the questions were chosen and nobody voted on that.

There are five Hidden Assumption boxes in this part. That is the range the method sets. Fewer and the series loses its signature. More and they stop landing.

Remember This

Every chapter ends with one of these. Three to five short paragraphs, key words in bold, the last one being the whole chapter in a sentence.

The test is simple. A reader who reads only these boxes should still come away with the complete argument.

One warning specific to this part

What This Document Is Not

This is not a retraction. I am not withdrawing the fourteen parts. If I thought a claim in them was false I would correct that claim, in that part, and say what changed and who caught it. That is the standing promise and it still holds.

This is something narrower and harder. It is the strongest case that can be made against the series as a whole — its framing, its evidence base, its method of scoring, its fairness rule and its likely effects — assembled by the person with the most to lose from it landing. Where an objection lands, I say so. Where it does not, I say that too, and I do not soften a bad objection to look humble. False modesty is a way of dodging criticism, not accepting it.

One more thing. Two chapters in this part are written as advocacy. Chapter Five puts the traditionalist case against the series at full strength. Chapter Six does the same for the feminist case. In both, I am arguing for a position, not reporting one. Part 6 and Part 7 of this series did the same thing on the main subject, and the same warning applies: while you are inside those chapters, you are reading a brief, not a survey.

1The Rule Turned Around

Fourteen parts have taken other people’s arguments apart. This one takes mine apart. Before that can mean anything, we have to agree what counts as a hit.

1.1 — What a case against actually is

Most books that criticise themselves do it in the last two pages. The author admits the work is incomplete. He says more research is needed. He thanks his critics in advance. Then the book ends and nothing in it has changed.

That is not a case against. That is a bow.

A real case against has to be able to do damage. It has to name a specific thing the document did, show why that thing was a mistake, and say what follows if the objection is right. If an objection is correct and nothing follows, it was not an objection. It was a mood.

So here is the test I am going to hold this part to. Every objection in it must come with an answer to one question: if this is true, what in the series is now worthless? Not “weaker”. Not “more complicated than I said”. Worthless — as in, a reader who believed it has been misled about something that matters to them.

Some objections in this part pass that test. Several do not, and I will say which. An honest prosecution includes the charges that fail, because a prosecution that only brings charges it can win is a performance.

Word Box

Steel man: the opposite of a straw man. A straw man is a weak, silly version of somebody’s argument, built so it is easy to knock down. A steel man is the strongest version of that same argument — often stronger than the version its own supporters usually give.

Why it matters here: this whole series runs on steel-manning, and Chapters Five and Six turn it on the series itself. If you find yourself thinking “no serious traditionalist would put it that well” or “no serious feminist would concede that”, that is the method working, not failing.

There is a reason to be suspicious of a document that steel-mans its opponents. Done badly, it flatters everyone and commits to nothing. Done well, it produces the one thing arguments almost never produce, which is a clear view of where the actual disagreement sits. Whether this series did it well is one of the things this part is trying to find out.

1.2 — The six ways a document like this can be wrong

People usually mean one thing by “wrong”. They mean a false statement. That is the least interesting failure and the easiest to fix. You correct the sentence.

A long piece of research can fail in at least six ways, and they get worse as you go down the list.

One: a fact is wrong. A number is misreported, a date is off, a study is described as showing something it did not show. Serious, embarrassing, easy to repair.

Two: the evidence will not bear the weight. Every fact is accurate, but the conclusion rests on them more heavily than they can take. This is much harder to see, because nothing in the text is false.

Three: the frame is wrong. The document divides the world into categories that do not carve it at the joints. Nothing inside the categories is false. The categories are the error.

Four: the question is wrong. The document answers the question it asked, and the question should not have been asked in that form, because asking it that way already granted something to one side.

Five: the method is self-serving. The rules the document follows are rules that happen to produce the kind of conclusion the author finds comfortable. Not by design. By selection — you keep the rules that keep giving you results you can live with.

Six: the effect is bad. Everything is true, the frame is fine, the question was right, and putting it into the world still made things worse — because of who reads it, what they use it for, and what it lets them stop doing.

Notice that only the first of those can be fixed with a correction notice. This part spends almost all its time on numbers two to six.

The Hidden Assumption

Everyone who argues about a book argues about its conclusions. The reviewer disagrees with the verdict. The supporter agrees with it. The comment section fights about it. All of that activity shares one unexamined premise: that being wrong is a property of answers.

It is not, or not mainly. By the time a document reaches its answers, the important decisions are already behind it. Somebody chose which question to ask. Somebody chose which things count as evidence. Somebody chose what would count as a good outcome and what would count as a bad one. None of those choices appear as claims in the text, so none of them can be disagreed with in the normal way. They are not stated. They are used.

This is why a document can be careful, honest, accurate in every line, and still do damage. It is also why the case against a piece of research is almost never made properly. The reader has been handed a page of answers and invited to fight about those, which keeps everybody busy at the shallow end.

The general form: an argument you are allowed to have is usually not the argument that decided the outcome. You saw this in Part 8, where the fight over whether the equality project “worked” turned entirely on which column was on the ledger, and nobody was fighting about the columns.

That box is the reason this part is ordered the way it is. It goes after the question first, the evidence second, the scoring third. The conclusions come last and get the least attention, because if the first three are sound the conclusions are mostly bookkeeping, and if they are not, arguing about conclusions is a waste of everyone’s afternoon.

1.3 — The standard of proof

An audit needs a standard. Without one, “I have considered the criticisms” means whatever the author wants it to mean.

Here is mine, and you can hold me to it, because it is written down and dated.

An objection lands if I cannot answer it without either changing a claim in the series, or admitting that a claim rests on something I chose rather than something I found.

An objection is survivable if it is true, but the series already says so somewhere, in the text, before this part — not in a footnote and not by implication.

An objection fails if answering it requires nothing more than pointing at what the series actually says, as opposed to what a hostile summary of it would say.

Chapter Nine has the scoreboard. Nine objections go in. Not all of them come out the same way, and the ones that land are not the ones I expected when I started writing this part.

How We Actually Know This

The claim that I did not know which objections would land is the kind of claim a writer makes and nobody can check. So here is the honest position on it.

The evidence for it: the fourteen parts are published, dated, and out of my hands. Chapter Seven lists six places where the writing moved me, and in each case the earlier part is on the record predicting one thing and the later part is on the record finding another. That pattern is checkable by anybody with a browser.

The evidence against it, which is the honest half: this document was written after all fourteen. Nothing stops me from writing “I did not expect this” about a conclusion I expected perfectly well. There is no way to verify a surprise. Treat every sentence in this part that begins “I was surprised to find” as a claim with no evidence behind it, because that is what it is.

1.4 — Who is speaking here

The method this series uses has a rule about declaring your position wherever it could pull the writing. This part is where that rule bites hardest, so let me put four things on the table.

I am an Indian man. I am writing about rules that mostly constrain Indian women. Every part of this series has been written by somebody who is not subject to the rules he is describing, and who has never once in his life had his movements, clothing or friendships negotiated by a family council. That is not a disqualification. It is a fact about the instrument.

I am the author of the thing being audited. Fourteen parts is a large amount of work and I do not want it to have been a waste. That is a motive, and motives do not stop operating because you have named them.

I publish under my own name on a site that is mine. Nobody commissioned this series and nobody can pull it. That removes one kind of pressure — no funder to please — and adds another. There is no editor who can tell me a chapter is bad.

And I have an audience. It is not enormous, but it exists, and it arrived because of the earlier parts. A writer with readers has a reason to keep producing the sort of thing that got him readers. If the honest conclusion of this part were “the series should not have been written”, I would like to think I would print it. I have no way to prove I would.

In Real Terms

Fifteen parts, about 300,000 words, roughly a thousand pages. That is about three times the length of a normal non-fiction book. Set next to a working life it is small — perhaps a year of evenings. Set next to a reader’s life it is enormous.

Put it this way. If a hundred people start Part 1, and each part loses a fifth of the ones who started it, then by Part 15 you are writing for about four people. This chapter is being read by a small number of unusually stubborn strangers, and pretending otherwise would be the first dishonest thing in the document.

Those four declarations are the whole of my standing. Everything after this chapter is argument, and you now know exactly who is making it and what he has riding on the outcome.

Remember This

A real case against has to be able to do damage. Every objection here has to answer one question: if this is true, what in the series is now worthless? An objection that changes nothing was a mood, not an objection.

A long piece of research can fail in six ways: a wrong fact, evidence that will not bear the weight, a wrong frame, a wrong question, a self-serving method, and a bad effect. Only the first can be fixed with a correction. This part is about the other five.

The assumption underneath every argument about a book is that being wrong is a property of answers. It is not. The decisions that matter — which question, which evidence counts, what a good outcome looks like — are never stated as claims, so they can never be disagreed with in the normal way.

I have declared four things that pull this writing: I am not subject to these rules, I wrote the thing being audited, nobody can edit me, and I have readers to keep.

The argument you are allowed to have is usually not the argument that decided the outcome.

2The Question Was Chosen

The series asked what the rules about women are for, and what happens if they change. That sounds like the neutral question. It is not, and this chapter shows exactly what it gave away before the first sentence.

2.1 — The question the series actually asked

Part 1 set the terms for everything that followed. It said the series would ask why rules about women’s dress, movement and sexual behaviour exist, what those rules are doing, and what actually happens — to a woman, and to the society around her — when they change.

Read that again and notice its shape. It is a what is it for question. It treats a rule the way an engineer treats a part he has found inside a machine he did not build. What does this do? What breaks if you remove it? What was it put there to solve?

Fourteen parts then did exactly that, carefully. And the care is the problem, because a question of that shape is not a neutral opening. It has already made three decisions, and none of them were argued for.

Word Box

Functional explanation: explaining something by what it does rather than by how it came about or whether it is right. “The heart exists to pump blood” is a functional explanation. So is “the dowry system exists to transfer property at marriage.”

Why it matters here: a functional explanation quietly turns a practice into a solution. Once you have called something a solution, you have implied there was a problem, and problems are things that need answers. The rule stops looking like something people do to other people and starts looking like something a society needs.

The first decision the question makes is that the rules are for something. That sounds obvious. It is not. Some arrangements are not for anything. They persist because the people who benefit have the power to keep them, and asking what problem they solve is like asking what problem a landlord solves. There is an answer — he houses people — and the answer does most of the work of the defence without ever presenting itself as a defence.

The second decision is that a society is the kind of thing that can have needs. Part 9 spent ninety-two pages treating a national birth rate as an object that can be too low, which means treating a country as a thing with interests, separate from the interests of the women who would have to bear the children. Part 9’s own last box admitted this. It admitted it in one box, on page eighty-something, at the end.

The third decision is the one nobody notices, and it is the largest.

2.2 — Who has to justify themselves

“What happens if the rules change?” is the question this series asked over and over. It is the spine of Part 5, Part 8, Part 9, Part 10 and Part 13.

Now try the mirror question. What happens if they stay?

The series never ran that one at the same depth. And there is no evidential reason for the asymmetry, because both are questions about consequences and both are answerable with the same tools. The reason is that one of them is the question people ask and the other is not.

This is a burden-of-proof problem, and it is worth being precise about it, because it is the single most powerful invisible thing in any public argument.

Word Box

Burden of proof: in an argument, the job of proving your case. Whoever carries it has to produce evidence. Whoever does not carry it wins by default if the evidence is unclear.

Why it matters here: in a courtroom the burden is assigned deliberately, in advance, and everyone knows who has it. In a public argument nobody assigns it. It settles on whoever is proposing the change, silently, and it never moves. That means uncertainty always helps whatever already exists. If nobody can prove anything, the rule stays.

Look at what that does to a series like this one. Fourteen parts of careful, honest reporting keep arriving at “the evidence is thin”, “the effect does not travel”, “this has never been tested outside one country”, “nobody has run the experiment”. Each of those is an accurate statement about the state of knowledge. Every single one of them, in the world outside the page, functions as a point for the side that wants nothing to change.

The series never chose that. It is a structural property of writing honestly about a contested arrangement while the arrangement is in force. Honest uncertainty is not neutral in its effects. It is neutral in its intentions, which is a different thing, and the difference is invisible from inside the writing.

How We Actually Know This

You can see the asymmetry inside the series itself, without taking my word for it. Count the chapters.

Part 10 spends a whole part testing the traditionalist’s prediction about what happens to a woman who breaks the rules. Four claims, unpacked, each run through a portability test — the check of whether an effect found in one place still shows up somewhere else, because an effect that does not travel belongs to the setting rather than to the act. That is the right way to treat a claim.

Now find the equivalent part testing the claim that the rules protect her. There is a version of it in Part 6, but Part 6 is written as advocacy, and advocacy is not a test. There is no part in this series that takes the sentence “these rules keep women safe” and audits it the way Part 10 audits “she ruins her own life”.

What this cannot show: whether the imbalance came from bias or from the state of the literature. Researchers have studied the costs of breaking the rules far more than the costs of keeping them, so a series built on the published record inherits that shape. Both explanations predict the same table of contents.

2.3 — What a serious version of the objection sounds like

Put together, the objection is not “you were unfair to one side”. It is sharper than that, and it does not depend on anyone’s motives.

It goes: a series that asks what a rule is for, and what changing it would cost, has built a machine that produces reasons for caution, no matter what data you feed it. The output is not evidence about the rules. It is evidence about the question.

That is a strong objection and I want to state the reply just as strongly, because I think the reply is right and the objection is still partly right, and both of those can be true.

The Argument — is “what is this rule for?” a neutral question?

Every serious study of a social arrangement starts by asking what it does. The question in dispute is whether that opening move is analysis or is already a concession.

It is already a concession

Asking what a practice is for assumes it is solving something. That framing has a direction built in: if you remove a solution, you should expect a problem to reappear. So the whole apparatus points at risk. Meanwhile the harms the rule is currently doing are the background, not the subject, because background does not need explaining. The framing does not have to be biased to produce biased output. It only has to be asymmetric, and it is.

It is the only question that leads anywhere

The alternative framing — these rules are simply domination and the analysis is complete — has been available for a century and has a poor record of shifting anything. Part 14’s central result depends entirely on the functional question. It was only by asking what dowry, seclusion and son preference are actually doing that the series could show that persuasion fails and substitution works. Drop the functional question and you lose the one finding in the series that could change a policy. You also lose the ability to show, as Part 14 did, that the stated function of chastity rules was solved cheaply thirty years ago and nothing moved — which is the most damaging fact in the entire series against the traditionalist.

Both, and the fix is a second question

The functional question is fine. Running it alone is not. A series that asked “what is this for?” should also have asked, at the same depth and with the same page count, “what is this costing right now, and who is paying?” That part was never written. Its absence is the finding, not the framing.

Where things stand: the third position is correct and I did not see it until this chapter. The functional question is not a concession by itself. Running it without its mirror is. In this series the mirror was run in fragments — Part 11 on men, Part 12 on enforcement, Part 3 on the machine — but never as its own part with its own budget of evidence.

What would settle it: writing the missing part and seeing whether it changes any conclusion. My honest expectation is that it would change the emphasis of most parts and the conclusion of none, which is either a defence of the series or a demonstration that the framing was doing more work than I thought. I cannot tell from here which.

Why people care so much: because “what is it for?” is how a defender of any arrangement would prefer the conversation to start, and “what is it costing?” is how a critic would. Choosing the opening question is most of the fight, which is why the fight is usually over before anyone notices it began.

2.4 — The woman at the protest

There is one more thing in the framing, and it is the least abstract thing in this part.

This series exists because of a specific person. A woman was photographed at a protest. The photograph travelled. Strangers who had never met her decided, on no evidence, that she sold intimate content online, and they said so at volume, with her face attached.

I did not write about her. I decided, early, that writing about her would repeat the injury, and I have said so in the text more than once. Instead I wrote about the norms underneath the attack — fifteen parts, a thousand pages, an examination of the rules that made “she is that kind of woman” a usable weapon.

Part 12’s last box already said the uncomfortable half of this: I used her as material, decided on her behalf that she would rather not be named, and put that decision on every cover. Here is the half that box did not say.

Turning her into a research question was itself a choice with a direction. What happened to her is not a puzzle. It is not epistemically interesting. It is a woman being hurt by people who wanted to hurt her, using a tool that was lying around. By treating it as the entrance to a fifteen-part inquiry, I converted an event that calls for a response into a subject that calls for a study. Those are different things, and the second one is much more comfortable for the person doing it.

The Hidden Assumption

Both sides of almost every argument about social rules agree on one thing without ever saying it: that understanding a thing is a step towards changing it. The traditionalist writes to explain why the rule is wise. The reformer writes to explain why it is unjust. Both are producing understanding, and both assume the understanding does work in the world.

Part 14 found that it mostly does not. Of the intervention shapes it scored, “tell people the truth and let them reason” was the one that behaved worst — so badly that the chapter classified it as a category error rather than a weak intervention. India’s own flagship persuasion campaign spent 79 per cent of the money it released to states on advertising the campaign.

Now put those two things next to each other. This series is 300,000 words of understanding, written by someone who has published a finding that understanding is the weakest lever available. If that finding is right, the series is not a small contribution to change. It is a large contribution to something else — and the honest name for that something else is not obvious to me.

The general form: the belief that explanation leads to change is held most firmly by the people whose only available tool is explanation. That is exactly the pattern you would expect if the belief were false and comforting rather than true.

That box is the most damaging thing in this chapter and it does not touch a single factual claim in the series. It is failure type six from Chapter One: everything true, and the effect still not what the writer thinks it is. Chapter Eight comes back to it with the price attached.

Remember This

The series asked a functional question — what are these rules for, and what happens if they change. That question is not neutral. It treats a rule as a solution, which implies a problem, which implies caution about removal.

The mirror question — what is this costing right now, and who is paying? — was never given its own part. Running the first without the second is the real objection, and it lands.

Because the burden of proof silently sits on whoever proposes a change, every honest finding of “the evidence is thin” works, in the world, as a point for keeping things as they are. The series never chose that. It is what honest uncertainty does when it is written about an arrangement that is currently in force.

The series began with a real woman being hurt. Turning that into a research question converted an event that called for a response into a subject that called for a study — which is far more comfortable for the person doing the converting.

Choosing the opening question is most of the fight, and the fight is usually over before anyone notices it has started.

3The Evidence Was Never Neutral

This series quoted a lot of studies. This chapter is about what those studies are actually made of — and about the fact that on this subject, more than most, people lie to researchers in a pattern.

3.1 — What the series rests on

Strip the fourteen parts down and the evidence comes in four kinds.

There are counts: censuses, national surveys, crime records, birth registers. Part 9 and Part 11 lean on these. They are the strongest thing here, because they were collected for administrative reasons by people with no stake in this argument.

There are experiments: a small number, mostly in psychology and economics. Part 13 found exactly one genuine randomised experiment in the whole subject — India’s 1993 rotation of reserved village council seats — and built a chapter on it.

There are correlations: the largest category by far. Partner counts and divorce rates. Education and marriage age. Female employment and fertility. Almost everything in Part 10 sits here.

And there are histories: legal records, colonial files, court judgments, campaign archives. Parts 2, 3, 7 and 12 run on these.

The objection in this chapter applies mainly to the third category and partly to the second. It does not much touch the first or the fourth. That distinction matters, and I will come back to it, because a chapter called “the evidence was never neutral” can easily be read as “throw it all out”, which is not what the evidence about the evidence says.

3.2 — The replication problem

Start with something that has nothing to do with women, rules or India.

Word Box

Replication: running somebody else’s study again, on new people, following their method, to see whether the same thing happens. It is the check that turns a result into a finding.

Effect size: how big a difference something makes, as opposed to whether the difference is real at all. A pill can have a real effect and a tiny one. Most arguments in public treat “real” and “big” as the same word. They are not, and the gap between them is where most of this chapter lives.

In 2015 a group of 270 researchers finished a project that had taken them three years. They took 100 studies published in three respected psychology journals and ran each one again, using the original materials where they could get them, with larger samples than the originals.

Ninety-seven of the 100 original studies had reported a clear positive result. When the studies were run again, 36 of them produced a clear positive result. The effects that did show up were, on average, about half the size the original papers had reported.

That was not a one-off. In 2018 a different team took 21 social science experiments that had been published in Nature and Science — the two most prestigious general science journals there are — between 2010 and 2015. They pre-registered their plans, had them checked by the original authors, and used samples about five times larger. Thirteen of the 21 replicated. Again the surviving effects were about half the original size.

And it is not old news. A large project published in 2025 attempted 274 claims from 164 papers, drawn from 54 journals in the social and behavioural sciences, published between 2009 and 2018. Around 55 per cent of the claims replicated. Weighted by paper, it was about half.

In Real Terms

Think of it as a shop. You walk in and every item on the shelf has a label saying it works. You buy a hundred of them, take them home, and about thirty-six do what the label said. Of those thirty-six, most work about half as well as advertised.

Now here is the part that matters for this series. You are not the person who bought a hundred. You are a reader who has been handed one item, in a chapter, by an author who found it convincing. You have no way of knowing which pile it came from.

Why does this happen? Not mainly because researchers cheat. It happens because of what gets published.

A study that finds nothing is boring. It is harder to publish, so people stop writing them up, so the published record is a filtered version of the research that was actually done. The filter selects for surprising results, and a surprising result is more likely than a boring one to be a fluke. On top of that, a researcher analysing a dataset has many small choices to make, and if he makes them one at a time while watching the result, he will drift towards the version that works without ever telling a lie.

Word Box

Publication bias: the tendency for interesting results to reach print while boring ones sit in a drawer. It does not require anyone to behave badly. It only requires editors, reviewers and authors to each prefer a finding to a non-finding, which they all do.

Why it matters here: it means the published literature on any question is not a sample of what is true. It is a sample of what was publishable, and those two things come apart most sharply on questions where a particular answer is exciting — which describes almost every question in this series.

So when Part 10 reported that a Norwegian team in 2016 replicated an American finding about regret after casual sex, that report was accurate and it was also doing something specific: it was noting that a result had survived a check, which most results in this field have not been given.

3.3 — The narrow sample problem

Now the second problem, which is worse for this series than the first.

A survey of the top journals in six branches of psychology found that 68 per cent of research subjects were from the United States, and 96 per cent were from Western industrialised countries. Those countries hold about 12 per cent of the world’s people.

In Real Terms

Imagine the human race as a village of a hundred people. Twelve of them live in the rich West. When the village’s scientists want to know what human beings are like, they interview those twelve — and about eight of the twelve are the ones from one particular household.

One commentary put it as a rate. A university undergraduate in a Western country is something like four thousand times more likely to end up in a psychology study than a randomly chosen person walking around outside.

The point is not that Western people are strange. It is that we do not know whether they are, because there is nothing to compare them with. And on the specific questions this series asks, the West is exactly the wrong place to sample from, because the West is the place where the rules under discussion have already been dismantled.

Part 10 is the clearest casualty. Its central table came from American survey data. It found that women with more sexual partners before marriage had somewhat higher divorce rates, that the pattern was not a straight line, and that nobody has produced a mechanism that explains it. Every word of that is about the United States.

Part 10 also ran a portability test on it and found the association essentially untested outside America, which is the correct thing to have done. But notice what remains. A reader in Ludhiana finished that chapter having read a great deal about the marriages of Americans, presented as the best available evidence about her own life, because it is the best available evidence about her own life. The alternative was not better evidence. The alternative was no chapter.

3.4 — The lying problem

The third problem is specific to this subject, and it is the one I find hardest to work around.

Almost everything in Part 10 rests on people answering questions about their own sexual history. There is a simple arithmetic check on whether they answer honestly.

How We Actually Know This

In any closed population, over any fixed period, the average number of opposite-sex partners reported by men and by women has to come out the same. Each encounter adds one to a man’s count and one to a woman’s. This is not a theory. It is bookkeeping.

In every national survey ever run, men report more. The gap has been measured in Britain, America and elsewhere for decades. So at least one of these is happening: men are overcounting, women are undercounting, or the surveys are missing a small group of women with very many partners.

In 2003 two psychologists tested it directly. They asked students about their sexual history under three conditions: connected to a machine they had been told was a lie detector, anonymously, and in a setting where the experimenter might see their answers. The sex difference in reported partners was largest under exposure, moderate when anonymous, and vanished under the fake lie detector — where women reported slightly more partners than men.

What this cannot show: which number is true. It shows that the number moves when the social pressure moves, which is enough to establish that at least one of the reported figures is not a measurement of behaviour.

Now put that next to the subject of this series. The rules being examined are precisely the rules that punish women for sexual history. A woman answering a survey in a society with those rules has an obvious reason to answer carefully.

So the measurement error is not random noise. It runs in the same direction as the thing being studied, and it is strongest exactly where the rules are strongest. That is the worst possible shape for an error to have. Random error makes findings weaker and blurrier. Error that tracks the independent variable can manufacture a finding out of nothing at all.

3.5 — What survives

Having made that case as hard as I can, I have to say honestly what it does and does not kill, because a reader who takes this chapter as permission to disbelieve everything has been misled in a new direction rather than rescued from an old one.

The Argument — does this sink the series’ evidence?

Three positions, and the distance between them is smaller than it looks.

Yes, substantially

Take away the studies that have never been replicated, the ones run only on Western samples, and the ones that depend on honest self-reporting about sex, and Part 10 is mostly gone, Part 5 loses its comparisons, and Part 4 loses most of its effect sizes. That is three of fourteen parts resting on a base the series itself has now described as unreliable. A book that leans on a literature must inherit that literature’s condition.

No, because of what it leaves standing

The problems named here attack one category of evidence: small correlational studies of behaviour, published in journals, based on self-report. They do not touch birth registers, census counts, court judgments, colonial land records, sex ratios at birth, school enrolment rolls, or the one randomised experiment. Every load-bearing conclusion in the series that survived to Part 14 rests on the sturdy categories. The panchayat result is a randomised experiment. The child sex ratio is a count of babies. The 79 per cent advertising spend is a parliamentary committee reporting on its own government’s accounts. None of those are going anywhere.

The damage is real but it is in a specific place

The failure is not that the series believed bad studies. It is that it presented sturdy and fragile evidence in the same voice, in the same typeface, in the same kind of sentence. A reader cannot tell from the prose whether a number came from a census or from 180 undergraduates in Ohio. The series has a box called “How We Actually Know This” designed exactly for that job — and used it perhaps seven or eight times per part, when the honest number would have been closer to one per claim.

Where things stand: the third position is right, and it is the objection I would least like to answer. The series did not overclaim in its conclusions. It underclaimed in its signposting, which produces the same effect on a reader while looking careful.

What would settle it: an evidence grade attached to every empirical claim in all fifteen parts — census, experiment, replicated study, single study, self-report — visible in the margin. That is a real piece of work and it would be the most useful correction anybody could make to this series. I have not done it.

Why people care so much: because on this subject both sides quote studies at each other constantly, and both sides are drawing from the same shallow pool. The traditionalist’s favourite finding about promiscuity and divorce and the reformer’s favourite finding about education and autonomy are frequently the same kind of object with the same weaknesses. Neither side wants to know that, because the weakness is symmetrical and the arguing is not.

One last thing before the summary, because it is easy to take the wrong lesson from a chapter like this. Nothing here licenses the response that the evidence is all rubbish and therefore anybody may believe what suits them. That response is not scepticism. It is scepticism used as a permission slip, and it always arrives on the side of whoever was already comfortable.

Remember This

The series rests on four kinds of evidence: counts, experiments, correlations and histories. This chapter’s objections hit the correlations hard, the experiments a little, and the counts and histories barely at all.

When 100 psychology studies were run again, 36 produced the same result, at about half the size. In a 2025 project across 54 journals, about half of the claims replicated. This is mostly caused by publication bias, not by cheating: boring results do not get printed, so the printed record is not a sample of what is true.

About 96 per cent of psychology’s research subjects come from countries holding about 12 per cent of the world’s people — and those are the countries where the rules in question have already been dismantled.

On this subject specifically, people misreport in a pattern. When students believed a lie detector was watching, the sex difference in reported partners disappeared. The measurement error runs in the same direction as the thing being measured, which is the worst possible shape for an error to have.

The series did not overclaim in its conclusions. It underclaimed in its signposting — sturdy evidence and fragile evidence were written in the same voice, and the reader had no way to tell them apart.

4The Scoreboard Was the Argument

Twice this series built a table and scored a side. Both times the result was decided before a single row was filled in — by the choice of what the columns were.

4.1 — The two scoreboards

Part 8 built a ledger of the equality project. Eighteen rows, each an outcome, with a column asking who the change had actually reached. Part 13 built a scorecard of both sides’ predictions against nine cases where history had run something like an experiment. Part 6 scored seven traditionalist arguments against four conditions. Part 14 scored four intervention shapes.

Those tables are the most quotable objects in the series. They look like the moment the arguing stops and the counting starts.

They are not that. A table of outcomes is an argument with the argument taken out, and the argument was: these are the things that count.

Part 13’s fourth hidden assumption said this out loud. Put fertility, marriage rates and care work on the scoreboard and the traditionalist wins comfortably. Put literacy, maternal mortality, income and life expectancy on it and it is not close the other way. Nobody has ever proposed a shared scoreboard, and neither side wants one, because a shared scoreboard would have to include the other side’s best columns.

Having written that, the series then went on using its own scoreboards for two more parts. That is the objection of this chapter, and it is mine, about my own work, and it lands.

Word Box

Proxy: a thing you measure because you cannot measure the thing you actually care about. Nobody can measure whether a child is well brought up, so a study measures school attendance. Nobody can measure whether a marriage is good, so a study measures whether it ended.

Why it matters here: a proxy is a substitution, and every substitution loses something. The danger is that after a while people stop discussing the thing and start discussing the proxy, because the proxy is the thing with numbers attached. Divorce is a proxy for marital failure. It is also, in a society where divorce is nearly impossible, a measure of how hard it is to leave.

That last example is not decoration. Part 10’s central table used divorce as the outcome, because divorce is what American surveys record. In a place where a woman cannot leave, a low divorce rate would show up on that table as a good result.

4.2 — Why some things get counted

Here is the deeper version of the problem, and it is not about bias. It is about what states can see.

Word Box

Legibility: a word borrowed from the study of how governments work. A thing is legible to a state if the state can see it, count it, and act on it from a distance — through a form, a register, a licence or a survey.

Why it matters here: a state cannot see whether a woman is respected in her house. It can see whether she is enrolled in school, whether she has a bank account, whether she died in childbirth, whether she voted. Everything the state can see becomes data, and data becomes evidence, and evidence becomes the scoreboard. What the state cannot see does not become nothing — it becomes something everybody has to argue about with anecdotes.

Now look at what that does to a series like this one, which is built almost entirely on published data.

What became a column in this seriesWhat never did
School enrolment and years of educationWhether school was a relief or another place to be watched
Female labour force participationWhether the job was chosen or was a second shift added to the first
Age at marriageWhether she wanted the marriage
Fertility rateWhether the children were wanted, and by whom
Maternal mortalityFear during a pregnancy that ended fine
Divorce rateMarriages that should have ended and did not
Sex ratio at birthWhat it is like to grow up known to be the second choice
Reported crimes and conviction ratesThe daily background rate of being managed
Contraceptive useWho decided
Seats held, offices wonBeing believed at home

Read the right-hand column again. Almost every item on it is a thing that, if you asked a woman what her life is like, she would mention before she mentioned anything on the left.

None of them are on any scoreboard in this series. Not because I decided they did not matter. Because there is no register anywhere that holds them, so there was no row to write.

The Hidden Assumption

Every scoreboard in this series — mine and everybody else’s — assumes something that nobody argues about because nobody notices it: that outcomes add up across people.

When a ledger says maternal deaths fell, it is adding the deaths that did not happen to a total. When it says literacy rose, it is adding readers. The arithmetic treats a gain to one woman and a loss to another as quantities that can cancel. That is what an average is.

But the thing under discussion in this series is not a quantity. It is an arrangement in which some people are constrained so that others get something. Averaging over it destroys exactly the information that the argument is about. A reform that lifts a hundred thousand urban graduates and leaves ten million rural women exactly where they were shows up on the ledger as progress. Part 8 put a “reached whom” column on its ledger precisely because I could feel this problem. One column is not a fix. It is an acknowledgement stapled to the side of a method that still averages.

The general form: a total is a decision about whose experience is allowed to cancel out whose. It looks like arithmetic, so it is never defended, so it is never attacked.

There is a reply to this, and it is not weak. Refusing to aggregate leaves you with anecdotes, and anecdote is the medium in which every bad argument about women has always been conducted — the cousin who ruined her life, the aunt who was perfectly happy. Numbers, for all their faults, are the only instrument that lets a claim be wrong.

4.3 — What a fair scoreboard would need

Suppose someone tried to build the shared scoreboard that Part 13 said nobody has ever proposed. What would it need?

It would need columns from both sides, agreed in advance, before anyone knew which way the numbers would fall. It would need each column weighted, and the weights argued for openly rather than smuggled in by which data happened to exist. It would need to report the spread and not only the average, so that a gain to a few could not hide a loss to many. And it would need somewhere to put the outcomes with no numbers at all — a column that says “not measurable, and here is what is at stake in it” rather than dropping them.

I did not build it. I am not sure it can be built. But notice that the reason it has never been built is not technical.

The Argument — could there be a neutral scoreboard?

If both sides could agree on what counts as a good outcome, the disagreement would become an empirical one, and empirical disagreements eventually end. So why has nobody done it?

It is possible and the obstacle is bad faith

Both sides can name outcomes the other side cares about. A traditionalist knows perfectly well that maternal death is bad; a reformer knows perfectly well that loneliness is bad. Nobody actually holds the view that their opponent’s entire column set is worthless. The reason no shared scoreboard exists is that each side does better in a world where the scoring is contested, because a contested scoreboard means you can always quote the column you are winning.

It is impossible and the obstacle is real

The two sides do not disagree about the weights on a shared list of goods. They disagree about what a person is for. If you hold that a woman’s flourishing is constituted by her role inside a family, and I hold that it is constituted by her freedom to leave it, we are not weighing the same items differently. We have different items. Asking us to agree on columns is asking one of us to lose the argument before the counting starts.

It is possible for a narrow band, and only there

There is a small set of outcomes almost nobody defends: girls dying of neglect, women dying in childbirth for want of a clinic, children married at twelve, acid, murder. That band is real and it is where every serious reform of the last century has actually happened. Outside it, the scoreboard idea collapses. So the honest conclusion is not “build a shared scoreboard” but “notice how narrow the shared part is, and stop pretending the wide part is an evidence problem.”

Where things stand: the third position is where I end up, and it is not a comfortable place, because it means most of this series has been applying evidence to a question that is not mainly evidential. The narrow band is where evidence decides things. Everything outside it is a disagreement about what a life is for, wearing the clothes of a disagreement about data.

What would settle it: somebody actually attempting the shared list and publishing the negotiation. Not the result — the negotiation. Where it broke down would be the most informative document anyone could produce on this subject.

Why people care so much: because whoever writes the columns wins. Every campaign on both sides is, underneath, a bid to make its own outcomes the ones that get counted — which is why so much energy goes into what should be measured and so little into measuring it.

The simplest way to see the whole problem is to take it out of this subject entirely.

In Real Terms

Think of two football teams who cannot agree on what a goal is. One counts balls in the net. The other counts passes completed. Both play hard for ninety minutes. Both walk off certain they won, and neither is lying.

Now notice that a spectator watching that match will spend the whole afternoon arguing about the players, when the only thing that decided the result was the rulebook nobody read out.

That is what four of the most quotable objects in this series are. They are not the moment the counting began. They are the rulebook, printed as though it were the result.

Remember This

This series built four scoreboards. They look like the point where arguing stops and counting starts. They are the opposite: a table of outcomes is an argument with the argument removed, and the removed argument was about which things count.

What gets counted is mostly what a state can see — school rolls, deaths, jobs, votes. That is called legibility. Everything a state cannot see, which includes most of what a woman would tell you about her own life, never becomes a row, so it never becomes evidence.

Every scoreboard assumes outcomes add up across people. That treats a gain to one woman and a loss to another as quantities that cancel — which destroys exactly the information the argument is about.

A shared scoreboard is probably possible only in a narrow band: deaths, child marriage, violence. Outside that band the two sides are not weighing the same goods differently. They have different goods.

Whoever writes the columns wins, and the fight over the columns is conducted almost entirely by people who think they are fighting about data.

5The Traditionalist’s Case Against This Series

Written as advocacy, the way Part 6 was. From here to section 5.5 I am arguing for a position, not reporting one. The objection is not the one you are expecting, and it is better than the one you are expecting.

5.1 — Not the objection you were expecting

The expected complaint is that the series is hostile to tradition. It is not, and a serious traditionalist would not waste breath on that charge. Part 6 gave the old rules a full advocate’s brief. Part 9 found that no country has ever reversed a fertility decline. Part 14 found that the reforming state’s flagship campaign spent most of its money advertising itself. On the scoreboard the series built, the traditionalist did not do badly.

That is exactly the problem.

The real objection is this: the series has been generous to our conclusions while quietly destroying our position. It agreed with us in several places and it did so on grounds we do not accept, in a language we did not choose, using a test we never proposed. A friendly verdict from a court with no jurisdiction is not a friendly verdict.

5.2 — You asked whether it works. We said it was right.

Fifteen parts have asked one question of every rule: what does it produce? Does it lower this, raise that, make her better or worse off, hold the birth rate up, keep the family together?

We never said any of that. Read what the tradition actually claims. It does not claim that modesty produces good outcomes. It claims that modesty is owed — that there is a right way to live, that a woman’s conduct is part of a larger order she did not invent and does not own, and that living rightly is what a life is for, whatever it produces.

Word Box

Consequentialism: the view that whether an action is right depends on the results it produces. If it makes things better, it is good; if worse, bad.

A constitutive good: something that is not a means to anything. It is part of what a good life is. Friendship is usually treated this way. Nobody defends friendship by showing it improves health outcomes, and anyone who did would have changed the subject.

Why it matters here: the whole series is written in the first language and the tradition speaks the second. Every chapter that carefully measures whether a rule “works” has, before its first sentence, ruled that rules must justify themselves by working.

Notice what happens once you accept that framing. Suppose our rules turned out to produce excellent outcomes on every measure. On the series’ own terms, we would then be right. But we would be right for a reason we reject — and a defence we reject is not a defence, it is a takeover with better manners.

Suppose instead the outcomes came back bad, as they do in several chapters. Then we are wrong. Either way, the question of whether a life owes anything to an order larger than itself has been settled without being asked, by the simple act of setting up an outcomes table.

Part 14 makes this vivid. It ran what it called a functional analysis of each rule — dowry as property transfer, seclusion as a substitute for policing, son preference as a substitute for a pension. Then it landed a test: paternity certainty was the stated function of chastity rules, a cheap test has existed for thirty years, and nothing changed. Therefore, it concluded, the stated function was never the function.

That is clever and it is an error. It only works if our rules were a hypothesis about paternity. They were not. Paternity is one reason among many that we have offered to people who asked for a reason in the language they would accept. That we could offer a practical reason to a practical questioner does not mean the practice was that reason. Ask a man why he does not read his wife’s letters and he may tell you it would damage trust. Solve the trust problem with a machine and he will still not read them. You have not discovered that his stated reason was a cover story. You have discovered that some things are not held for reasons of that kind.

5.3 — The tools are the worldview

Consider what counts as evidence in this series.

Surveys, in which a stranger asks a woman about her private life and writes the answer on a form. Effect sizes, which are averages over populations of interchangeable individuals. Randomised experiments, in which people are sorted by lot in order to see what a policy does to them. Ledgers of measurable outcomes.

Every one of those instruments was built by a civilisation that already believes what the series is supposedly investigating: that a person is an individual unit, that lives can be compared and totalled, that private conduct is a legitimate object of public measurement, and that an arrangement is answerable to its measured effects.

How We Actually Know This

You do not have to take this as philosophy. The series proves it against itself in Chapter Three of this very part.

The evidence there is that about 96 per cent of behavioural research subjects come from Western industrialised countries holding roughly 12 per cent of the world’s people. The series presents that as a sampling problem — the wrong people were asked.

Read it again as the objection it actually is. It is not only that the subjects were Western. The method was built there, tested there, and tuned to detect the things that vary there. An instrument developed to measure individual attitudes will find individual attitudes wherever you point it, including at a household that does not think of itself as a collection of individuals with attitudes.

What this cannot show: that the instrument is therefore wrong. A thermometer designed in Europe still measures temperature in Punjab. The traditionalist’s claim is that this instrument is more like a questionnaire than a thermometer, and that claim is not settled by either side.

5.4 — The unit error

Part 13 is the one we should be angriest about, and it is also the best chapter in the series, which is an unpleasant combination.

It went looking for natural experiments — places where the rules changed sharply enough that you could compare before and after, or compare across a line. Iran. Saudi Arabia. Afghanistan. The Indian village councils. It graded them, honestly, and it warned that “look at Sweden, look at Afghanistan” fails five of its own six tests.

But the whole exercise assumes that societies are the kind of thing that can be compared like this. That there is a variable called “the rules about women” which can be turned up in one place and down in another, with everything else held roughly steady, so that the difference in outcomes is caused by the setting.

There is no such variable. What the series calls “the rules” is not a dial fitted to an otherwise identical machine. It is woven into inheritance, worship, kinship, land, language, food, mourning and the names people carry. You cannot change it and hold the rest constant, because there is no rest. Part 13’s own third hidden assumption half-admits this — it says the reform is not the intervention, because what is delivered is text plus enforcement plus local reading plus substitution. Push that admission one step further and the comparison collapses entirely.

In Real Terms

Suppose you want to know what removing salt does to a dish. You can cook two identical pans and leave salt out of one. Clean comparison.

Now suppose you want to know what removing the third movement does to a symphony. There is no second orchestra playing an otherwise identical symphony. The movement is not an ingredient added to a base. The base is the movements.

We say a moral order is the second kind of thing and this series has treated it as the first, for fifteen parts, because the first kind is the kind you can put in a table.

And one more, briefly, because it costs the series more than it looks. Every part of this series has been written in English, drawing on a literature in English, published on the English-language internet. The traditions being examined transmit themselves in Punjabi, in Hindi, in ritual, in proverb, in the way a grandmother says a sentence, and almost none of that is in any database. The series then reports, honestly, that the evidence for the traditionalist position is thin. Of course it is thin. It was never written down in the language the search was conducted in.

5.5 — Stepping back out

That is the case at full strength. Now I stop arguing for it and say what I think it is worth, which is the same thing Part 6’s ninth chapter did to Part 6.

Two of those objections land and two do not.

The consequentialist objection lands. It is the strongest criticism in this entire document. The series does assume that an arrangement must justify itself by its effects, it never argues for that assumption, and a large body of serious moral thought rejects it. I do not agree with the tradition’s answer, but the objection is not about the answer. It is that I never put the question on the page. That is failure type four from Chapter One — the question was wrong — and it is the second time in this part that the same failure has been identified.

The instrument objection half lands. It is true that survey-and-effect-size research encodes a set of assumptions about persons. It is not true that this makes the findings worthless, and the traditionalist cannot have it both ways: the same instruments produced the fertility findings in Part 9 that the traditionalist quotes with enthusiasm. An instrument you accept when it agrees with you is not an instrument you have an objection to.

The Argument — does asking “does it work?” already decide the case?

This is the live one, and it has consequences beyond this series.

Yes — the framing is the verdict

To ask whether a practice produces good outcomes is to treat it as a means. Anything that is genuinely held as an end will fail that test, not because it is bad but because it was never entered in that competition. Every religious practice, every mourning rite, every promise kept at a cost would fail it too. A test that condemns all of those is not detecting badness. It is detecting a category.

No — because these rules make outcome claims constantly

The tradition does not confine itself to “this is right”. It says, loudly and daily, that girls who dress a certain way get attacked, that women who have several partners cannot stay married, that a society that stops having children will die. Those are predictions. Part 10 tested four of them. Once you make a prediction you have entered the outcomes competition voluntarily, and you do not get to withdraw when the result arrives.

Both, and they are addressed to different audiences

The tradition speaks two languages. To the faithful it says this is right. To the doubting, the young and the courts it says this works. The series audited the second language and left the first untouched. That is a legitimate scope — but it should have been declared as a scope in Part 1 rather than discovered in Part 15.

Where things stand: the third position is correct and it costs me something to write it. The series never claimed to settle whether the rules are right. It claimed to test what they do. That is a narrower project than fifteen parts and a thousand pages make it feel, and the size of the thing has been doing work that the argument has not earned.

What would settle it: nothing empirical. This is the boundary of what evidence can reach, and one useful function of an evidence-based series is to show exactly where its own instrument stops working.

Why people care so much: because whoever gets to choose the language of a public argument has usually won it. “Does it work?” is the language of the modern state, which is why reformers use it, and why traditions that once spoke only of what is owed have learned to answer in it — and lose something every time they do.

So the traditionalist walks away with one direct hit and one glancing blow, which is a better afternoon than he is usually given. The direct hit is not about any finding in this series. It is about the sentence the series never wrote down on page one.

Remember This

The traditionalist’s real objection is not that the series was hostile. It is that the series was generous on grounds the tradition rejects. A favourable verdict from a court with no jurisdiction is not a favourable verdict.

The series asked of every rule: does it work? The tradition claims something different — that certain conduct is owed, whatever it produces. By setting up an outcomes table, the series settled that question without asking it.

The instruments themselves — surveys, effect sizes, randomised trials, ledgers — were built by a civilisation that already holds the individual view of persons. That objection half lands: it is true about the instruments, and it is not available to anyone who quotes the same instruments when they agree.

Part 13’s comparisons assume there is a dial called “the rules about women” that can be moved with everything else held steady. There is no such dial. But the tradition makes outcome predictions constantly, and having made them, it cannot withdraw from the test when the results come in.

The series audited the language of “this works” and never touched the language of “this is right” — and it should have said so on page one instead of page nine hundred.

6The Feminist’s Case Against This Series

Advocacy again, the way Part 7 was written as an audit. From here to section 6.5 I am arguing for a position. The charge is not that the series was unfair. It is that fairness was the instrument.

6.1 — Not “you were unfair”

The weak version of this objection is that the series took the traditionalist side too seriously. Nobody serious is going to argue that, and it would be wrong. Part 11 went after men as the missing comparison group. Part 12 showed that two independent sets of criminals have priced the double standard into their business models. Part 14 showed the reforming state spending its budget on advertising itself. The series is not soft.

The charge is different, and it is structural. The series adopted a rule — every side gets its strongest case — and applied it evenly to two parties who are not evenly placed. Applying the same rule to unequal parties does not produce equal treatment. It produces an advantage, and the advantage is always in the same direction.

Three mechanisms, and none of them require the author to be biased.

6.2 — The first mechanism: uncertainty has a home

Chapter Two already conceded the core of this and then moved on quickly, so let me hold it still.

Whoever proposes a change carries the burden of proof. That is not a rule anyone voted for. It is simply what happens. So every honest finding of “we do not know”, “the evidence does not travel”, “this has not been tested outside America” is a point scored, and it is always scored for the same team.

Count how often those sentences appear across fifteen parts. They are the house style. The series is proud of them — Chapter Ten of every part is a two-page list of what we do not know, and it is described in the method as the most valuable chapter in the document.

It may well be the most valuable. It is also, in a world where the rules are currently in force and enforced by families rather than by argument, a fifteen-times-repeated statement that nothing has been proved. A father does not need to win the argument. He needs it to remain open.

Word Box

False balance: giving two positions equal space and equal treatment when the evidence behind them is not equal. It is a well-known failure in journalism — a hundred scientists and one contrarian, one on each side of the desk.

Why it matters here: the version in this series is subtler and harder to see, because the evidence really is thin on both sides. Where genuine uncertainty is presented evenly, the presentation is accurate and the effect still favours whoever benefits from delay. Accuracy and neutrality are not the same property, and this series has repeatedly claimed the second by demonstrating the first.

6.3 — The second mechanism: the funded side has more evidence

Now the deeper one.

The series’ rule says: give each side its strongest case. But a side’s strongest case is not a fact about how good its position is. It is a fact about how much work has been done on that position — how many researchers studied it, how many surveys asked about it, how many centuries of writing there are to draw on.

What has been studied? The consequences of women’s behaviour. There are decades of research on partner counts and divorce, on age at marriage, on female employment and fertility, on what happens when a woman does something. That is the literature Part 10 is built out of, and it is thick.

What has not been studied? Almost everything in Part 11. That part had to be invented from scratch because the comparison group — men — had barely been examined. Part 11 is a good chapter, and the reason it is a good chapter is that the field was empty. The series treated that emptiness as a discovery. It is also a measurement of where the research money and the research attention have gone for a century.

How We Actually Know This

You can see this inside the series without leaving it. Part 10 tests four traditionalist claims about what happens to a woman, using a large American survey, a Norwegian replication, and a body of work going back decades.

Part 11 tests the equivalent claims about men and finds nothing comparable. Its evidence is a marital rape exception, a set of population projections, two censuses and a Supreme Court judgment — a legal and demographic record, not a behavioural literature, because a behavioural literature was never built.

What this cannot show: whether the imbalance is because researchers were biased, or because funders were, or because the question about women is genuinely more studied everywhere for reasons unrelated to sexism. All three predict the same empty shelf.

What it does show, definitively: a rule that says “give every side its best evidence” hands more to the side whose evidence exists. That is not a claim about anyone’s motives. It is a claim about shelves.

6.4 — The third mechanism: Part 6 exists

Part 6 of this series is titled The Case for the Old Rules. It is sixty-five pages. It is written as an advocate’s brief, with seven arguments developed at full strength, and its ninth chapter scores them.

Look at what that object is when it leaves the author’s hands.

It is the most articulate, best-evidenced, most respectable defence of the old rules that exists in accessible English, on a free website, published by an Indian writer under his own name. It is better than anything the tradition has produced for itself in decades, because the tradition does not usually need to argue — it has enforcement instead.

The series’ defence is that Chapter Nine of Part 6 scores those arguments and finds most of them conditional or weak. But a document does not travel with its scoring chapter attached. It travels in screenshots. Nobody has ever forwarded a scorecard.

And the arguments in Part 6 are not abstractions. Their real-world content is a girl not being allowed to go somewhere, a marriage arranged at a speed she cannot object to, a phone checked, a cousin sent along. When the series writes those arguments at full strength, it is not conducting a thought experiment. It is manufacturing ammunition and then, in a later chapter, expressing reservations about the ammunition.

The Hidden Assumption

Everything in this part so far — my objections and the traditionalist’s — has assumed something neither side would think to state: that the reader is a neutral third party who has not yet decided.

The whole apparatus depends on it. Steel-manning both sides is useful for someone weighing them. A verdict box saying “nobody knows” is useful for someone waiting to find out. An honest list of what we do not know is useful for someone who intends to update.

But who actually reads a document like this? Not a jury. The readers are inside the thing being described. Some of them are women living under exactly these rules, for whom “the evidence is unclear” is not information but a description of the argument they have already lost at home. Some are men who make these decisions about other people, for whom Part 6 is not one chapter of fifteen but a supply of sentences. Some are already convinced in one direction and are reading for material.

For none of those people is the document a source of information. It is an instrument in a fight they are already in. Fifteen parts have been written as though addressed to an undecided reader who, on the evidence of who actually reads long documents about contested subjects, may barely exist.

The general form: a text written for the undecided will be used by the decided, and its careful even-handedness becomes, in their hands, a supply depot open to both sides — with more shelves stocked, as section 6.3 showed, on one side than the other.

That is the objection at full strength. It is the one I have found hardest to answer, and I have not answered it. I have only got as far as being able to state it clearly.

6.5 — Stepping back out

Now out of the advocate’s voice, and the same treatment Part 7’s ninth chapter gave the feminist claims it audited.

The mechanism objections in 6.2 and 6.3 land completely. They are not about my intentions, they do not require me to have been careless, and I cannot answer them by pointing at anything in the text. Honest uncertainty helps the status quo. A strongest-case rule rewards the better-documented side. Both are true, both are structural, and neither has a fix inside the method.

The Part 6 objection is where I want to push back, and the argument is worth having in full.

The Argument — should the strongest case for a coercive position be written down?

This is the practical question underneath the whole method, and it is not confined to this series.

No — you are supplying the other side

Arguments are not inert. A well-made case for restricting women’s movement is used to restrict women’s movement, by people who did not have those words before. The author’s disapproval does not travel with the text. Nobody who was going to use Part 6 will read Part 6’s ninth chapter. Writing it may be the single most consequential thing this series has done, and the consequence is not the one the author intended.

Yes — the alternative has been tried and it failed

The case for the old rules was not invented by Part 6. It is the operating assumption of most households on the subcontinent, transmitted without ever being written down, which is precisely why it has never been examined. An argument that is never stated cannot be answered. Writing it down converts an atmosphere into a claim, and a claim can be tested — which is exactly what Part 10 then did to it, and what Part 14 did when it showed the stated function was solved thirty years ago and nothing moved. You cannot land that blow on a position you have refused to state.

Yes, but not for free

The dispute is not really about whether to state it. It is about packaging. A brief that can be forwarded on its own is a different object from an argument embedded where its answer cannot be detached. Part 6 was published as a standalone PDF with a cover, a title, and its own web page. That was a design decision and nobody made it as a decision.

Where things stand: the third position is right and it is a criticism of the format, not of the method. The argument for stating the case survives. The argument for stating it as a separately downloadable sixty-five-page document with an attractive cover does not survive as easily, and I made that choice without ever once thinking about it as a choice.

What would settle it: knowing how Part 6 has actually been used. That is knowable in principle — where it is linked, what is quoted from it, whether the scoring chapter ever appears alongside — and I have not looked. Not looking is itself a decision I should not be allowed to describe as neutral.

Why people care so much: because this is the oldest question in public argument and it has no stable answer. Every serious person who has ever written down an opponent’s best case has had to decide whether they were clarifying a fight or arming one, and nobody has ever been able to check.

Two chapters, two sets of objections, and notice how differently they behave. The traditionalist attacks the framing and lands once, hard. The feminist attacks the mechanics and lands three times, and none of the three can be fixed by writing better. That difference is itself information about what kind of document this series is.

Remember This

The objection is not that the series was unfair. It is that the same rule applied to unequally placed parties produces an advantage, always in the same direction.

Three mechanisms. Uncertainty has a home: every honest “we do not know” is a point for whoever benefits from nothing changing. The funded side has more evidence: a strongest-case rule rewards the position that has been studied, and one position has been studied for a century while the other left an empty shelf that Part 11 had to fill by hand. Part 6 exists: the best-argued defence of the old rules in accessible English is now a downloadable document with a cover, and the chapter that scores it does not travel with it.

Underneath all of it sits the assumption that the reader is an undecided third party. Most readers of a document like this are inside the fight it describes. For them it is not information. It is a supply depot, with more shelves stocked on one side.

The first two objections land completely and have no fix inside the method. The third is a criticism of packaging, and it is correct about the packaging.

Fairness is a rule about the writer. It is not a property of the effect.

7Where I Changed My Mind

Six places where the writing moved me. Each one is a correction to something I believed at the start — and the version of me that believed those things is the one who chose the questions.

7.1 — Why this chapter is in the document at all

A writer telling you he changed his mind is doing something suspicious. It is a claim about his own honesty, made by him, with no way for you to check it. It also flatters him: only an open-minded person changes his mind, so the confession is a compliment wearing plain clothes.

So I want to be careful about what this chapter is claiming. It is not claiming that I am open-minded. It is producing a list of specific, dated, checkable places where an earlier part of this series says one thing and a later part says another, and drawing an unfriendly conclusion from the pattern.

How We Actually Know This

The evidence here is unusually good for a claim about a writer’s private beliefs, because the belief was published before the test was run.

Each of the six items below has this shape: an earlier part is on record with an expectation, in print, on a public site, with a date. A later part is on record with a finding that does not match it. The gap between the two is the evidence. I cannot go back and edit the earlier part without leaving a corrected edition behind, and the standing rule of this series is that corrected editions say what was changed and who caught it.

What this cannot show: why the belief changed. Reading a study and being persuaded looks exactly like reading a study and finding it convenient. Nobody has access to the difference, including me, which is why Chapter One told you to discount every sentence in this part beginning “I was surprised”.

7.2 — The six

One. I expected the partner-count finding to be a clean story. Before Part 10 I assumed the association between a woman’s number of partners before marriage and her later divorce risk would either be a straightforward causal effect or a straightforwardly spurious one. It is neither. The pattern is not a straight line at all — the divorce rate does not simply rise as the number rises, and the researcher who assembled the best version of it has published that no proposed mechanism explains its shape. I expected to resolve that chapter. It resolved into a shrug that is more interesting than either answer would have been.

Two. I expected enforcement to be male. I began Part 11 expecting to describe fathers and brothers policing daughters. The daily reality in the data is that the enforcement is largely administered by women on women. In the Indian household survey work, a mother-in-law’s view of contraception predicts a couple’s use of it better than the wife’s own view does. The lethal end of enforcement is delegated to the most legally expendable male in the family. That is a different machine from the one I sat down to describe.

Three. I expected an asymmetric rule to imply an asymmetric fact. If a rule falls on daughters and not sons, the obvious reading is that somebody believes something different about daughters. The better explanation turned out to be much duller and much worse: the rule falls where it can be enforced. A household can reach a daughter at home. It cannot reach a son in another city. That single sentence reorganised Part 11 and it did not come from me — it came from noticing that the pattern of enforcement tracks reach, not belief.

Four. I expected natural experiments to be reasonably common. Part 13 was planned as a survey of many cases where history had run the trial for us. After grading them against six tests, exactly one genuine randomisation exists in this entire subject: India’s 1993 reservation of village council seats, which rotated by village serial number. One. The rest are comparisons, some of them good, none of them experiments. I had assumed the shelf was half full. It has one item on it.

Five. I expected persuasion campaigns to be weak. I did not expect them to fail as a category. India’s flagship campaign for the girl child released ₹446.72 crore to states between 2016 and 2019, and 78.91 per cent of it went on media advocacy — a parliamentary committee reported that to Parliament in December 2021. The campaign to change minds spent four rupees in five telling people the campaign existed. Part 14 ended up classifying “tell people to stop” not as a weak intervention but as a category error, and I did not begin that part expecting to write that sentence.

Word Box

Crore: ten million, in the Indian way of counting. ₹446.72 crore is about 4.47 billion rupees. Indian budget figures are almost always written this way, and a reader outside India will misread them by a factor of ten million if nobody says so.

Six. I expected money to move behaviour. Korea has spent more than two hundred billion dollars on raising its birth rate. Hungary has spent around 5 per cent of its national output. Both got real effects, both small, and Hungary’s rate has fallen back to roughly where it started. Cash is the intervention every government reaches for. It is also the one with the shortest memory: the effect stops when the money stops.

In Real Terms

Six changes across fourteen parts is not a lot. It is roughly one every two parts, or one every 20,000 words.

Put it the other way and it sounds worse. If you had asked me six specific questions on the day I started, I would have got all six wrong. Not slightly wrong — wrong in the direction of the ordinary, expected answer, which is exactly the direction a person is wrong in when he has not looked yet.

7.3 — What that says about Part 1

Here is the part that costs something.

Part 1 did not just open the series. It set the questions for everything after it. It defined the portability test that Part 10 and Part 13 both run. It decided that the series would examine what the rules are for and what happens when they change. It agreed the shape of all sixteen parts.

And Part 1 was written by the version of me who would have got all six of those questions wrong.

That is not a small observation. Every later part could be perfectly reasoned and the whole structure would still be carrying the fingerprints of a set of expectations that the evidence went on to contradict six times. The corrections are visible, because they happened inside chapters and got written up. The framing errors are not visible, because framing does not produce a testable claim. It produces a table of contents.

Look back at the list and notice what kind of error they all are. Every one of them is me expecting the simpler, more individual, more belief-driven version of the truth. A clean causal effect rather than an unexplained shape. Men enforcing rather than a whole household machine. A belief about daughters rather than a fact about reach. Plenty of evidence rather than almost none. Persuasion working a little rather than not being an intervention at all.

That is a consistent bias and it has a name in the ordinary sense: I kept expecting the world to be made of individuals with reasons. Part 14’s second hidden assumption says beliefs are downstream of structures. I now think that is right. But I only got there by being wrong about it six times in public, and the series is built on a frame chosen before any of that happened.

The Argument — does a writer who changes his mind become more reliable or less?

This matters beyond me, because it decides how you should read anybody’s long project.

More reliable

A writer whose conclusions never move is either extraordinarily lucky in his priors or is not testing anything. Visible correction is the only externally checkable sign that evidence is doing work. You cannot see inside a writer’s head, but you can see whether Part 10 contradicts Part 1, and you can see whether he printed it.

Less reliable

Six documented reversals in fourteen parts is a measured error rate for this author on this subject, and it is high. It applies to the claims he has not yet had a chance to be wrong about — which, by definition, are the claims still standing at the end. A person who has been wrong six times out of six is not a person whose framing you should trust; he is a person whose framing has simply not been tested yet, because framing never is.

Neither — it measures the field, not the man

All six errors are in the direction of the popular, individual-level explanation. That is not a personal defect. It is what an educated reader of the English-language literature on this subject would believe, because that is what the literature emphasises. The reversals measure the gap between what gets written about and what turns out to hold. Anyone doing this honestly would have produced roughly the same six.

Where things stand: the second and third positions are both true and they are not in conflict. The error rate is real and it does transfer to untested claims. It also is not idiosyncratic — it is the standard prior of anyone who reads this material in English, which makes it more dangerous rather than less, because a shared error does not feel like an error. It feels like context.

What would settle it: someone else running the same fourteen questions from a different starting position and seeing whether they get corrected in the same six places. Nobody is going to do that, so this stays unresolved.

Why people care so much: because “he changed his mind” is used as a trump card by both sides in every public argument — as proof of honesty by supporters and proof of unreliability by opponents — and it is genuinely evidence for both.

7.4 — The inference I do not like

Put the two halves together and you get something I would rather not write down.

The six corrections are the visible errors. Visible errors are the ones that happened at the level of claims, because claims are the level at which the evidence can push back. Nothing pushed back on the framing, the choice of questions, or the decision about what a good outcome would be — not because those were right, but because nothing in a research process ever tests them.

So the fair summary is not “the author was corrected six times and is now better calibrated”. It is: the author was corrected six times at the only level where correction was possible, and the level that was never tested is the level that determined the shape of all fifteen parts.

That is Chapter One’s hidden assumption arriving with a bill attached.

Remember This

Six documented reversals across fourteen parts: the partner-count shape, who actually enforces, reach rather than belief, how few real experiments exist, persuasion as a category error, and money’s short memory. All are checkable, because the expectation was published before the test.

Every one of the six is an error in the same direction: expecting the world to be made of individuals acting on beliefs, rather than households acting within reach of each other.

Part 1 set the questions for all sixteen parts, and Part 1 was written by the version of me that got all six wrong. The corrections are visible because they happened at the level of claims. The framing was never tested, because framing does not produce a testable claim.

Changing your mind is evidence of honesty and evidence of a measured error rate at the same time, and the error rate transfers to everything still standing.

I was corrected at the only level where correction was possible, and the level that decided the shape of the whole thing was never in play.

8What Even-Handedness Cost

The rule that every side gets its strongest case is the thing I am proudest of in this series. This chapter puts a price on it — four prices, and one of them I did not see until I wrote this down.

8.1 — What the rule bought

Before the bill, the purchase, because a chapter that only lists costs is not an audit either.

The strongest-case rule bought three things and they are not small.

It made the series usable by people who disagree with it. A traditionalist reader can get through Part 10 without being insulted, which means he can get as far as the finding that the marriage-market damage does not travel — and that finding is a problem for him. An argument nobody hostile will read has an audience of the already-convinced, which is another way of saying no audience.

It forced the writing to be right. You cannot state an opponent’s best case and then knock it down with a sentence, because you have just made the sentence look thin. Half the useful findings in this series exist because a lazy answer stopped working once the objection was written out properly.

And it produced the one thing public argument almost never produces: a clear view of where the disagreement actually sits. Part 5’s four competing accounts. Part 6’s scorecard. Part 13’s table of both sides’ predictions. None of those could have been built by someone who had decided the answer first.

Now the bill.

8.2 — Price one: articulate positions win

“Give every side its strongest case” contains a hidden entry requirement. To have a strongest case, you have to be a side. To be a side, somebody has to have written your position down.

Chapter Six put one version of this: the funded position has more evidence. Here is the wider version.

The positions that appear in this series are the ones with literatures — the religious tradition, the reforming state, academic feminism, evolutionary psychology, demography. Each of those has advocates, books, journals and money.

Now name the positions that do not appear. The woman who thinks the rules are stupid and complies anyway because complying is cheaper than the alternative, and who has no interest in either camp. The mother-in-law who enforces and would not describe herself as enforcing. The young man in a village with no marriage prospects, who appears in Part 11 only as a risk to other people. None of these are “sides”. They have no journals. So they enter this series as objects of study and never as arguments.

That is not a fixable flaw in the method. It is what the method is. A rule that says “represent every position at full strength” will always be a rule about positions that have been articulated, and articulation is expensive.

Word Box

The default: what happens if nobody does anything. In a form, it is the box already ticked. In a system of social rules, it is the arrangement currently running.

Why it matters here: defaults win far more often than their arguments deserve, because they are the outcome of inaction as well as of decision. Every argument that ends inconclusively is an argument the default has won. This is not a claim about persuasion. It is arithmetic about how outcomes are produced when nobody is persuaded of anything.

8.3 — Price two: refusing to advise is a position

Part 10’s ninth chapter refused to give advice. It said so openly, and it explained why refusing is itself a position rather than an absence of one.

I want to go further than that chapter did, because it stated the problem and then moved on as though stating it had handled it.

Consider what a young woman is actually doing when she reads Part 10. She has a decision in front of her — about a relationship, a phone, a photograph, a city, a marriage. She has read seventy-six pages establishing that the traditionalist’s warning does not survive a portability test, that the wellbeing effects reverse depending on what her community believes, and that the divorce association is essentially untested outside America.

Then the author declines to tell her what he thinks she should do.

What has she got? She has got an accurate map of a landscape and no direction. And the landscape she is standing in already has a direction built into it, because her family has a view, her neighbourhood has a view, and the cost of doing nothing is zero while the cost of doing something falls entirely on her.

So the refusal does not leave her free. It leaves her exactly where she was, with better information about why she is there. On any question where the default is one-sided, a writer who declines to advise has not abstained. He has voted, quietly, for whatever was going to happen anyway.

The Argument — should this series have given advice?

The method says do not pick a side for the reader. The objection is that the reader is not standing on neutral ground.

Yes, it should have

The series has spent a thousand pages establishing what the evidence does and does not support. Withholding the conclusion at the last step is not humility, it is a refusal to be accountable for the one sentence anyone could act on. Doctors do not present a patient with a survival table and decline to recommend. And the cost of not advising is not shared evenly: it falls on the person with the decision, who is almost never the person with the power.

No, and the asymmetry is exactly why

The author is a man who has never been subject to these rules, writing to women who will bear every consequence of following his advice. He carries none of the risk. Advice under that arrangement is not generosity, it is a stranger spending someone else’s safety. Worse, the evidence in this series does not support advice — Part 10’s own finding is that the wellbeing effects reverse depending on the community. There is no general answer, so any general answer would be invented, and an invented answer delivered with a thousand pages of authority behind it is more dangerous than no answer.

Advise on the reasoning, not the decision

The dispute assumes advice means “do this”. A third option exists and the series never took it: tell the reader which questions decide her case. Not “leave” or “stay”, but: your outcome depends far more on what your particular community believes than on anything in the research; the warning you have been given is a prediction and here is how to check whether it has ever come true near you; the people advising you carry none of the cost. That is directive without being presumptuous, and it is actionable.

Where things stand: the third position is right and the series did not do it. Not once in fifteen parts is there a section addressed to a reader with a decision, telling her which of the findings actually bear on her situation and which are background. That is a real gap and it is the most useful thing anyone could add to this work.

What would settle it: nothing evidential. This is a question about what a writer owes a reader, and the answer depends on who you think the reader is — which is precisely what Chapter Six’s hidden assumption says nobody knows.

Why people care so much: because “I am just presenting the evidence” is the most comfortable position available to anybody writing about other people’s lives, and it is comfortable exactly in proportion to how little the writer stands to lose.

8.4 — Price three: it can be quarried

A document that gives every side its strongest case is a document from which any side can extract support.

This is not hypothetical. Part 6 is a defence of the old rules, written well. Part 9 contains the finding that no country has reversed a fertility decline, which is the single most quoted fact in pronatalist politics. Part 4 contains effect sizes for sex differences. Part 7 contains an audit of feminist claims that finds some of them weakly supported.

Every one of those is honest. Every one of them is also a paragraph that can be lifted, screenshotted and used with the surrounding argument removed. The series’ answer is that the arguments are all in the same document and a reader can see the whole. That answer assumes the whole gets read. Chapter One already worked out roughly how many people finish a fifteen-part series, and the answer was: very few.

The honest description is that this series has produced a well-stocked quarry, marked its own load-bearing stones carefully, and has no control over what anybody builds.

8.5 — Price four: the length is the filter

This is the one I did not see until I wrote this part.

The method’s central claim is that hard things can be written so that anyone can follow them. Short sentences. Everyday words. Every term explained where it appears. No assumed knowledge. I believe that claim and this series demonstrates it: there is no sentence in the fifteen parts that requires a degree.

But simplification is not compression — the method says so explicitly. The simple version is longer. Spend enough words buying comprehension and you have produced a thousand pages, and a thousand pages has an entry requirement of its own. It is not vocabulary. It is time.

In Real Terms

Twenty to twenty-five hours of reading. Now think about who has that.

A woman doing paid work and then a second unpaid shift at home has, by every time-use survey ever run in India, less discretionary time than almost anyone else in the country. She is the person this series is about. She is the person least able to read it.

The reader who does have twenty-five spare hours, an English reading habit and an interest in evidence is, on average, comfortable, educated, urban and male. The prose was made accessible. The object was not.

So the accessibility of the method is real at the level of the sentence and undone at the level of the artefact. Every one of the eight rules of the voice is aimed at the sentence. Nothing in the method is aimed at the total. Fifteen parts of scrupulously simple English add up to something only a particular kind of person can consume, and that person is not the subject of the book.

Remember This

The strongest-case rule bought three real things: it made the series readable by people who disagree, it forced the arguments to be right, and it showed where disagreements actually sit.

Price one: to have a strongest case you have to be a side, and being a side requires somebody to have written you down. The people in this series with no journals appear only as objects of study.

Price two: refusing to advise is not abstention. Where the default is one-sided, declining to advise votes quietly for whatever was going to happen anyway. The series never once addressed a reader with an actual decision in front of her.

Price three: an even-handed document is a quarry. Every side can lift a stone, and the answer “read the whole thing” assumes a reader who does.

Price four, and it is the one I missed: the prose was made accessible and the object was not. Twenty-five hours is an entry requirement, and the person this series is about is the person in the country with the least spare time.

9What Would Have To Be True For Me To Be Badly Wrong

Nine chapters of objections, scored. Then the four ways this series could fail in a way that actually matters — and the assumption sitting underneath the whole confession.

9.1 — Four ways to be wrong that would matter

Chapter One set a test: an objection has to say what becomes worthless if it is right. Here are the four failures that would do real damage, stated so that you could in principle check them.

Failure one: the causal arrow in Part 9 runs backwards. Part 9 treats falling birth rates as something that happens when women get options. Suppose the arrow mostly runs the other way — that the collapse in fertility is driven by housing costs, insecure work and delayed adulthood, and women’s changing choices are largely downstream of those. Then Part 9 spent ninety-two pages attributing to a cultural shift something caused by an economy, and the entire pronatalist argument it takes seriously is aimed at the wrong target. What would show it: a country where the material conditions were fixed and the birth rate recovered while the cultural change stayed. No country has produced this, which is either evidence for Part 9 or evidence that nobody has tried the experiment.

Failure two: the enforcement thesis is an artefact of who gets studied. Part 11’s spine is that the rule falls where it can be reached. The evidence is that households police daughters at home and cannot police sons in other cities. But household surveys are conducted in households. A method that visits homes will find the control that happens in homes. What would show it: good data on how men in other cities are actually constrained — by remittance expectations, by marriage brokerage, by the threat of being cut off. If that constraint turned out to be heavy and simply invisible to household surveys, Part 11’s central claim shrinks to a claim about measurement.

Failure three: persuasion works on a lag longer than any study. Part 14 grades “tell people to stop” as a category error, on evidence drawn from campaigns measured over five to ten years. Almost every large moral change in history — on slavery, on child labour, on public smoking — took two or three generations, and the persuasion phase looked useless throughout. What would show it: a case where a campaign produced no measurable effect for twenty-five years and then a large one, with the mechanism traced. I do not know of one that has been properly documented, and the absence may be about the length of research funding rather than about the world.

Failure four, and the serious one: the series is a way of not acting. Everything in it is true, and its effect on a reader is to convert a situation that calls for a decision into a subject that calls for further reading. What would show it: nothing, from inside. This is the failure that cannot be detected by the person committing it, which is what makes it worth putting first in your mind and last in the document.

9.2 — The scoreboard

Chapter One set three verdicts. An objection lands if I cannot answer it without changing a claim or admitting something rests on a choice. It is survivable if it is true and the series already says so in the text. It fails if answering it only requires pointing at what the series actually says.

The objectionVerdictWhat answering it would take
The functional question was run without its mirror — what is this costing now, and who pays?LandsWriting the missing part at the same depth. My expectation is it changes emphasis everywhere and conclusions nowhere, which I cannot verify from here.
Honest uncertainty is not neutral — every “we do not know” scores for the arrangement in forceLands, no fixNothing available. This is structural. It is a property of writing truthfully about a rule while the rule is running.
A real woman’s injury was converted into a research subjectLandsNothing. It is done. Part 12 named half of it; this part names the rest.
Sturdy and fragile evidence were written in the same voiceLandsAn evidence grade on every empirical claim in all fifteen parts. Real work, and the single most useful correction anyone could make.
The scoreboards’ columns decided their results, and totals let one woman’s gain cancel another’s lossLandsReporting spread as well as average, and a column for outcomes with no numbers. Part 8’s “reached whom” column was a gesture at this, not a fix.
Consequentialism was assumed, never argued — the series tested “does it work” and called it neutralLands hardestDeclaring the scope on page one: this series audits outcome claims and does not touch claims about what is owed.
The research instruments encode the worldview under examinationHalf landsNot available to anyone who quotes the same instruments approvingly, which both sides do constantly.
Societies are not comparable units, so Part 13’s cases are not experimentsSurvivableNothing — Part 13 says this itself, sets six tests, and fails five of them on the popular comparisons.
A strongest-case rule rewards whichever position has been studied and fundedLands, no fixNothing inside the method. The empty shelf Part 11 had to fill by hand is the measurement of it.
Part 6 is a downloadable, well-designed brief for coercive rules, and its scoring chapter does not travel with itLands, on packagingPublishing the case and its audit as one inseparable object. A design decision I never made as a decision.
The document assumes an undecided reader; most readers are inside the fightLandsKnowing who actually reads and quotes it. Knowable in principle. I have not looked.
Six reversals show the framing was set by the least informed version of the authorLandsSomeone else running the same questions from a different start. Nobody will.
Refusing to advise is a vote for the defaultLandsA section telling a reader with a decision which findings bear on her case. Never written, in fifteen parts.
Fifteen parts of simple English add up to an object only the comfortable can readLandsA short version. One that is genuinely short, not a summary that points back at the long one.

Fourteen objections. Twelve land, one half lands, one is survivable, and none fail outright. That last figure should bother you, and it bothers me.

A prosecution in which nearly every charge sticks is not a rigorous prosecution. It is a document written by someone who selected charges he was willing to concede. Chapter One promised that an honest audit includes the charges that fail, and this one produced almost none — which tells you the selection happened earlier, in the choosing, where selections always happen and never show.

9.3 — The one that survives everything

Strip the fourteen down and one combination is left standing that nothing in the series answers.

Part 14 found that explanation is the weakest lever available for changing an arrangement like this. Part 15, Chapter Eight, found that the artefact takes twenty-five hours to read, which selects for a reader who is comfortable, educated and not subject to the rules. Part 15, Chapter Six, found that the readers who do arrive are mostly not undecided.

Put those three together and you get a specific, unflattering prediction: this series will be read mainly by people it cannot help, in a format that works least, on a question its own findings say is not settled by argument.

Every step of that is drawn from the series’ own results. I cannot refute it with anything in the series, because it is built out of the series.

The Hidden Assumption

Everything in this part — every objection, every concession, every “this lands” — rests on a premise so comfortable that I did not notice it until the scoreboard came out at twelve to nothing.

The assumption is that being able to state an objection shows you are not subject to it.

It is what makes self-criticism feel like progress. I named the burden-of-proof problem, so I am no longer captured by it. I described the way an even-handed document supplies both sides, so my document is now less of a supply depot. I identified that the length filters the audience, and having identified it, I have somehow answered it.

None of that follows. A confession is a description, not a repair. The burden of proof still sits where it sat. Part 6 is still downloadable. This part is 20,000 words long, which makes the length problem it diagnoses worse by exactly 20,000 words. There is no sentence anywhere in this document that changes a single thing about the fourteen parts it audits, and I have not committed here to changing any of them.

The general form: naming a bias is the cheapest possible response to it, and it produces the same feeling as fixing it. That is why it is so popular, and it is why an author’s own audit — however honest — is worth less than one line of correction actually applied to the text.

That box is the reason the previous section exists rather than a conclusion. If this part had ended on “twelve objections land and I have faced them”, the facing would have been the whole product.

9.4 — What I am actually committing to

So let me put the only thing on the page that is not a description. Four commitments, checkable, and Part 16 will report on them.

One. The evidence grade. Every empirical claim across the fifteen parts marked by what it rests on — census, experiment, replicated study, single study, self-report. This is the correction with the highest value and it is a large piece of work.

Two. Part 6 and its audit republished as one object that cannot be separated.

Three. A short version. Genuinely short — the length that fits in the time a person with two shifts actually has.

Four. The missing part, or an honest statement that I am not going to write it: what these rules are costing right now, and who is paying, at the depth Part 10 gave to the other direction.

If Part 16 arrives and none of those exist, then this chapter was what it suspected itself of being, and you will have the evidence in your hands.

Remember This

Four failures that would matter: the arrow in Part 9 running backwards, the enforcement thesis being an artefact of household surveys, persuasion working on a lag longer than any study, and the series functioning as a way of not acting. Only the last cannot be checked from outside, which is why it is the dangerous one.

Of fourteen objections, twelve land, one half lands, one is survivable, and none fail. A prosecution where every charge sticks is not rigorous — it is curated.

One combination survives everything, and it is built entirely from this series’ own findings: it will be read mainly by people it cannot help, in a format that works least, on a question its own results say argument does not settle.

Underneath all self-criticism sits the assumption that naming a problem shows you are not subject to it. It does not. A confession is a description. This part is 20,000 words, which makes the length problem it diagnoses worse by 20,000 words.

Four commitments are on the page instead of a conclusion. If Part 16 arrives and none of them exist, you will know exactly what this chapter was.

10An Honest List of What We Do Not Know

Every part of this series ends here. This time the subject is the series itself — what genuinely cannot be established about it, why not, and the smaller list of things that are simply true.

10.1 — Genuinely unknown

The useful part of a list like this is never the gap. It is the reason for the gap, which is usually more informative than the missing fact would have been.

Who reads this, and what they do with it. I do not know. Nothing in fifteen parts required me to find out, and I have not. Some of it is knowable — where the parts are linked, which passages get quoted, whether Part 6 travels with its scoring chapter. Why it is unknown: because I have not looked, and because I preferred not to. An author who does not check how his work is used keeps the pleasant version of the answer.

Whether the framing changed any conclusion. Chapter Two argues that the functional question shaped everything. It might have shaped nothing. Why it is unknown: framing does not generate a claim that can be tested. It generates a table of contents. The only test is writing the series again from a different opening question, which nobody will do, including me.

Whether I would have published a conclusion that destroyed the project. I said in Chapter One that I would like to think so and cannot prove it. Why it is unknown: the situation never arose. A disposition that has never been tested is not a fact about a person. It is a hope he has about himself.

What the unrepresented would say. Chapter Eight listed the positions that never appear as arguments — the woman who complies because compliance is cheaper, the enforcing mother-in-law who does not consider herself an enforcer, the unmarried man at the bottom of the class structure. Why it is unknown: they have no literature, and this series is built entirely from the published record. Finding out would require going and asking, which is a different kind of work than any part of this series has done.

How much of the fragile evidence would survive checking. Chapter Three showed what happens to psychology and social science findings in general when somebody runs them again. Almost none of the specific studies quoted in this series have been through that. Why it is unknown: replication is expensive, unglamorous and rarely funded, so it happens to famous findings and not to the ordinary ones a series like this actually leans on.

Whether persuasion works on a long lag. Part 14’s verdict rests on campaigns measured over five to ten years. Why it is unknown: research budgets run in three-year cycles and moral changes run in generations. The gap between those two numbers is not a fact about persuasion. It is a fact about how research is paid for, and it means the question has never really been asked.

Whether the series has done more good than harm. Why it is unknown: there is no comparison world. I cannot observe the version of the last year in which these thousand pages were not written and see what those readers did instead. This is the hardest kind of ignorance, because it applies to the only question that finally matters and there is no instrument for it at all.

How We Actually Know This

Notice the shape of that list. Almost none of the items are unknown because the evidence is contradictory. They are unknown because nobody built the instrument.

Nobody funds replications of ordinary findings. Nobody runs studies longer than a funding cycle. Nobody surveys the people with no organised position. Nobody can construct a comparison world.

Part 13’s method chapter made the same point about natural experiments and Part 11 made it about men. When a whole category of question is unstudied, that is rarely because the question is hard. It is because no institution had a reason to pay for the answer.

What this cannot show: that the answers would be interesting. Some unstudied questions are unstudied because they are dull. You cannot tell which from the outside, which is the second-order version of the same problem.

10.2 — Solid

Shorter, as always, and worth more.

The replication figures. Of 100 psychology studies rerun by a 270-person collaboration, 36 produced a significant result where 97 of the originals had; the surviving effects were about half the original size. Of 21 social science experiments from Nature and Science, 13 replicated, again at about half size. A 2025 project across 54 journals put around 55 per cent of claims through successfully. Why it is solid: these were pre-registered, designed to be checked, run by teams with no stake in any particular result, and their methods are public. They are among the few studies in this document built specifically to be audited.

The sampling figures. Roughly 96 per cent of behavioural research subjects come from Western industrialised countries; about 68 per cent from the United States alone; those countries hold about 12 per cent of the world’s people. Why it is solid: it is a count of published papers, and anybody can recount them.

That reported sexual histories cannot all be true. In a closed population over a fixed period, men’s and women’s average number of opposite-sex partners must be equal. In every survey, men report more. Why it is solid: it is arithmetic, not research. It does not tell you who is misreporting. It establishes beyond argument that somebody is.

That this series contains six documented reversals. Why it is solid: the parts are published and dated. The expectation was in print before the finding was.

That Part 6 exists as a separable object. A sixty-five-page brief for the old rules, with its own cover, its own page and its own download. Why it is solid: one click confirms it.

That no part of this series addresses a reader with a decision. Why it is solid: it is a search across fifteen documents, and the result is zero.

That India’s flagship campaign for the girl child spent most of its released money on advertising itself. A parliamentary committee reported in December 2021 that of ₹446.72 crore released to states from 2016 to 2019, 78.91 per cent went on media advocacy. Why it is solid: it is a legislature auditing its own government’s spending and finding against it, which is the direction institutional evidence almost never runs.

In Real Terms

That last figure again, in ordinary money. Out of every hundred rupees the campaign actually got into the states’ hands, about seventy-nine went on telling people the campaign existed and about twenty-one went on everything else.

If a family spent its monthly food budget that way, four weeks in five would go on announcing dinner.

And one thing that is solid without being empirical. Where nobody is persuaded, the arrangement already running continues. That is not a claim about psychology and it needs no study. It is what a default is. Every argument in this series that ends without resolution ends with the rules still in force, and that outcome was determined by the structure of the situation rather than by anything either side said.

Remember This

Almost everything on the unknown list is unknown for the same reason: nobody built the instrument. Not funded, not surveyed, not measured longer than a grant cycle, not asked of people with no organised position.

I do not know who reads this, whether the framing changed any conclusion, what the unrepresented would say, how much of the fragile evidence would survive checking, or whether the series has done net good — and that last one has no comparison world, so it is not merely unmeasured but unmeasurable.

What is solid is narrow and mostly external: the replication figures, the sampling figures, the arithmetic showing reported sexual histories cannot all be true, the six dated reversals, and a parliamentary committee’s finding that four rupees in five went on advertising.

And one non-empirical certainty: where nobody is persuaded, the default runs. Every unresolved argument in fifteen parts ended with the rules still in force, and that was decided by the structure, not by the arguing.

The gaps in this list are not mysteries. They are places where no institution had a reason to pay for an answer.

Sources & further reading — Part 15

The Series, Part by Part

Fifteen parts, roughly a thousand pages. This table is the record of what each one claimed, so that a reader of this audit can check the charges against the things being charged.

PartTitle and what it established
OneThe Question Underneath the Question — set the framing for all sixteen parts, and the portability test used later to check whether an effect found in one place shows up in another.
TwoWhere the Rules Came From — the historical origins, and the warning that origin does not settle worth.
ThreeThe Indian Machine — how the rules are actually transmitted and enforced here; turns the where-it-came-from objection on the series itself.
FourThe Bodies Themselves — sex differences, with an effect-size table and heavy caveats about what an average difference does and does not license.
FiveThe Paradox of the Free Countries — four competing accounts of why outcomes in liberal societies are not what either side predicted.
SixThe Case for the Old Rules — written as advocacy; seven arguments at full strength, then scored against four conditions.
SevenWhat Feminism Actually Claims — written as an audit; a claims scorecard and an Indian timeline from 1848.
EightThe Ledger of the Equality Project — eighteen rows of outcomes, with a column asking who each change actually reached.
NineThe Arithmetic of Descendants — fertility, pronatalism and the finding that no country has reversed a decline.
TenWhat the Evidence Says Happens to a Woman — the core warning unpacked into four claims and tested; infection travels, marriage-market damage does not.
ElevenThe Half Nobody Studied — men as the missing comparison group; the rule falls where it can be enforced, not where belief points.
TwelveThe Community With No Edges — enforcement once the community loses its boundaries; two patterns of blackmail-by-photograph, priced differently by two independent sets of criminals.
ThirteenWhen History Ran the Experiment — nine cases graded against six tests; one genuine randomisation exists in the whole subject.
FourteenWhat Would Actually Work — four intervention shapes scored; persuasion classified as a category error, visibility as the one tier-one success.
FifteenThe Case Against This Series — this document. Fourteen objections, twelve of which land, and four commitments instead of a conclusion.
SixteenPart 16 starts here. The synthesis — and the report on whether the four commitments in Chapter Nine were kept.

Glossary

Every term this document treats as needing explanation, in plain English, in one place.

TermMeaning
AggregationAdding outcomes across many people to get a total or an average. It treats one person’s gain and another’s loss as quantities that cancel.
AuditChecking something against the original records rather than against the account the doer gives of himself.
Burden of proofThe job of proving your case. Whoever does not carry it wins by default when the evidence is unclear.
Category errorTreating something as a kind of thing it is not — for example, scoring a moral duty as though it were a policy with results.
ConsequentialismThe view that whether an action is right depends on the results it produces.
Constitutive goodSomething that is part of what a good life is, rather than a means to anything else.
CroreTen million, in Indian counting. ₹1 crore is ten million rupees.
DefaultWhat happens if nobody does anything. In a system of rules, it is the arrangement currently running.
Effect sizeHow big a difference something makes, as opposed to whether the difference is real at all.
False balanceGiving two positions equal treatment when the evidence behind them is not equal.
Functional explanationExplaining something by what it does, rather than by how it arose or whether it is right.
Hidden assumptionA premise both sides of an argument accept without stating, and therefore never examine.
LegibilityWhether a state can see, count and act on something from a distance, through forms and registers.
Natural experimentA real-world event that splits a population in a way an experimenter could not arrange, allowing a before-and-after or side-by-side comparison.
Portability testChecking whether an effect found in one place still appears in another. If it does not travel, it belongs to the setting rather than to the act.
ProxySomething you measure because you cannot measure the thing you actually care about.
Publication biasThe tendency for interesting results to reach print while boring ones stay in a drawer, so the published record is not a sample of what is true.
RandomisationAssigning people or places by lot, so that the groups differ only by chance and any later difference can be attributed to the treatment.
ReplicationRunning somebody else’s study again, on new people, to see whether the same thing happens.
Self-reportEvidence made of what people say about themselves on a form. Reliable for some things and not for anything they are punished for.
Steel manThe strongest version of an argument you disagree with. The opposite of a straw man.
Straw manA weak, distorted version of an opponent’s argument, built so it is easy to knock down.

Download The complete book · 4.0 MB