33 yes/no questions cannot identify 8 billion people. Here are 77 that nearly do.

33 yes/no questions cannot identify 8 billion people. 77 nearly do.
In short
  • A 2013 GitHub repo asked for 33 yes/no questions that uniquely identify every person alive. It has 361 stars, 39 forks, and two questions. Nobody finished it.
  • 33 is the wrong number, and the README says why in the sentence right before it. Assigning 8 billion serial numbers costs 33 bits. Asking questions about facts people already have costs about 65, because those answers land pseudo-randomly and the birthday problem applies.
  • At 33 perfect independent bits, 4.85 billion people still share their answer pattern with somebody.
  • I built the list. It runs to 77 questions and 65.6 Shannon bits. That's good enough for "almost everyone" and not good enough for "everyone" - P(nobody in 8 billion collides) is about 0.57.
  • Three models reviewed it adversarially and forced ten corrections, including one where all three independently told me the same block was worth 6.5 bits when I had 9.19. They were right.
  • Identical twins named Anna and Anne answer 60 of the 77 identically, including all six given-name questions. No number of extra questions fixes that.

In November 2013 someone opened a repo called 33-questions and wrote a very good README. It has 361 stars and 39 forks. It was last pushed on 25 November 2013, two days after it was created. The question list is still this:

1. Do you identify yourself as male?
2. Do you currently live in one of the following countries: China, India,
   The United States, Indonesia, Brazil or Pakistan?
3. ...

Thirteen years, nobody finished it. I spent a couple of days on it and I think I know why, and it isn't that the questions are hard to think of. The target was impossible from the first paragraph.

The finished list, the full working with every correction marked in place, and the measurement scripts are here: MarcinDudekDev/77-questions. Every number below tagged as measured came out of those scripts.

The error is in two consecutive sentences

Here is the README, quoted exactly:

We could give everybody on the planet a unique series of 33 1s and 0s, and identify anyone by their personal series. But that would be boring.

What if, instead assigning 1s and 0s, we had 33 'Yes' or 'No' general questions that, when answered correctly, uniquely identified everyone on the planet.

Those are two different problems, and the second one costs about 32 bits more than the first.

Assigning serial numbers is cheap because you control the assignment. You hand out 8 billion labels, each one used once, and log2(8×109) = 32.9 bits is enough. Nothing collides because you never let it.

Nobody is assigning anything here. The answers fall out of facts people already have - their birthday, their name, their mother's birthday. Those land pseudo-randomly across the space of possible patterns, so two people can land on the same one, and you get the birthday problem at planetary scale. With N people spread over M = 2b patterns, the expected number of colliding pairs is N²/2M. Setting that below 1 needs b = log2(N²/2) = 64.8 bits.

Log-scale chart of how many of 8 billion people still share an answer pattern, against bits collected. At 33 bits, 4.85 billion. At 65.6 bits, about 1. At 71.4 bits, 0.02.
Expected number of people who still share their answers with somebody, N(1 − e−N/M), for N = 8 billion. The curve is flat until well past 33 bits, which is exactly the region the original repo was aiming at.

At 33 perfect, independent, 50/50 questions, 4.85 billion people out of 8 billion still share their answer pattern with somebody. 33 questions doesn't nearly work. It fails almost completely.

The gap between 33 and 65 is the answer to the question the repo was asking. It isn't a detail you tune away with cleverer questions.

"Unique" has three different prices

This one I got wrong for three drafts, and a reviewer caught it by noticing I was quoting one metric and tabulating another. There are three thresholds and they're 6.6 bits apart:

TargetBits needed
Expected colliding pairs < 164.80
Expected people who aren't unique < 165.80
P(nobody at all collides) ≥ 0.9971.43

The last row is what "uniquely identified everyone on the planet" plainly means, and my finished list doesn't reach it. At its 65.6 bits, the probability that all 8 billion patterns are distinct is 0.57. A coin flip. Even at the 68.1 bits I claimed before the review, it was 0.90, so roughly one draw in ten still had a collision in it.

So "77 questions gets almost everyone" is true. "77 questions uniquely identifies every living human" is false, and it stays false under assumptions that are already too kind.

You can't ask your way past a fact

Every question mines some underlying fact, and no set of questions about a fact can extract more bits than the fact contains. I measured the big one by enumerating every valid date from 1926 to 2026 and weighting it by an approximate world age distribution:

H(your birth date) = 14.76 bits. That's a hard ceiling on every date question combined, however many you invent.

That kills the obvious strategy. The original draft I started from wanted to keep adding date parities: day odd, month odd, day-of-year odd, sum of day and month even. You run into the wall at 14.76 and everything after that is worth nothing. It's also why the answer is never "just ask 65 date questions".

Names have the same shape of limit. You can't ask about the 7th letter of a 4-letter name.

The way past the ceiling turned out to be more people, not more cleverness about one person. Your mother's birth day and month is a fresh 8.5-bit pool that your own birth date says nothing about. Twelve questions about both parents' birthdays take the colliding count from 8.6 million down to 2,674. Those twelve are worth more than the address, sibling and body-and-birth blocks put together.

Two questions worth exactly zero, and one I wrongly wrote off

The fun part of this problem is that a question can look independent and be mathematically determined by two questions you already asked. I brute-forced the candidates instead of arguing about them.

"Is the sum of your birth day and month even?" - zero bits, given day parity and month parity. Checked over all 372 day/month pairs, zero mismatches. All three of those questions were in the original draft's "best 20 to try first" list.

"Was your parent's age at your birth an even number?" - zero bits. Given your own birth year and the parity of theirs, the parity of the gap is fixed. My selector scored it at +0.0000 and dropped it.

Then the one I got backwards. I wrote off "do the digits of your birth year add up to an odd number?" as a duplicate of the year-parity question and cut it. Measured over 1926-2026 the two are independent, mutual information 0.0009 bits, and the digit-sum question is worth a full bit. My greedy search picks it ninth. It's question 9 in the final list.

The dependency is real, but it's conditional. Add a tens-digit question and the digit sum collapses: for 19xx the digit sum is 10 + tens + last, for 20xx it's 2 + tens + last, so its parity is exactly parity(tens) XOR parity(last), with zero exceptions over 1900-2099. That's what made two of my hand-written parents'-birth-year questions worth nothing, and it's why the birth-date block gets away with asking both - that block has no tens-digit question in it.

Then a result I didn't expect. A-M letter cuts at different positions in a name are effectively independent: mutual information between "1st letter is A-M" and "2nd letter is A-M" is 0.0045 bits, between 2nd and 3rd it's 0.0007. Vowel questions are not: "1st letter is a vowel" and "2nd letter is a vowel" share 0.278 bits. Names have strong phonotactic structure, and an A-M cut is coarse enough to average right over it while a vowel test sits directly on top of it. So the list uses A-M cuts everywhere and no vowel questions at all.

I also measured the cost of insisting every question be answerable by a human without arithmetic. It's zero, within measurement noise. There's no tradeoff between rigour and readability here, which was a relief.

Where I was wrong

When the list was finished I handed it to three models - Grok, Sol running in Cursor, and Fable - and asked each to attack it. Between them they forced ten corrections. I'd rather write these down than quietly fold them in, partly because a couple of them are the kind of mistake I'd make again.

I introduced the exact error I'd criticised

I spent a whole section of my notes complaining that the original draft contained zero-information questions. Then I hand-wrote the two parents'-birth-year blocks instead of running my own selector over them, and put in two questions worth exactly zero bits. The fix was to delete my hand-written block and let greedy_parent.py pick, measuring each candidate's gain conditional on your own birth year already being known.

I budgeted in the wrong entropy

I did the whole budget in Shannon entropy. Collisions don't depend on Shannon entropy. They depend on Rényi-2, the collision entropy, because expected colliding pairs = N²/2 × Σp² and Σp² = 2−H₂. H₂ ≤ H₁ always, with equality only for a perfectly uniform distribution, and names are the opposite of uniform.

On the date block the gap is small: H₁ = 11.760 against H₂ = 11.588, because that block is nearly uniform by construction. On the name blocks it's brutal. A review estimate using real Chinese surname frequencies put one name block's collision entropy at about 4.25 bits against the 5.95 I'd credited, a 29% loss.

If the name blocks loseEffective bitsPeople still colliding
0% (uniform, unreal)65.451.1
10%62.836.3
20%60.2139
30% (the Chinese-surname figure)57.59152

So the honest headline isn't "0.2 people still colliding", which is what I had. It's somewhere between about five and about a hundred and fifty, and the width of that range is set almost entirely by one thing I never obtained: real name-frequency data.

All three reviewers said the same number, and I argued

I measured each parent's birth year in isolation. Mother 4.50 bits, father 4.69, credited as 9.19. Every one of the three reviewers came back independently with roughly 6.5 (Sol said 6.70, Fable about 6.5, Grok about 6) and pointed at the same thing: assortative mating. Parents' ages at your birth are correlated, fathers run about three years older, and measuring the two in isolation double-counts. Taken separately the two years have ceilings of 4.77 and 4.97 bits, so 9.74 between them. Model the age gap with a standard deviation of 2 to 4 years and the joint ceiling drops to 7.8-8.8. The budget now carries a 2.19-bit deduction for that overlap, which puts the pair at 7.00.

They were right. What bothers me is the order I did it in: I only moved after measuring the joint myself. Three independent models converging on the same correction should have been enough to make me measure it first, not to make me defend the number until I'd checked. I wrote about scoring models on restraint a few weeks ago, and this is the human version of the same failure.

Two smaller ones

I claimed my ordering was prefix-optimal, meaning any prefix of the list is the best identifier of that length, and I never tested it. It was false by up to 2.00 bits at the worst truncation point. Reordering so the parents' date blocks come before the parents' name blocks cuts that to 0.34 bits and costs nothing else.

And I understated the identical-twin population twofold, by taking "0.35% of deliveries" and using it as "0.35% of people". A twin delivery produces two people. It's 0.70%.

Anna and Anne

Sections one through five treat this as a budget problem. Buy enough bits and the collisions go away. They don't, and the reason isn't statistical.

Take identical twin sisters. Same birth day, same parents, same address, same family name, same siblings, named Anna and Anne.

Grid of 77 numbered dots. Sixty are dark, marking questions identical twins answer identically by construction. Seventeen are gold, marking the ones the list allows to differ.
Sixty of the seventy-seven are forced identical for a twin pair. Every high-value block is in the dark group.

The six given-name questions are the one thing in the list that should separate them, and they don't. Both names are 4 letters. Both start with A. Both have N as the 2nd and the 3rd letter. Both end in a letter in A-M. I checked all six mechanically and Anna and Anne produce the same answer to every single one. Sofia and Sonia differ on exactly one of the six. Anna and Alma differ on two.

Adding more questions of the same kind cannot help, however many you add. That's the part of the bit budget that averages hide: 65 bits is enough on average, and the average is not where the problem lives. Identical twins are about 0.70% of people, so roughly 56 million people in 28 million pairs, all sitting in the one region of the answer space that generic questions never reach.

The only questions that touch it are ones that target within-pair difference. That's why an otherwise 50/50 list ends with three questions split 2%, 1% and 45%. Question 76, "were you delivered first", is worth 0.081 bits averaged over everyone, which is nearly nothing. For the roughly 160 million people who were born as part of a multiple birth it's the most valuable question in the list, and for the 56 million identical twins among them it's one of only two or three that can separate a pair at all.

What I'd trust and what I wouldn't

The date measurements are solid - full enumeration of every valid date from 1926 to 2026, weighted by an approximate world age distribution. Changing the age weights moves the splits by fractions of a percentage point.

The name-independence result uses English dictionary words as a proxy for names, because I had no name-frequency corpus on the machine. The structural finding it supports, that coarse A-M cuts barely interact while vowel questions do, should survive a change of corpus. The exact splits for real names will not.

Six of the sixteen blocks - both parents' given names, the mother's birth surname, the address, family structure and body - are estimated rather than measured. Their discounts are reasoned guesses. The 54.41-bit total for the core 59 questions therefore carries maybe plus or minus 2 bits of real uncertainty, which moves the "2,674 people" figure by a factor of a few either way. It doesn't move the headline.

And the whole thing assumes people answer correctly. At a 2% per-question error rate, the chance of getting all 77 right is 21%. In practice that's a bigger threat to this list than any correlation I've spent days measuring.

The 77 questions

Spelling rule for every name question: use the spelling on your main government ID or birth certificate. No Latin form? Transliterate phonetically. Ignore hyphens, apostrophes and spaces. "First given name" and "family name" mean whichever your own culture treats as such. When a question asks about a letter your name is too short to have, answer no. When you don't know, answer no, and never substitute a fact from another question.

They're grouped by where the answer comes from, and within that by how much each block is worth, so stopping early still leaves you holding the best questions you could have asked. Percentages are the share of people expected to answer yes.

Sex

1 question · no lookup

#QuestionYes
1Do you identify as male? (under ~12: were you recorded male at birth?)50.4%

Your age and birth date

12 questions · your birth certificate

#QuestionYes
2Were you born in 1995 or later? (at or below the global median age, ~31 in 2026)53.8%
3Were you born in January-March or July-September?49.9%
4Is your day-of-year an odd number? (1 January = 1)50.1%
5Were you born in January-June?49.6%
6Is your birth year an even number?50.5%
7Is your birth day of the month the 16th or later?50.7%
8Is your birth day of the month in 8-15 or 24-31?50.7%
9Do the digits of your birth year add up to an odd number?49.1%
10Were you born on a Monday, Tuesday or Wednesday?42.9%
11Is your birth day of the month an odd number?51.0%
12Were you born on a Tuesday, Thursday or Saturday?42.9%
13Is the last digit of your birth year 0, 1, 2, 3 or 4?50.5%

Your given name

6 questions · no lookup

#QuestionYes
14Does your first given name start with a letter A-M?51.8%
15Is the 2nd letter of your first given name A-M?49.6%
16Is the 3rd letter A-M? (shorter name: answer no)49.6%
17Is the 4th letter A-M? (shorter name: answer no)54.4%
18Is the last letter A-M?53.0%
19Does your first given name have an odd number of letters?49.7%

Your family name

6 questions · no lookup

#QuestionYes
20Does your family name start with a letter A-M?~51%
21Is the 2nd letter of your family name A-M?~50%
22Is the 3rd letter A-M? (shorter: no)~50%
23Is the 4th letter A-M? (shorter: no)~54%
24Is the last letter A-M?~53%
25Does your family name have an odd number of letters?~50%

Your mother's birth day and month

6 questions · ask your mother

#QuestionYes
26Was your mother born in January-March or July-September?49.9%
27Is her day-of-year an odd number?50.1%
28Was she born in January-June?49.6%
29Was she born on the 16th or later?50.7%
30Was she born on a day in 8-15 or 24-31?50.7%
31Was she born on an odd-numbered day of the month?51.0%

Your father's birth day and month

6 questions · ask your father

#QuestionYes
32Was your father born in January-March or July-September?49.9%
33Is his day-of-year an odd number?50.1%
34Was he born in January-June?49.6%
35Was he born on the 16th or later?50.7%
36Was he born on a day in 8-15 or 24-31?50.7%
37Was he born on an odd-numbered day of the month?51.0%

Your mother's birth year

5 questions · ask your mother

#QuestionYes
38Is your mother's birth year an even number?50.0%
39Is her birth year divisible by 4, or 1 more than a multiple of 4?50.0%
40Was she under 23, or 33 or older, when you were born?51.4%
41Is the last digit of her birth year 0, 1, 2, 3 or 4?50.0%
42Was she under 28 when you were born?51.6%

Your father's birth year

6 questions · ask your father

#QuestionYes
43Is your father's birth year an even number?50.0%
44Is his birth year divisible by 4, or 1 more than a multiple of 4?50.0%
45Was he between 25 and 34 when you were born?45.8%
46Is the last digit of his birth year 0, 1, 2, 3 or 4?50.0%
47Was he under 28 when you were born?28.0%
48Is the tens digit of his birth year odd?50.0%

Your mother's given name

4 questions · ask your mother

#QuestionYes
49Does your mother's first given name start A-M?~51%
50Is its 2nd letter A-M?~50%
51Is its 3rd letter A-M? (shorter: no)~50%
52Is its last letter A-M?~53%

Your father's given name

4 questions · ask your father

#QuestionYes
53Does your father's first given name start A-M?~51%
54Is its 2nd letter A-M?~50%
55Is its 3rd letter A-M? (shorter: no)~50%
56Is its last letter A-M?~53%

Your mother's birth surname

4 questions · ask your mother

#QuestionYes
57Does your mother's birth surname start A-M?~51%
58Is its 2nd letter A-M?~50%
59Is its 3rd letter A-M? (shorter or unknown: no)~50%
60Is its last letter A-M?~53%

Deeper into your own names

4 questions · no lookup

#QuestionYes
61Is the 5th letter of your first given name A-M? (shorter: no)~58%
62Is the 6th letter of your first given name A-M? (shorter: no)~65%
63Is the 5th letter of your family name A-M? (shorter: no)~57%
64Is the 6th letter of your family name A-M? (shorter: no)~63%

Where you live

4 questions · your address

#QuestionYes
65Is your house or building number even?~50%
66Is the last digit of your house or building number 0-4?~50%
67Is the numeric part of your postal code even?~50%
68Does your street name start with a letter A-M?~50%

Family structure

3 questions · your family

#QuestionYes
69Do you have at least one older sibling by the same mother?~56%
70Do you have at least one younger sibling by the same mother?~52%
71Did your mother bear an even number of children in total?~48%

Body and birth circumstance

3 questions · yourself

#QuestionYes
72Are you taller than the median adult of your sex in your country?~50%
73Were you born before noon, local time? (unknown: answer no)~50%
74Does your town or city of birth start with a letter A-M?~51%

Twin distinguishers

3 questions · your birth record

#QuestionYes
75Were you born as part of a multiple birth (twin, triplet or more)?~2%
76If so, were you delivered first? (not a multiple: answer no)~1%
77Was your birth weight above 3.2 kg / 7 lb? (unknown: answer no)~45%

If you stop early

Stop afterBitsNarrows you to
1312.81 in 7,000
2524.71 in 27 million
4943.31 in 11 trillion
6154.41 in 24 million billion
7765.8everyone but a handful, and never identical twins

Those are counts of answer patterns, not of people. With 8 billion actual people the last row still leaves somewhere between a handful and a couple of hundred sharing a pattern with somebody, because real answers cluster. The 65.8 in that row is the straight sum of the blocks; the 65.6 I've been quoting is what's left after deducting the overlap between the sibling questions and the parents' birth years.

Where this lands

Getting the general population to uniqueness is a budget problem, and 77 questions very nearly solves it. Getting everyone there is a different problem, and no list of yes/no questions about inherited facts solves it, because the questions interrogate facts that twins share. That last step needs an attribute that's unique by design rather than by accident: a national ID, a phone number, a biometric hash. At which point you've stopped asking questions and started reading an identifier.

Which is roughly what the original README said in its first sentence, before it talked itself out of it.