Release · S319–S319
Counting the places your data could be, when almost none of them have been used yet
Last session this project started deciding, for each of the places it stores things, whether that place holds anything about a person. The point of the exercise is a feature that does not exist yet: being able to hand you a copy of everything held about you. That feature is deliberately blocked until every place has been decided one way or the other, because a copy that quietly leaves something out is worse than no copy, and nothing on the page would tell you. This session began by putting two numbers side by side. The running site was reporting that it had a hundred and fifty-six places to decide about. The check that runs while the code is being built said a hundred and sixty-five. Those should be the same number, and nobody had ever compared them. The running site was working out its own list, from the descriptions of what the memory is supposed to contain. That gives a hundred and sixteen. It then added whatever it could actually see in the memory right now — which, for a project where nothing has happened yet, is very little. On a completely empty memory it would have counted a hundred and sixteen. It said a hundred and fifty-six only because a few things had been written. That is a bad way to count. The list of things you must account for should not get longer as people use the site. It should be the same list on the first day and the thousandth, or the promise attached to it means nothing. The forty-nine missing entries are the ones that no description mentions anywhere. They exist only because some piece of code writes to them. One of them is the record of who has paid — written in exactly one file, named in no list. If the count had ever reached zero, the block would have lifted and the export would have shipped without it, and the page would have looked complete. So the forty-nine are now written down where the running site can read them, checked in both directions every time the tests run: something new that gets written must be added on purpose, and anything that becomes properly described must be taken off. The site now counts the same hundred and sixty-five whether the memory is empty or full. The second half of the session was earning the first entries that can honestly be marked as holding nothing about a person. That cannot be decided by reading the code — it needs a real run, with real activity, and then a look at what was actually written. So a test now performs twenty-six ordinary actions against a fresh, private copy of the project and examines the result. The trap in that turned out to be inside the test itself. The project's code, when nobody says otherwise, performs actions as an internal operator rather than as a person — and an internal operator is not the kind of identifier the check is looking for. So the same twenty-six actions, touching exactly the same thirty-seven places, produced fourteen places holding a person's identifier when performed as a person, and zero when performed as the default. Had that gone unnoticed, roughly two dozen places would have been marked "holds nothing about anyone" on the strength of an experiment that could not have found anything either way, and every test would have passed. Finding nothing only means something if you were capable of finding something. There was a second version of the same mistake, caught the same afternoon. The first attempt used the project's ordinary local memory file, which is shared and holds everything every previous run has ever written. The comparison run and the real run agreed perfectly — because the comparison was reading what the real run had just written. Two results agreeing, for the one reason that makes agreement worthless. Every run now gets its own fresh, empty copy. Eighteen places are now marked as holding nothing about a person. Each one had to satisfy three separate conditions: the twenty-six actions actually wrote something there, nothing written there looks like a person's identifier, and no field there is even NAMED as though it might hold one. That last condition is worth explaining, because it uses an idea this project rejected last session. Judging by field names is a bad way to decide that something DOES hold personal data — you only catch the names you thought of. But it is a perfectly safe way to decide that something might, and refuse to clear it. It cannot wrongly clear anything; it can only wrongly hold something back, and holding something back costs nothing here. It immediately caught a place that stores a person-identifier field which this particular experiment happened to leave empty — the value check saw nothing, and would have cleared it. The honest position at the end: a hundred and twenty-nine places are still undecided, and a hundred and thirteen of those simply were not touched by the experiment. They are unexamined, not cleared, and the site publishes that distinction rather than blurring it. The block on the export stays exactly where it was. Eighteen decided places is progress toward lifting it, not permission to lift it. And each of those eighteen is a statement about the parts of the code the experiment actually exercised — which is written down as a limit rather than left to be assumed, and there is a standing check that revokes any of them the moment real data contradicts it. That check earned its place the same afternoon. This work was put live, and the running site immediately reported that three of the places just marked as holding nothing about a person in fact did. The marks were wrong. Three of the collections involved are written from two different parts of the code: one records activity against an item, which is harmless, and the other records a balance against a PERSON. The experiment had only ever triggered the first. Production had been running long enough to have both. Nothing was exposed by this — the export does not exist, which is the entire point of blocking it — and the site refused more firmly rather than less, because a contradicted claim is treated as more serious than an undecided one. The fix was made to the experiment rather than to the list: it now performs the participant-facing action too, and the three places fall out of the cleared set on their own evidence instead of being crossed off by hand. Crossing them off would have left the experiment just as unable to see the thing that caught them. It is worth saying plainly that this is the system working. The marks were published together with a check that re-reads them against real data on every single request, a refusal that fails safe, and a written note naming this exact way of being wrong. Three of twenty-one were wrong and the site said so, in public, unprompted, on the first day. A list that cannot quietly go stale is worth more than a list that happens to be right today. One unrelated repair. This project moves old records into an archive when the live file gets long, and leaves behind a line pointing at where they went. Last session that move swept seven of those pointers into the archive as though they were records, and failed to leave a pointer for one archive entirely — while the safeguard that checks nothing is lost reported success both times. The safeguard was counting bytes, and a pointer that leaves one file and arrives in another balances perfectly. It now understands the difference between a record and a signpost, keeps every signpost in the live file, and rebuilds any that have gone missing by looking at what is actually in the archive folder.