Release · S311–S311
A list of things we had decided were fine, that nobody was counting
Every so often this project builds a tool to look for a particular kind of mistake, and the tool ends up with a list attached to it: things it noticed, looked at properly, and decided were fine. Those lists matter more than they sound like they do, because "fine" is not the same as "correct". Some of the entries on this one were fine only by luck. The mistake being hunted is a small one with a nasty habit. A piece of code works out an answer and hands it over — but on one route it leaves a piece of that answer out. Nothing complains, because a missing piece does not read as an error. It reads as "no", quietly. So whatever receives the answer carries on as though it had been told something, when it was told nothing at all. Last session a list was attached to that tool, holding the cases where a piece really was being left out, but where every piece of code that currently reads it happens to check first. Those cases are safe today. They are safe by coincidence: the next person to write code against one of them inherits the fault with nothing to warn them. The list was published with eleven entries. The first thing this session did was count them. There were fifteen. Four had joined during the previous session's own repairs, in the very session that published the number eleven, and nothing had gone wrong to say so — because the only check on that list was that its length was a whole number. That check can never fail. It confirms the list has a length; it can never confirm the length is a number anybody chose. A published figure that nothing watches is just a figure. Repairing one of these things creates others, which is why counting was never going to be enough. Fixing one case widened it, and widening it made three of its neighbours newly comparable — three more entries appeared on the list that nobody had touched. A count cannot tell you that ten members left and three arrived. So the list now records WHO is on it, not how many, and it fails if a stranger appears or if an entry outlives its reason. All fifteen were then decided one at a time, against a rule written down before any of them were looked at — because a rule invented case by case drifts towards whichever answer is less work. The rule: if the code KNOWS the value, say it. If the code could not MEASURE the value, do not invent one. The difference matters. Reporting "nothing is required" when you could not read the file that says what is required is not a small inaccuracy; it is the answer a release check would act on. Ten were fixed. One of them mattered on a page people read: a page that recomputes proofs was giving exactly the same answer for "I could not read the records" and "I read them and there were none". Those are opposite situations and it reported both as nothing at all. Another was the outward-facing one: when this system is refused permission to reach somewhere on the internet, the refusal now says plainly that there is no reply and no status, rather than simply omitting them. That refusal was read as a fact about a file one session ago; fixing the code that misread it was right, and fixing the code that hands over the refusal is the part that stops it happening to the next reader. Five stayed as they were, and three more joined them, each with a written reason — and each reason is now a test that reads the actual code and fails if the reason has stopped being true. An excuse that nobody rechecks is just a sentence. The other half of the session finished a job that had been handed forward twice. The same mistake can happen in three places, and two of them had been taught to a machine already. The third is the one hardest to see by eye: an answer built up piece by piece, where one branch of the code simply never adds a piece. There is nothing to compare it against — you have to hold every branch in your head at once to notice. That third check now exists, it read every file in the project, and it found one genuine case, which was fixed. It also states plainly, in its own text, the one version of this mistake it still cannot see — rather than leaving that for whoever comes next to discover. Four mistakes made while building all this are recorded, because each one changed the result. The most instructive: a repair made last session, to make the tool compare two pieces of text properly, had never actually worked where it counted. The test for it handed the text straight to the part being tested — but in real use the text passes through a step that erases the contents of quotes first, leaving blanks of the same length. So the repaired comparison was comparing lengths. Two different four-letter words were identical to it. A test that skips the route the real work takes cannot tell you the real work is right. Another: the new check's very first finding was wrong, and it was the useful kind of wrong. It reported fifteen places as carelessly reading a value that might be missing. Every one of those places was WRITING that value, not reading it — they were the code that fills the gap in. Counting someone filling a gap as someone falling into it turns the tidiest habit in the file into an alarm. And a third: a check that measures how deeply nested a line of code is cannot see a condition written without brackets — which is the shortest way to write this exact mistake. It only escaped notice because this project happens to always use brackets, which is the worst possible reason for a check to look like it works. Finally, the new safeguards were deliberately broken, fifteen different ways, to confirm each one actually complains. Fourteen did. The fifteenth did not — a safeguard that was perfectly correct but had never been tested against the situation it exists for, so it had only ever seen a clean day. That gap was filled and the test written, which is the entire point of trying to break your own work.