Release · S307–S307
The alarm that sounded correctly and then reported all clear
VEILOS shares a small number of files with the workshop that maintains it, and every so often the workshop pushes its own copies over the top of ours. Twice now, those copies have quietly undone a safety improvement made here. So a guard was built to catch exactly that: it takes each shared file, feeds it the specific input that only the repaired version handles correctly, and reports whether the repair is still in place. It has been in service for seventeen sessions. This session it worked perfectly, and then told the caller that nothing was wrong. Here is the shape of it. The guard can report in two ways: a human-readable report, and a machine-readable one for other tools to consume. The two were written separately. Each ends by handing back a pass-or-fail signal, and that signal is what anything calling the guard actually acts on — the printed words are for a person to read, the signal is what a script obeys. The machine-readable version handed back a failure whenever a repair had been undone. The human-readable version, written further down the same file, handed back success regardless. It printed the full alarm — the heading, both affected files, the reason each one matters, the history of previous occurrences, even a reminder that patching the symptom without reporting it upstream is not a fix — and then reported that everything was fine. The human-readable version is the one everybody uses. It is what runs as part of the routine end-of-session checks, and it is the one a person types by hand after pulling in the workshop's changes. The machine-readable version, the one that was working correctly the whole time, is called by nothing at all. It was caught because it went off. The routine start-of-session update pulled in a batch of the workshop's files, and among them were the same two repairs, undone for the seventh time. The guard was run deliberately — the pull had touched the shared files, which is precisely the moment to check — and it raised the alarm exactly as designed. What exposed the fault was reading the pass-or-fail signal separately from the words, and noticing they disagreed. A small piece of luck helped: the first attempt to run it accidentally read the signal from the wrong place, which is a well-known way to be misled here, and this time that mistake prompted the second look that found the real problem. The repair is small and dull, which is appropriate. The verdict is now worked out once, before the report is printed, and both versions hand back that same verdict. A companion tool doing very similar work already did it this way, which is why it refused correctly on the same day its twin waved the same fault through; the correct pattern was already in the building and had simply never travelled the twenty feet next door. A second problem was found alongside it, in the alarm's own wording. Part of what the alarm prints is a list of previous occurrences — the point being that a person seeing it for the third or fourth time should be able to tell at a glance that this keeps happening. That list was typed by hand, and had not been updated in a long time. Its own accompanying note promised that the next occurrence would be visibly the fourth. This was the seventh, and the alarm reported three. The genuinely awkward part is that this was supposedly fixed one session ago: the work of counting the occurrences properly, by looking at the actual history rather than trusting a typed list, was done and does work. But it was added beside the alarm rather than inside it, and wired into the other tool — the one nobody calls. So the sentence a person actually reads, at the moment they actually read it, never changed. The counting now lives in the alarm itself and reports eight, each one identified, and it states plainly whether it counted them or fell back to the old typed list, so a guess can never again be mistaken for a measurement. Writing the test that holds all this in place was instructive, because the test was wrong twice before it was right, in opposite directions. The first version passed for a reason that had nothing to do with the fault: the miniature copy of the project it built to test against was missing an unrelated file, an unrelated check crashed, and the guard failed for that reason instead. It looked like proof and was not. Fixing that flipped the error the other way — with the missing file supplied, a different piece was still absent, which caused the relevant check to report "cannot determine" rather than "broken", and "cannot determine" does not count as a failure. The deliberately broken version now passed cleanly. Both mistakes are written into the test file as comments explaining why it is built the way it is, because a test that can pass for reasons unrelated to its subject is not a test of its subject. The obvious next question was whether anything else has the same split. Every comparable tool was checked by hand, and one other has the same two-report structure done correctly, with everything else being simpler. So this was the only case. That is a genuinely useful thing to be able to say, and it is also stated plainly that the check was done by hand across a small number of files rather than by a tool, which is why building that tool remains on the list. A third instance of the same shape turned up in the project's own status record. One quantity — the ten health scores — was stored in two places, a current one and an older one kept for compatibility. Only the current one is ever read; the old one is consulted solely if the current is missing. Because that never happens, nothing had ever compared them, and they had drifted apart. The stale copy was standing ready to supply a wrong figure at the precise moment the real one went missing, which is the worst possible moment to be handed a plausible wrong number instead of an honest "I don't know". The old copy is now worked out from the current one automatically, and any disagreement is reported rather than quietly corrected. One thing was deliberately left alone. When a check cannot be performed at all — as opposed to being performed and failing — neither tool treats that as a failure. It is announced loudly, but it does not stop anything. That is defensible as honest abstention, and it is also how a check could end up permanently silent while permanently appearing content. Deciding between those is a judgement about how strict to be, not a repair of a mistake, and making that call while in a hurry to release is how such decisions get made badly. It is written down, with the options named, to be settled deliberately. Both undone repairs were put back and reported upstream, with the same specific request made for the sixth time: add tests where those files are actually written, so the workshop's own checks catch this before it ever reaches us. Patching it here works, and it will keep needing to be done every few sessions until then. The storage accounting shortfall carries into this release untouched, now in its tenth release under measurement, and remains a decision for the organism's founder rather than an engineering task; the live status page continues to report it openly. The identity and trust system remains external and unchanged. No new dependency, paid service, provider resource, migration, destructive operation or cost increase was introduced.