I shipped a fix, reported it as an improvement, then measured it and found it did nothing at all
The premise of this series is that a claim is not evidence. Post one was about kubemend’s verification pipeline having its own bugs. Post two was about me trusting a status report instead of checking the artifact.
This one is about a nastier version of that: a check that reports success because it is structurally incapable of reporting anything else. Not a check with a bug in it. A check with no failure mode.
I wrote five of them in about two weeks. Here they are, worst first.
The one I actually published
kubemend’s eval harness pushes a commit to a GitOps repo, then inspects the
cluster to confirm the fault it just injected is visible. That races Argo CD’s
reconciliation interval, and losing the race shows up as a SymptomTimeout in
scenarios that have nothing else in common. So I added a wait between the push
and the inspection, using argocd app wait --sync.
The next sweep showed timeouts dropping from 3/5 to 2/5. I wrote that down as the improvement it looked like.
Then I measured it. argocd app wait --sync returns in zero seconds, in
both places I called it. It returns the moment sync status reads Synced -
and immediately after a push, the app is still Synced, to the previous
revision. It never waits for anything. It cannot. The function I’d shipped, and
reported on, was a no-op with a plausible name.
The 3/5 to 2/5 movement was noise. I’ll come back to that.
The pattern, four more times
The harness needs live Argo credentials before a sweep starts, or every
iteration fails identically after paying for its own reset and fault
injection. The obvious probe is argocd account get-user-info, so that’s what
I used. It exits 0 for a completely invalid token, and prints Logged In: false while it does it. Exit code zero. My check was reading the exit code.
It validated nothing, for any input.
The gitea token check had the same shape, inverted. It used /api/v1/user,
but that token is minted with scopes: ["write:repository"], so /user
answers 403 for a perfectly valid credential. This one couldn’t pass - it
regenerated the token on every single run, which looks like working software,
because the regenerated token is fine. Both probes are the same defect: the
output is uncorrelated with the thing being tested. One’s stuck on yes, one’s
stuck on no.
Then there’s the lint filter, which is the one that actually embarrasses me. I
was running the project’s lint task through a grep to cut the noise.
ruff format --check failed, and my filter dropped the line. Worse, the
format failure aborted the task before mypy ran, so a real type error never
got a chance to print either. Two genuine failures, one filtered and one never
reached, and a clean-looking terminal. CI caught it, which is to say a human
caught it, which is to say I’d built a check whose output I had personally
taught to be reassuring.
The last one didn’t even involve a check I wrote. Two stacked pull requests
both touched test_validator.py, and the squash landed the fix while
silently dropping the three tests that proved it worked. The suite went from
466 to 463 and stayed green, because tests that don’t exist don’t fail. A
helper class survived in the file, quietly serving nothing. The tests are back
now, and they’ve got the names they should have had all along:
test_a_starved_co_tenant_does_not_make_an_over_quota_proposal_look_like_it_fits
test_headroom_still_allows_a_proposal_that_genuinely_fits
test_the_app_being_changed_is_never_counted_as_its_own_neighbour
Before restoring them I reverted the fix and confirmed they fail. That’s the whole point of this post, so it would have been embarrassing not to.
Why I believed the no-op worked
Because a number moved, and I read the number at a resolution it doesn’t have.
These sweeps are nine scenarios at n=5 against a live cluster with a deliberately weak model. A scenario moving by two iterations is inside the noise floor. When I re-baselined after a batch of real fixes, the total went from 28/45 to 25/45 - it got worse - and I never found a cause. The harness change was inert, the cluster was healthy, no scenario failed for a newly identified reason.
So neither 28/45 nor 25/45 is quotable as a baseline, and I’m not quoting them as one. What that episode really established is that several conclusions earlier in my own diagnosis document were read at exactly that resolution, including the one about the sync wait. A no-op and a real fix produce the same distribution of n=5 results. I couldn’t have told them apart with the method I was using, and I didn’t.
The rule
Verify the probe against a known-bad input before you trust it.
Every one of these dies immediately under that rule. Point the token check at a garbage token and watch it pass. Break the formatting on purpose and watch the filter swallow it. Delete the fix and watch the tests still go green. Time the wait and watch it return instantly. None of this is sophisticated. It’s just the difference between “I ran the check and it said yes” and “I know what this check says when the answer is no.”
The corollary, which is what cost me the lint incident specifically: check exit codes, not filtered output. Anything that reshapes a tool’s output before you read it is a place where a failure can go to die.
What’s real, because it was measured
argocd app get --hard-refresh after a push cuts push-to-observable-symptom
from 55 seconds to 26 seconds, and removes the dependence on Argo’s poll
interval - which was the actual mechanism behind the cross-scenario timeouts.
That’s the change that shipped, under a name that describes what it does. It is
a refresh and not a sync: it makes Argo re-read git sooner and applies nothing
itself, so it does not widen what the harness is able to do to a cluster.
I also tested and disproved two other hypotheses I liked, including one about event-list truncation hiding a fresh Kubernetes event. Sixty-six events returned, the event was present, hypothesis dead. Writing down a disproof you were hoping would be a cause is most of the value of keeping the document at all.
The part I didn’t get to claim
The milestone this work belongs to has three acceptance criteria. Two are met. The third - a trustworthy cheap-tier baseline number - is not, and the plan now says so in those words rather than quoting a figure I’ve just spent a week demonstrating I can’t trust. There’s also a new infrastructure-error classifier that is unit-tested, correct by construction, and has never once fired in a real sweep. It is recorded as field-unproven, because that is what it is.
An agent harness whose entire premise is that the model’s self-report doesn’t count is a slightly ridiculous place to keep discovering that my own reports don’t count either. But it is the same failure at a different layer, for the same reason: it’s much easier to run a check than to establish that the check can fail.
The diagnosis document, including the correction where I retract the sync-wait
claim in place rather than quietly deleting it, is in
evals/reports/cheap-baseline/diagnosis.md.