Aug 6, 2026 · v0.3.1
Two False SAFE Verdicts
Two false SAFE verdicts, found and fixed. Both had the same shape: coverage measured a different program than the tests gate ran, then reported the changed lines as covered. A false SAFE is the one defect this product cannot have, so upgrade rather than pin.
Two false SAFE verdicts, found and fixed. Both had the same shape: coverage measured a different program than the tests gate ran, then reported the changed lines as covered. A false SAFE is the one defect this product cannot have, so upgrade rather than pin.
Everything here is a fix. No command, flag, contract or report field changed.
Fixed
- A leading `NAME=VALUE` on `testCmd` silently disabled coverage. The tests gate runs the override through
sh -c, which honours the assignment; the coverage runner tokenized the command itself and treatedPYTHONPATH=.as a module name.coverage run -m PYTHONPATH=. python3 -m pytestimports nothing, writes no data file, and the verdict degraded toUNPROVEN. That is the safe direction, but it capped every project using the documented shadow-bypass remedy atUNPROVEN, which no amount of test-writing could lift. Leading assignments are now hoisted into the child environment. (#95, PR #97) - A console entry point ran under the wrong interpreter.
coverage run <script>executes the file as source in the *current* interpreter and never honours its shebang; the shell that runs the gate does. A venv'spytestwas therefore measured under whateverpython3we happened to spawn, with a different set of installed packages. The resolver now models what the shell actually does: it stops at the shell's first PATH match, requires the exec bit, declines on Windows, and accepts only two shebang shapes. (#98, #99, PR #102, PR #103) - Shebang arguments were dropped, which was itself a false `SAFE`.
#!/usr/bin/python3 -sis the Fedora and RHEL packaging default, and-Emakes Python ignorePYTHONPATH. The gate execs the script so the kernel applies the flag and imports the *installed* copy; coverage dropped the flag, honouredPYTHONPATH, imported the *shadow* copy and measured it as covered. The changed file lands inmeasuredFiles, so the shadow-bypass guard stays silent and the fusion readsSAFEfor a change no test executed. Any shebang carrying arguments now declines. (PR #103) - The shadow-bypass floor blamed the wrong thing. When measurement failed it told users to prefix their test command with
PYTHONPATH=., which could not have helped and was not the cause. The reason now names the real remedy. (PR #103) - Coverage declines now say what to do instead. A console entry point that is not a Python script (a native launcher on Windows, a pyenv, asdf or nix shim) can never be handed to
coverage run, so the decline is permanent rather than transient. The message names module form,python -m, as the fix. (#100) - Eight tests reported PASSED without running. Guards inside test bodies returned early when
coverage.pyor pytest was missing, which vitest counts as a pass. Converted toit.skipIf. Measured against a coverage-lesspython3shim:19 passed | 3 skippedbefore,12 passed | 10 skippedafter. (#101)
Changed
- A `testCmd` naming a console entry point that cannot be resolved now reports `UNPROVEN` rather than `SAFE`. This is a verdict change in the safe direction, and the reason it is not a major: the previous
SAFEwas, in the cases this covers, not something the measurement had established. - The default `testCmd` moved from `pytest -q` to `python3 -m pytest -q`. Module form is the only form coverage can reliably wrap.
- The CLI mascot and palette. The mascot is now Tabslot, and the chrome is monochrome cream, so the only colours on screen are verdicts, severities and diffs. Cosmetic; no output contract changed.
Docs
- An MCP tab, with per-client setup for nine clients. Claude Code, Claude Desktop, Codex, Cursor, Gemini CLI, VS Code, Windsurf and the generic case each get their own page with a copy-only prompt block you hand straight to the agent. (PR #94)
- One canonical answer for the `PYTHONPATH` question.
verdicts.mdxsaid to export it,verify-diff.mdxgave no form at all, and nine agent prompts said to prefix the test command without saying what the command should look like. All of them now name module form, which is not a style preference: coverage has to run the same program the gate ran, and a bare console script is only runnable under coverage when it resolves to a Python file. (#96)
