Claude Code accepts, without a single error, settings it will never apply: a Write() permission, a hook on bash, a settings key placed in the wrong file. Cursor and Copilot do the same with their own files. deadweight reads the agent configuration of a repository and reports what the agents ignore. I introduced it on 30 September. Since then it also checks Cursor, Copilot and AGENTS.md, it has grown to 53 groups of checks, and it hands a model the two defects no parser can see. Here is what changed, and how it is built.
The problem: a silent refusal
A coding agent does not tell you when it ignores part of your configuration. No error, no warning: the session starts, and the rule you wrote does not count.
Three examples on the Claude Code side, all documented by Anthropic:
A
Write(src/**)permission is never consulted. Path rules are checked againstEdit(...)andRead(...)only: it should have beenEdit(src/**).A hook whose matcher is
bashnever fires. Matchers are case-sensitive, and the tool is calledBash.A settings key placed in the wrong file is ignored. Each key has a scope (project, user, managed): the same line that works in one file is dead in another.
None of these breaks anything. That is why they last: a configuration that ignores part of what you wrote always looks like it works.
Claude Code is rarely alone in the repository
Of 300 public repositories configured for Claude Code, nearly one in two also carries another agent's files: AGENTS.md, Cursor, Copilot. Each of these tools silently skips a misnamed or misplaced file, and some read the others' files.
Since 0.20, deadweight checks those files too. A Cursor rule saved as .md when Cursor only reads .mdc. A Copilot instructions file without the .instructions.md suffix. An AGENTS.md that Claude Code does not read because a CLAUDE.md sits next to it. And the lines of CLAUDE.md that only Claude Code can follow: Cursor applies that file to every conversation, and an instruction such as "run /compact" becomes an order it cannot carry out.
AGENTS.md is the most misleading case. Since version 2.1.277, Claude Code reads it on its own, but only when no CLAUDE.md exists in the directory or above it. As soon as there is one, Claude Code reads CLAUDE.md instead, and AGENTS.md no longer counts, unless CLAUDE.md imports it with an @AGENTS.md line or the "Project instructions" setting asks for both. A CLAUDE.local.md counts too: the developer who creates one for personal notes switches off, without knowing it, the whole team's AGENTS.md on their machine.
For Gemini CLI, the tool only provides documentation so far, no checks.
Where each rule comes from
deadweight sorts its rules by origin, and the severity follows. A rule taken from a vendor's documentation is quoted, with the name of the page. A rule that comes from a measurement carries its date and the tool version it was checked against. House conventions are labelled as such and only surface as a notice, unless the project chooses to make them stricter.
How it is built
The audit is a Python package with no dependency at all: one module per family of checks, plus file parsing, vocabulary, the report and the floor. With no network call, it runs in any CI, on Linux, macOS and Windows.
Each check receives the audit state as a parameter, and a check that crashes produces a "check failed" finding instead of taking the whole report down. The split into modules, in 0.21, was verified output against output: byte-identical on 200 public repositories re-cloned at their recorded commit.
The script parses YAML front matter and expands globs itself, so there is nothing to install. Those home-made parsers cost a lot. A trailing comment stripped the wrong way produced 24 false "unknown model" findings on one sample. A buffer not flushed produced 422 on a single public repository. Both are fixed, and the code keeps a note of why.
The floor covers the whole package
The CI floor (--set-floor, --check-floor) records the fingerprint of the auditor that measured it. Since 0.21, that fingerprint covers every module, not just the entry point. If the auditor changes, the comparison is refused: two different instruments do not produce two states of the same project, and comparing them silently mistakes a change of tool for progress.
How we know it is right
The first sample, told in the previous article, had 84 wrong or misclassified errors out of 98. The method got stricter from there. Seven samples of 150 public repositories were drawn at random, with a fixed seed and no repository ever reused. The instrument is frozen before each draw, nothing is fixed during the measurement, the pass mark (9 correct errors out of 10) is set in advance, and every error is read by hand.
Samples 4 to 7 gave the following values:
sample 4: 69 correct errors out of 82;
sample 5: 126 out of 244, because a single repository produced 100 of the 112 false ones;
sample 6: 1,416 out of 1,425, and still 127 out of 136 with the dominant repository set aside;
sample 7: 42 out of 48.
The 9 out of 10 mark is not yet held consistently. What converges is the number of new ways of reading a file that each sample reveals: 13, then 7, then 1, then 0.
The checks for other agents followed the same rule before release: 30 correct warnings out of 31, 53 out of 53 and 3 out of 3 depending on the check, then 877 out of 878 on 513 repositories never seen before. Since 0.21, 202 public test cases, most of them born from a false positive once found on a real repository, run on every change on all three systems.
What a parser cannot see
Two defects escape any mechanical reading: two instructions that cannot both be followed, and one rule copied into two files that will drift apart. Since 0.21.1, a separate command hands them to a model, called through your own Claude Code session.
Each finding must quote both passages word for word, and each quote is checked against the file it names: a quote that is not there gets the finding rejected. Measured result: 50 correct contradictions out of 50, 49 correct duplicates out of 50. But of 24 defects that maintainers declared fixing in their own commits, it finds only 4 at their spot, 6 counting the same conflict seen in another file.
Two successive runs agree on only about three findings in four. So this result never enters the floor: a number that moves from one run to the next would fail the CI while nothing had changed.
Nor is the model asked whether a rule is useful. A model judging a rule finds clear and useful the rule it would have followed anyway, which is exactly the rule that does nothing. That is measured by running the agent with and without the rule, not by reading it.
What it does not do
deadweight fixes nothing. It proposes, and the maintainer decides: the script deliberately has no --fix option, because a configuration encodes decisions the tool does not know about.
It counts what is on disk, not what the agent actually loaded. To know that, --runtime reads the log of an InstructionsLoaded hook that you add yourself: the tool never installs one. Finally, each measurement is dated by Claude Code version: when Claude Code changes, some must be redone. The tool says so.
Try it
claude plugin marketplace add e-xode/deadweight
claude plugin install deadweight@e-xode --scope projectThen run /deadweight:config-auditor in your project. It is free, under the MIT license: https://github.com/e-xode/deadweight
If your team has rolled out coding agents across several repositories and nobody really knows any more what gets loaded, that is the kind of work I take on: https://www.e-xode.net/en/contact

Comments (0)