Tool Definition Score¶
Every pull request that changes the MCP tool definitions is scored before it
merges. The gate exports the tools/list of the base and of the head, scores
both with the Tool Definition Quality Score (TDQS) v1.3 using one pinned model,
and fails when the mean TDQS of the tools falls. A pull request that leaves the
tool definitions unchanged passes without a single model call.
The scores are the same kind a public MCP registry publishes for this server. How close the gate's scores come to the registry's is measured below, in Calibration against the registry.
What is scored¶
- The definitions a host sees. Each side is exported by asking a
provisioned server for
tools/listover stdio: name, title, description, input schema, output schema and annotations of every tool. The base is exported from a detached worktree of the base revision, with the interpreter and dependencies of the environment running the gate; only the package itself comes from the base. Both exports are hashed, and two equal hashes mean nothing to score. - The specification. TDQS v1.3 at commit
b9881b0cfec88969e42672c92544487ca191a992of the specification repository. Its four prompts are fetched from that commit, checked against the SHA-256 hashes intools/tdqs/version.json, and cached outside the repository, inagentic-hil-tdqsunderXDG_CACHE_HOME,LOCALAPPDATAor~/.cache. The prompt texts are never stored in this repository, because the specification repository grants no license to copy them. A prompt that no longer matches its hash stops the gate. - The model.
claude-opus-5-5, asked through the Claude Code command line 2.1.288, whichtools/tdqs/package-lock.jsonpins by version and integrity. Each tool is one call, the coherence of the whole set is one more, and an answer that cannot be read is asked again up to two times. The calls run with no tools, no MCP servers, no settings, no project memory and no saved session, each in a fresh empty directory. An answer from any other model is refused. - The arithmetic. Each tool's TDQS weighs purpose clarity 25, usage guidelines 20, behavioral transparency 20, parameter semantics 15, conciseness and structure 10, and contextual completeness 10. Description quality is 60 % of the mean TDQS plus 40 % of the lowest one, so a single weak tool costs the whole set. The overall score is 70 % description quality plus 30 % coherence. Every rollup is computed on exact fractions and rounded half up to one decimal once. Tiers start at 3.5 (A), 3.0 (B), 2.0 (C) and 1.0 (D).
The decision¶
The mean TDQS over all tools decides, compared as its rollup: the exact mean rounded half up to one decimal once. The overall score, description quality, coherence and the lowest TDQS are reported next to it as information. They do not decide because they are not steady enough on identical text: the overall weighs the single lowest tool at 28 % and the judged coherence at 30 %. Over three runs of the same definitions the lowest TDQS moved from 1.7 to 2.1 and the overall from 2.8 to 3.0; over three runs of another set, coherence moved from 3.3 to 3.5. The mean over every tool moved by less than 0.05 unrounded in both, and it still falls when descriptions get worse.
| Result | Exit code | When |
|---|---|---|
| PASS | 0 | The definitions are unchanged; or the first pair holds (head mean TDQS at least base mean TDQS); or the first pair drops but the median of each side over three pairs does not |
| BLOCK | 1 | The first pair drops and two more pairs confirm it: the head's median mean TDQS over the three pairs is below the base's |
| INVALID | 2 | The definitions changed and could not be scored completely: no token, a prompt that fails its hash, an answer still unreadable after its retries, a different model answering, or a saved report that no longer matches the exports |
One scoring run of the same definitions can still differ from the next by 0.1 on the mean TDQS once it is rounded (the three calibration runs below gave 3.2, 3.2 and 3.3, from 3.207 to 3.255 unrounded, against 2.8, 3.0 and 2.9 on the overall score), so a drop on one pair alone never blocks.
The size of the definitions is reported next to the scores and is never part
of the decision: description characters, tools/list bytes and the bytes of
each tool's definition, for base, head and the change between them.
The report is written to .tool-definition-score/tool-definition-score.json,
with a Markdown summary beside it in tool-definition-score.md; in CI the
summary is also the job summary. When the mean TDQS drops on the first pair,
the report names the tools that fell, the dimensions that fell with the model's
justification for each, and what each tool's fall alone costs the unrounded
mean TDQS; the coherence dimensions that fell and a fallen minimum are listed as
information.
Running it locally¶
The gate needs the development install (python -m pip install -e '.[dev,can]')
and the claude command on PATH, logged in. The pinned command line is one
npm ci away:
cd tools/tdqs && npm ci && cd ../..
export PATH="$PWD/tools/tdqs/node_modules/.bin:$PATH"
python tools/tool_definition_score.py --base origin/master
Another installed version of Claude Code also works; the report then carries a
warning that the scores were calibrated with 2.1.288. On Windows the gate
refuses a claude.cmd batch wrapper, which cannot pass the system prompt
intact, and needs the native claude executable.
| Option | Meaning |
|---|---|
--base REV |
the revision to compare against, for example origin/master |
--head REV |
a head revision; without it the working tree is the head |
--report-dir DIR |
where the report goes (default .tool-definition-score) |
--token-env NAME |
the variable holding a Claude Code OAuth token; without it the local login is used |
--concurrency N |
model calls at once (default 6) |
--cache-dir DIR |
where the fetched specification is cached |
--check REPORT |
check a saved report against the current exports instead of scoring |
--calibrate CAPTURE with --runs N |
score a capture of the registry's published definitions N times and write calibration.json |
One pair of the 44 tools takes about two and a half minutes and 90 model calls at the default concurrency; a confirmed drop takes three pairs.
The end-to-end test of the gate against the real model is left out of every
test run unless AGENTIC_HIL_TDQS_MODEL is exactly 1:
AGENTIC_HIL_TDQS_MODEL=1 python -m pytest tests/test_tool_definition_score_model.py
In CI¶
The tool_definition_score job in .github/workflows/ci.yml runs on pull
requests only, since only a pull request has a base to compare. It installs the
pinned command line with npm ci, scores the merge of the pull request against
origin/<base branch>, runs the end-to-end test when the model was called, and
uploads the report as the tool-definition-score artifact whatever happened.
Required CI needs the job, and accepts it skipped only on a push.
The model is reached with the repository secret CLAUDE_CODE_OAUTH_TOKEN, a
token printed by claude setup-token for the account that pays for the calls.
Only the scoring step and the end-to-end step receive it. A pull request from a
fork receives no secrets: with its tool definitions unchanged it passes as
usual, with them changed it is INVALID, because nothing could score it.
tools/tdqs/version.json records everything a score depends on: the model, the
command line version, the specification commit, the prompt hashes, the
weights, the tier thresholds, the retries and the confirmation rule. Changing
any of it changes the record's digest, and the calibration record below carries
the digest it was made with, so a test fails until the calibration is made
again.
Calibration against the registry¶
The Glama registry
published TDQS scores for release 0.22.1-dev.0, scored on 2026-10-01. The
definitions it scored are those of 5762f9b in every field but one: its copy
carries no annotations.title on any of the 44 tools, while this server's
tools/list carries one on each. So the calibration scores the registry's own
definitions, and both evaluators judged the same text.
The gate scored those definitions three times; per tool and per dimension, the
median of the three is set against the registry's published score. Nothing is
adjusted: the gate's scores stay what the pinned model gives, and the
differences are recorded here and in tools/tdqs/calibration.json with the
version digest, the scored set hash and the definition differences.
The whole set¶
| Measure | Registry | Gate runs | Gate median |
|---|---|---|---|
| Overall | 3.1 (B) | 2.8, 3.0, 2.9 | 2.9 (C) |
| Description quality | 2.8 (C) | 2.6, 2.8, 2.7 | 2.7 (C) |
| Coherence | 3.8 (A) | 3.3, 3.3, 3.3 | 3.3 (B) |
| Mean TDQS | 3.6 | 3.2, 3.2, 3.3 | 3.2 |
| Lowest TDQS | 1.7 debug_symbol_info |
1.7, 2.1, 1.8 debug_symbol_info |
1.8 |
Scored once on 5762f9b's own tools/list, with the titles, the gate gives
2.9 (C) overall, 2.7 description quality, 3.3 coherence, 3.2 mean TDQS and
2.0 lowest (debug_symbol_info).
Per dimension¶
The mean absolute difference between the gate's median and the registry's score, over the 44 tools:
| Measure | Mean absolute difference | Mean signed difference | Gate higher | Gate lower | Equal |
|---|---|---|---|---|---|
| TDQS | 0.40 | -0.35 | 4 | 37 | 3 |
| Purpose clarity | 0.48 | -0.30 | 4 | 17 | 23 |
| Usage guidelines | 0.23 | -0.18 | 1 | 9 | 34 |
| Behavioral transparency | 0.59 | -0.55 | 1 | 24 | 19 |
| Parameter semantics | 0.25 | -0.16 | 2 | 9 | 33 |
| Conciseness and structure | 0.50 | -0.41 | 2 | 19 | 23 |
| Contextual completeness | 0.68 | -0.68 | 0 | 30 | 14 |
The largest differences¶
| Tool | Registry | Gate | Difference | Dimensions that differ (gate minus registry) |
|---|---|---|---|---|
bench_run_start |
4.4 | 3.6 | -0.8 | purpose clarity -1, behavioral transparency -1, parameter semantics -1, conciseness and structure -1, contextual completeness -1 |
debug_start_session |
2.9 | 2.1 | -0.8 | purpose clarity -1, behavioral transparency -1, parameter semantics -1, conciseness and structure -1, contextual completeness -1 |
flash_firmware |
4.5 | 3.7 | -0.8 | purpose clarity -1, usage guidelines -1, behavioral transparency -1, parameter semantics -1, conciseness and structure +1, contextual completeness -1 |
test_reactor_run |
4.3 | 3.5 | -0.8 | purpose clarity -1, behavioral transparency -2, contextual completeness -1 |
project_config_set |
4.0 | 3.3 | -0.7 | usage guidelines -1, behavioral transparency -1, parameter semantics -1, conciseness and structure -1, contextual completeness -1 |
artifact_upload |
3.3 | 2.7 | -0.6 | behavioral transparency -1, parameter semantics -1, conciseness and structure -1, contextual completeness -1 |
classify_last_error |
3.4 | 2.8 | -0.6 | purpose clarity -1, behavioral transparency -1, conciseness and structure -1, contextual completeness -1 |
debug_continue |
3.3 | 2.7 | -0.6 | usage guidelines -1, behavioral transparency -1, conciseness and structure -1, contextual completeness -1 |
debug_stop_session |
2.8 | 2.2 | -0.6 | purpose clarity -1, parameter semantics -1, conciseness and structure -1, contextual completeness -1 |
get_last_report |
3.4 | 2.8 | -0.6 | purpose clarity -1, behavioral transparency -1, conciseness and structure -1, contextual completeness -1 |
Likely causes¶
- A stricter judge, by a constant offset. The gate scores with
claude-opus-5-5; the registry does not say which model it scores with, and its justifications read differently. The gate is lower on 37 tools and higher on 4, by 0.35 TDQS on average, but it orders the tools much as the registry does: per tool, the gate's median and the registry's score correlate at 0.89, and for 32 of the 44 tools the difference lies within 0.3 of that 0.35 offset. A judge that disagreed about the tools would scatter; this one sits one step lower on nearly all of them. The offset comes mostly from contextual completeness (lower on 30 tools, higher on none), behavioral transparency (lower on 24, higher on 1) and conciseness and structure, while usage guidelines and parameter semantics match the registry on 34 and 33 tools. - Dense descriptions. The largest differences, 0.6 to 0.8, fall on
descriptions that carry several rules in one or two sentences or point to a
resource for details:
bench_run_start,flash_firmware,test_reactor_run,project_config_set. Each is one or two points lower on three to five dimensions rather than far lower on one, and below the registry in every run. No tool differs by more than 0.8, and the four the gate rates higher differ by at most 0.4. - Run-to-run variation. The gate's own three runs differ from each other by 0.10 TDQS per tool on average and by up to 0.6, and gave the same coherence in all three; the median of three narrows that further, but the registry's single published score carries its own variation, which one score cannot show.
- The request around the prompts. The command line puts a billing line and one sentence naming the SDK ahead of the system prompt, and a reminder with the date ahead of the prompt itself, and samples at its default temperature. None of that can be switched off, and its effect on the absolute level is not known. It is the same for both sides of a comparison.
- Not the definitions. The calibration scored the registry's own
definitions, so the missing
annotations.titleis not a cause here.
Across the whole set the offset shows in every rollup: the mean TDQS is 3.2
against the registry's 3.6, the overall 2.9 against 3.1, and coherence 3.3
against 3.8. The lowest tool is the same, debug_symbol_info, at 1.8 against
1.7.
The gate compares a head with its own base under the same model, so an offset the pinned model holds against the registry applies to both sides and leaves the decision alone. What it cannot see is a change the two models would judge in opposite directions; the close per-tool agreement above makes that rare.
Every tool¶
| Tool | Registry | Gate (median) | Difference | Gate runs |
|---|---|---|---|---|
artifact_upload |
3.3 | 2.7 | -0.6 | 2.7, 2.7, 2.7 |
bench_run_start |
4.4 | 3.6 | -0.8 | 3.6, 3.6, 3.8 |
bench_run_status |
3.9 | 3.8 | -0.1 | 3.8, 3.7, 3.8 |
bench_run_stop |
4.2 | 4.1 | -0.1 | 4.1, 4.1, 4.1 |
can_buses_list |
4.4 | 4.2 | -0.2 | 4.2, 4.2, 4.1 |
can_read |
3.5 | 3.8 | +0.3 | 3.8, 3.8, 3.9 |
can_send |
3.2 | 3.2 | 0.0 | 2.9, 3.4, 3.2 |
can_session_start |
2.9 | 2.6 | -0.3 | 2.6, 2.5, 2.7 |
can_session_stop |
2.7 | 2.4 | -0.3 | 2.4, 2.4, 2.4 |
classify_last_error |
3.4 | 2.8 | -0.6 | 2.8, 2.7, 2.8 |
com_ports_list |
4.4 | 4.1 | -0.3 | 4.1, 4.2, 4.1 |
com_read |
3.8 | 4.1 | +0.3 | 4.1, 3.5, 4.1 |
com_session_start |
2.9 | 2.7 | -0.2 | 2.7, 2.7, 2.7 |
com_session_stop |
2.6 | 2.4 | -0.2 | 2.2, 2.4, 2.6 |
com_write |
2.9 | 2.9 | 0.0 | 2.9, 2.9, 2.9 |
debug_clear_breakpoints |
3.8 | 3.4 | -0.4 | 3.4, 3.4, 3.4 |
debug_continue |
3.3 | 2.7 | -0.6 | 2.7, 2.9, 2.7 |
debug_dump_symbol_ihex |
3.0 | 2.7 | -0.3 | 2.7, 2.7, 2.7 |
debug_get_session_status |
3.2 | 2.7 | -0.5 | 2.7, 2.7, 2.7 |
debug_get_stop_reason |
3.1 | 2.7 | -0.4 | 2.7, 2.7, 2.7 |
debug_halt |
3.0 | 2.5 | -0.5 | 2.6, 2.5, 2.5 |
debug_list_breakpoints |
3.7 | 3.3 | -0.4 | 3.5, 3.1, 3.3 |
debug_set_breakpoint |
3.4 | 2.9 | -0.5 | 2.9, 2.9, 2.9 |
debug_start_session |
2.9 | 2.1 | -0.8 | 2.1, 2.2, 2.1 |
debug_stop_session |
2.8 | 2.2 | -0.6 | 2.2, 2.2, 2.2 |
debug_symbol_info |
1.7 | 1.8 | +0.1 | 1.7, 2.1, 1.8 |
debug_symbol_value |
3.8 | 3.3 | -0.5 | 3.3, 3.3, 3.1 |
debugger_info |
3.9 | 3.8 | -0.1 | 3.8, 3.8, 3.8 |
debugger_probes_list |
4.4 | 4.2 | -0.2 | 4.1, 4.2, 4.2 |
flash_firmware |
4.5 | 3.7 | -0.8 | 3.7, 3.7, 3.9 |
get_last_report |
3.4 | 2.8 | -0.6 | 2.8, 2.8, 2.8 |
hardware_lease_status |
3.9 | 3.8 | -0.1 | 3.9, 3.8, 3.8 |
hardware_recover |
3.3 | 3.7 | +0.4 | 3.9, 3.7, 3.6 |
probe_target |
3.4 | 3.0 | -0.4 | 3.0, 3.0, 3.0 |
project_config_adopt_hardware |
4.1 | 3.7 | -0.4 | 3.7, 3.7, 3.7 |
project_config_create |
4.3 | 3.7 | -0.6 | 3.7, 3.3, 3.7 |
project_config_describe |
4.6 | 4.2 | -0.4 | 4.2, 4.2, 4.4 |
project_config_reload_description |
4.4 | 3.8 | -0.6 | 3.8, 3.8, 3.8 |
project_config_set |
4.0 | 3.3 | -0.7 | 3.3, 3.2, 3.6 |
reset_target |
3.6 | 3.2 | -0.4 | 3.7, 3.1, 3.2 |
server_upgrade |
4.1 | 4.1 | 0.0 | 4.2, 4.1, 4.1 |
test_reactor_run |
4.3 | 3.5 | -0.8 | 3.5, 3.5, 3.7 |
test_reactor_status |
3.4 | 2.8 | -0.6 | 2.8, 2.8, 3.0 |
test_reactor_stop |
3.5 | 2.9 | -0.6 | 2.9, 2.9, 2.9 |
Making the calibration again¶
Capture the registry's published scores as a JSON file of the form
{"release": ..., "server": {...}, "tools": [{"name", "definition", "qualityScore"}]}
from the registry's page for the server, then:
python tools/tool_definition_score.py --calibrate registry.json --runs 3 --report-dir calibration-run
cp calibration-run/calibration.json tools/tdqs/calibration.json
The record keeps only the registry's numbers, tiers and scoring time, not its justification texts. It refuses a run whose scored definitions are not the registry's.