Skip to content

MCP Tools

The complete tool surface Agentic HIL exposes over MCP, what each group does, and how a run is put together from it. Host registration syntax is in MCP host configuration; the permissions each of these calls is judged by are in the authoritative configuration.

MCP Entry

Every MCP host starts the same local stdio server from the firmware project root, using the reviewed absolute path of its persistent installation:

/absolute/path/to/persistent/agentic-hil mcp-stdio

Host configuration schemas are not portable: VS Code uses servers, Claude Code uses mcpServers, Codex uses TOML, and OpenCode uses a command array. agentic-hil agent-install --agent <agent>, which agentic-hil setup --agent <agent> runs first, performs secure user-level registration for Claude Code, Codex, and OpenCode. See MCP host configuration for the remaining hosts. agentic-hil mcp-config --output .mcp.json generates only a machine-local Claude-compatible form with an absolute executable path; keep it uncommitted.

mcp-stdio discovers the authoritative file from its project working directory or, started outside any project, from the folder the host names in its MCP roots. The authoritative configuration describes where it is found and how an absolute-path override is selected.

The tool surface

Group Tools Notes
Debugger debugger_info, debugger_probes_list, probe_target, reset_target Probe-ID listing runs the backend's own enumeration on pyOCD and STM32CubeProgrammer; OpenOCD has none, so on an OpenOCD bench the ids come from this host's USB serial inventory, which names an ST-Link and no other adapter, and discovered_by says which enumeration answered. An OpenOCD entry whose interface_cfg names another adapter still refuses not_supported. reset_target modes run and halt work on every backend; init also runs the target's reset-init event script and is OpenOCD-only: stlink and pyocd refuse it with not_supported rather than halting instead. On pyocd probe_target connects with the target pack's DebugCoreStart sequence disabled, as the reads do, so probing a core reset_target halted leaves it halted
Firmware flash_firmware, artifact_upload artifacts are validated, rechecked, and copied to private process staging before flashing; allow_reset is additionally required when reset_after_flash is requested
Serial com_ports_list, com_session_start, com_session_stop, com_write, com_read named ports only, buffered background reader
CAN can_buses_list, can_session_start, can_session_stop, can_send, can_read PEAK, SocketCAN, or a process bridge; buses with shares: use a named participant on every session call and run declaration
Diagnostics get_last_report, classify_last_error structured error classification with likely causes
Project setup project_config_create generates this workspace's configuration from attached hardware when it has none; takes no arguments and writes every permission true except allow_raw_debugger_commands and allow_mass_erase, which are false so that flashing works, so the bench is workable from the file it produces. Regenerating an existing one carries that file's permissions onto the entries still in it, gives an entry discovered for the first time those same defaults, and reads a probe under the same locks and audit trail as every other hardware call
Project config project_config_describe, project_config_set, project_config_adopt_hardware, project_config_reload_description field-wise changes to an existing configuration, gated by allow_config_description_write (what the bench is) and allow_config_permissions_write (every permission key: the permissions: blocks plus artifacts.allow_upload and debug.allow_all_symbols). describe needs no permission and says which keys are open in this state; set takes named keys with scalar values; adopt_hardware reads the attached probe and fills in the identity keys that are still unset, through set; reload_description makes a changed description the one this running server answers out of, re-reading no permission at all and needing no grant because it writes nothing
Debug sessions debug_start_session, debug_stop_session, debug_get_session_status, debug_set_breakpoint, debug_list_breakpoints, debug_clear_breakpoints, debug_continue, debug_halt, debug_get_stop_reason, debug_symbol_info, debug_symbol_value, debug_dump_symbol_ihex typed GDB/MI sessions through the backend's own GDB server, OpenOCD's, pyocd gdbserver, or on stlink the ST-LINK_gdbserver STM32CubeCLT installs, each on a port reserved for the session, with the same tools, payloads, audit and teardown on all three; unexpected breakpoints and target exceptions are returned as structured stop reasons; symbol allowlist and dump-size limits come from the debug: config section. debug_continue is the one call that resumes a halted target and the one gated by allow_debug_execution; every other call in this row reads or holds the session instead and needs no more than allow_probe already covers. debug_stop_session reconfirms the halt before it tears the session down, and reports safe_state_confirmed/halt_not_confirmed rather than a clean stop when it cannot. It also proves the session's breakpoints off the target before a server that has to be ended ahead of GDB's detach is ended (pyocd gdbserver and ST-LINK_gdbserver, neither of which can be asked to leave the core halted once its client goes), and reports breakpoints_removed_confirmed: false with breakpoints_not_removed rather than a clean stop where the backend's own list does not come back empty. The session runs GDB with asynchronous MI so that a running target can be interrupted: a debug_continue that reaches its timeout halts the target and says so, debug_halt on a target that is already stopped answers the stop it is in from the session's own record instead of waiting for one that never comes, and a GDB that refuses -gdb-set mi-async on is refused at debug_start_session with gdb_async_unsupported before anything on the target was touched. The three symbol tools answer three questions about one allowed symbol: debug_symbol_info where it is and how large it is, debug_symbol_value what it currently holds (hex in memory order, and at 1, 2, 4 or 8 bytes also value_unsigned and value_signed, read in the byte order the image declares in its ELF header, with byte_order naming that order and byte_order_from saying whether it was read off the image or assumed), and debug_dump_symbol_ihex those same bytes written to a file. All three resolve the symbol out of the ELF's debug information first and fall back to the ELF's own symbol table where there is no type to work with, which is what makes an assembly-defined object such as a vector table resolvable at all; resolved_from names the route that answered, and a symbol neither route describes is refused as before. debug_symbol_value is bounded by debug.max_dump_size_bytes like the dump, and leaves no file behind. All three are also the members that run on the stlink and pyocd backends, where STM32CubeProgrammer's -r and pyOCD's savemem read target memory without a session: same arguments, same refusals, same result fields, with the symbol resolved offline from the ELF a confirmed flash_firmware put on the board using debug.gdb_executable. On both, the value tool parses the backend's own memory dump back out of a private file it creates and removes, and a dump that does not cover the bytes that were asked for is a failed read rather than a short answer; on pyocd the Intel HEX is written here too, because savemem writes a raw binary. debug_symbol_info opens no probe at all on either, because an address and a size are properties of the image. The reads that do reach the board connect with the least intrusive connect each tool documents, STM32CubeProgrammer's mode=HOTPLUG and pyOCD's --connect attach: a read that reset or halted the target would destroy or freeze the RAM it was asked for. On pyocd the read also disables the target pack's DebugCoreStart sequence for its connect, because that sequence lets a halted core run, so a core reset_target halted stays halted across the reads after it. debuggers.<name>.connect_mode reaches neither, being the flash's key on stlink and nothing at all on pyocd. While a session is open on pyocd or stlink, the three read through it instead, so a read and a breakpoint share one connection to the probe. Everything else in this row runs on openocd and pyocd, and on stlink wherever debuggers.<name>.gdb_server_executable names ST-LINK_gdbserver or it is found beside the configured STM32_Programmer_CLI; an stlink entry without one refuses it with not_supported, and that refusal names both ways out rather than only the backend: the server, or the same probe under type: openocd with the interface script for it, plus what that move costs
Run boundary bench_run_start, bench_run_stop, bench_run_status declares the devices of a multi-call run and holds them for its whole duration; without it each call holds its device only for its own duration
Test reactor test_reactor_run, test_reactor_status, test_reactor_stop runs a reviewed test plan, through the same code agentic-hil test-reactor drives: the same preflight, the same lock declaration, the same per-step permissions, the same report. The plan is its own declaration, so test_reactor_run needs no bench_run_start around it and holds every device the plan names from before its first step to after its last; asked for while this session's own bench_run_start is open, it is refused as run_already_active with bench_run_stop as the way out. test_config_path selects another plan and has to resolve inside workspace_root; detach: true answers at once with a run handle instead of holding the call open, which test_reactor_status reads and test_reactor_stop ends after the step it is in. What a plan may contain is in Running hardware tests, and is published to an agent as the MCP resources agentic-hil://reference/test-plan (where a plan lives, how its path resolves, which version admits which step, every step with its keys, the comparators, two worked plans) and agentic-hil://reference/test-plan-schema (the schema the first is generated from)
Recovery hardware_lease_status, hardware_recover hardware_lease_status reads who holds the bench and any open incident, with its cleanup_reasons, quarantine_guidance and whether it stands, and changes neither: it is the record agentic-hil lease-status prints. hardware_recover clears a quarantine, gated by permissions.allow_recover. Reasons that name a call which never reached the hardware clear with no arguments; any other reason needs operator_statement. The agent asks the operator in chat and passes back what they said, which the recovery ledger keeps verbatim under an attestation label kept distinct from the operator's own signature. There is no confirmation boolean and will not be: a flag an agent sets for itself is not a confirmation, while a sentence about a bench has to come from somebody. The agentic-hil recover --confirm-safe-state --quarantine-id <id> line stays in every refusal as the operator's direct route, and the audit_broken families still take it
Installation server_upgrade lifts this installation to the newest release, gated by permissions.allow_upgrade. There is no version argument, so it can only go forward: an agent that could name a version could install one that reads this file's permissions differently. It replaces the package on disk and not the code this server is running: a successful result carries previous_version, version, running_version and restart_required: true, and an operator restarting the MCP server is what loads it. Refused while a run or session holds the bench (upgrade_in_open_run), on Windows, where a running process's files are locked and agentic-hil upgrade at a shell is the way (upgrade_cli_only_on_host), and on a pinned installation (upgrade_blocked_by_pin, which names an extras-preserving reinstall command and never runs it)

A typical loop: build firmware → bench_run_start naming the probe and the port → flash_firmware with reset_after_flash: true when a fresh boot is required → com_session_start → stimulate via com_write/can_send → assert on com_read/can_read → bench_run_stop → on failure, classify_last_error. A sequence worth keeping is written as a plan instead and run with test_reactor_run, which is the same loop stated once in a file that diffs like code.

Devices are named as {"kind": "debugger" | "uart" | "can", "id": "<config entry>"}. The lock is keyed on the hardware behind the entry, so two entries describing one physical unit (a probe and its virtual COM port under a shared resource_id) collapse to one lock. The DUT is not a device: it is what the devices drive.

Every result is structured JSON (ok, error_type, summary, likely_causes, report_path, log_path), and each hardware action is validated against the authoritative configuration, executed with a timeout, and logged to .agentic-hil/logs/. Which calls a host should treat as read-only or destructive is published per tool in the MCP annotations object; see Tool Annotations.

Agent-facing routing (which tool answers which request, and what to do with a refusal) lives in AGENTS.md.