Skip to content

Safety Model

Leave an agent alone with real bench hardware and still trust the board, the host, and the logs afterwards. The agent can edit every file inside the workspace, so nothing there is trusted. All authority sits outside, beyond its reach:

Trusted policy boundary: the agent-writable workspace talks to Agentic HIL over MCP stdio only; the authoritative config and state_root (leases, quarantine, audit chain) live on the operator-controlled host outside the workspace Trusted policy boundary: the agent-writable workspace talks to Agentic HIL over MCP stdio only; the authoritative config and state_root (leases, quarantine, audit chain) live on the operator-controlled host outside the workspace

The action gate

Every hardware action, from every entry point, walks the same gate:

Every tool call passes a deny-by-default permission gate, then validation, then a cross-process owner lease, before executing with a pinned executable and a timeout. Success appends to the SHA-256 audit chain and returns a structured JSON result; a broken audit chain quarantines the bench until an operator recovers it. Every tool call passes a deny-by-default permission gate, then validation, then a cross-process owner lease, before executing with a pinned executable and a timeout. Success appends to the SHA-256 audit chain and returns a structured JSON result; a broken audit chain quarantines the bench until an operator recovers it.

The gate is in the tool rather than in the agent's host because a probe, a port and a CAN adapter carry no permission model of their own, so enforcement has to sit directly in front of the hardware, the way a database engine and not its client holds the grants. A host's permission system is a second layer in front of this one and answers a different question; docs/security-design.md has why both exist.

Permissions and validation

  • What is denied by default is presence: a device does not exist on this bench until the authoritative configuration declares it, and a call naming any other one is refused before a driver is opened: unknown_device where a run declares it, and com_port_not_configured or can_bus_not_configured where a port tool or a bus tool names it, each of those two naming the ones the configuration does declare. On a device it does declare, a generated configuration grants every permission that has a tool behind it (allow_flash, allow_reset and allow_debug_execution on a probe, allow_write on a port and on a bus, allow_all_symbols and allow_upload on the project) and holds allow_raw_debugger_commands and allow_mass_erase false, the interlocked pair that has no tool behind it and refuses flashing while either is true. An operator takes one grant back with agentic-hil revoke <key> and reopens it with agentic-hil grant <key>, one named permission at a time, from their own shell and from nowhere else. Over MCP the direction is fixed: an agent narrows its own authority and never widens it, because a permissions write there may carry false and never true. permission_denied results are authoritative and agents are instructed to stop (see AGENTS.md).
  • Opening a debug session and inspecting a halted target need nothing beyond allow_probe (or, from version 2 on, nothing at all): the same read-needs-no-grant rule that covers the rest of this file, because an attach halts the core the way any other read does. Resuming it is a separate act and a separate grant, allow_debug_execution, checked on debug_continue and nowhere else it could be reached from. A session that ends without the target's halt reconfirmed is not reported as a clean stop: it is held open for retry the way any other unconfirmed hardware effect is, safe_state_confirmed: false and halt_not_confirmed: true in the result, rather than let ending the session itself be the undocumented way to leave the core running. The breakpoints a session set are proved off the target on the same terms. A GDB that detaches takes them off itself, which is the proof wherever the debug server is still there to carry the detach; where the server has to be ended first so that it never resumes the core (pyocd gdbserver and ST-LINK_gdbserver), the session deletes them and reads the backend's own list back before that, and a removal the backend does not confirm is reported as breakpoints_removed_confirmed: false and held open for retry rather than returned as a clean stop.
  • Reading needs no permission, and exclusivity carries that load instead. The risk in a HIL tool is not damage but a false green: a test disturbed unnoticed produces a result nobody should trust. Whoever holds a board cannot be disturbed, and whoever reads while no run holds it disturbs nobody. listen_only for CAN remains the way to observe a target provably undisturbed. A port opened with assert_dtr and assert_rts false keeps both lines released for the session, which is the least a serial observation does to a target, and it is not proved untouched by that: on Linux, measured with an FT232R over 65 opens, the open itself asserted DTR once for 239 to 943 microseconds before it was released for the rest of the session and at close, so a board that wires DTR to reset sees that pulse.
  • listen_only is enforced per adapter rather than assumed, because the three adapters obtain the mode in three different ways. peak sets it through PCAN's BusState.PASSIVE, re-asserts it once the channel is initialized, and reads PCAN_LISTEN_ONLY back from the driver. socketcan reads the kernel's control mode and never sets it: the mode belongs to the CAN netdev and only ip link puts it there, so an interface that is up without listen-only refuses the session rather than being joined by an ACKing node. A process bridge is sent the flag and must answer listen_only: true from its open result; forwarding confirms nothing. An adapter that cannot be held to it refuses with can_listen_only_unsupported (before the bus is touched) or can_listen_only_unconfirmed (asked, not confirmed, adapter closed again), never with a silent downgrade; can_buses_list reports listen_only and listen_only_enforcement per bus. Transmitting on such a bus is refused rather than attempted: can_send, a test plan's can_send step and a broker participant's send all answer can_listen_only_mode before any driver is called, whatever permissions.allow_write says: the mode is settled first and the permission is never reached, because a bus is not made transmit-capable by granting a permission on it. Enforcement and verification reach exactly as far as the driver's or the kernel's own report of its mode; beyond that this product makes no claim.
  • Configured executables and OpenOCD scripts are pinned at startup and must resolve outside the workspace; firmware artifacts are validated, hashed, and staged in a private process directory before any backend consumes them.
  • The MCP launcher that gets registered is pinned the same way and must be owned by the operator or by root, executable, and writable by nobody else, where group write counts as another writer only when the group is not the owner's own user-private group (its name is the account's, its gid is the account's primary gid, and it has no other member), so the group-writable console script a default Debian or Ubuntu umask 0002 produces is accepted as it stands while world write, a foreign group, and ownership by a third account stay refused with the failing condition and the path element named.
  • Serial/CAN writes and reads are size- and buffer-capped; debugger calls run with timeouts and OpenOCD's TCP servers disabled.
  • One live owner per project and physical probe/port/bus across all entry points; a second process gets resource_busy, and a crash or an unknown effect opens an incident on the resource (TROUBLESHOOTING.md). An incident is raised for a genuinely unknown physical state: a failure that proves it never reached the hardware (a missing toolchain, an unplugged probe or adapter, an unopenable port, a script refused before the adapter opened, a read whose backend reports that no target answered) refuses with a named error, target_contacted: false and retry_safe: true instead. The proof has to be the backend's, not the tool's category: a read killed at its deadline still counts as contact, because an SWD attach halts the core and a killed process never ran its own shutdown.
  • An incident is not a padlock on the bench. A run that hits one aborts with the verdict failed (recovering the bench never un-fails a test), and the abort drives a recovery action: reap, reset into halt where recovery.auto_recover and the probe's allow_reset permit it, then a probe re-read, all reported in the run result's recovery block. The calls that are the remedy (probe_target, reset_target, flash_firmware) run while an incident is open and clear it when they answer the reason it names, with a ledger line.
  • A gate is only owed where the missing proof cannot come back on its own, so the standing, human-visible quarantine is the audit halt and nothing else. A target's state is proven by the next reset into halt and an answering probe; a serial handle's and a CAN adapter's by their own next open, which the operating system refuses by itself if the handle is really stuck. Those incidents end when the call that raised them ends, leaving a no_standing_state line in the recovery ledger that names every reason they were held for. The evidence chain is the one thing no later contact rebuilds, because no reset writes a report that was never written: an audit_broken incident holds the bench, refuses the stimulus class (com_write, can_send, session starts, new runs) with resource_quarantined, and is cleared only through agentic-hil recover --confirm-safe-state or, where permissions.allow_recover says so, hardware_recover carrying the operator's own words as operator_statement. Both routes answer nothing_to_recover on everything else.
  • A result that names a cleanup_reason carries quarantine_guidance for it either way: what was attempted, what is confirmed, what remains unknown, and the physical check to perform (for the operator, or for the agent to relay to them).
  • A debugger result is judged by the backend's own report of the run and not by the words anywhere in its transcript: an OpenOCD run that exits 0 and prints the success marker its command string echoes after the operation is a success, because OpenOCD stops evaluating that string at the first command that fails and could never have reached the echo otherwise. The failure-worded lines such a run printed come back on the successful result as backend_warnings, verbatim, so a noisy build stays visible without a reset the board performed being reported as a failed one. A run without the marker, or one that exited non-zero, is classified exactly as before.

Exclusivity, incidents, audit

  • Exclusivity per physical device is machine-wide and held for the whole run, kept in one agreed place (~/.agentic-hil/device-locks) rather than beside a configuration, because two sessions with different state_root values are still one bench. A contender is refused with device_busy naming the holder; waiting happens only when asked for (--wait-s) and is bounded. Ownership is the operating system's lock, so a crashed run frees its devices immediately, with no quarantine and no manual recovery.
  • Canonical reports and the tamper-evident audit chain live under the operator-pinned state_root; workspace logs and reports are untrusted mirrors, verified against the chain on read.

The complete threat model and design rationale: docs/security-design.md. What each permission opens, and which ones a generated configuration writes false: the authoritative configuration.