Skip to content

Running Hardware Tests

Two entry points drive the same permission-gated tools without an MCP client: the test reactor executes a reviewable YAML plan, and the pytest plugin puts the bench behind a fixture.

Test Reactor

Write a hardware test as a plan, not a script: one reviewable YAML file describes flash → stimulate → break → dump, and the reactor guarantees it is either executed exactly as written or rejected before the first hardware action. No half-run plans, no leftover breakpoints, no orphaned debug or UART sessions. The same file behaves identically on your bench and in CI, and it diffs like code in a pull request.

The test reactor executes a strict, sequential YAML or JSON test plan against the named devices in the authoritative config. Every step names the configured entry it drives with one key (device: <name>), and the configuration is what knows whether that name is a debuggers, a com_ports or a can_buses entry. The version 2 keys debugger:, port_id: and bus_id: stay valid as readable aliases; a step writes one or the other, never both, and a step that names its device twice is refused rather than resolved by precedence. Every device kind the config models can take part in a plan, and each kind answers for its own actions, permissions, session order and cleanup. The name may be omitted on a debugger step while the config declares exactly one probe; with several, the plan must name one, because picking a board for the author is how the wrong board gets flashed. Every probe carries its own permissions, so a step is judged by the grants of the board it names. Probes other than the bound one run on their own service under one shared project lease. Typed debug sessions run on OpenOCD, pyOCD and stlink, each through its own GDB server; on stlink that is ST-LINK_gdbserver, which STM32CubeCLT installs. An stlink entry without it still runs flash and UART plans, and so can a plan that only reads target memory, on a backend that serves those reads with no session.

A step action is a string a device class declares on the method that implements it, so a capability is one decorated method and a device kind is a class of them. A string a kind does not declare is answered not_supported naming the kind, by construction rather than by a special case. delay is declared once on the base class and is therefore served by every kind: it drives nothing, but it is still a step and appears in the run's result, because a wait that left no trace would be a gap in the record of what a test did.

Feedback steps say what they are claiming, not merely that they read something. uart_read with a comparator: reads until the claim is met or its timeout passes: equals: for a whole value rather than a substring (any one complete line of what was received matches as soon as its terminator arrives, and the whole of what was received matches on the last pass, once the timeout has run out and nothing further is coming, so a value the board never terminates cannot pass while it is still being printed), pattern: for a Python regular expression under re.search, and pattern: with range: for a number captured by the pattern's one capture group and held to inclusive bounds. A claim that goes unmet fails with the tail of what the port did say and, for a range, with the value it did capture, so a reading that was merely out of range and a board that said nothing read differently. Without a comparator the step is a plain read. The comparator is an object so that preprocessing a bench eventually needs can be added as further keys without invalidating plans written today.

can_read takes the same comparator over the medium a bus has. A frame arrives complete and with an identifier on it, so the claim is two things at once: id: (optionally widened by id_mask: into a family of identifiers) says which frame the plan is waiting for, and equals:, pattern: or pattern: with range: says what that frame must carry, read against the payload as hexadecimal. The identifier is required rather than optional, because a bus carries every node's traffic and a payload matched without saying whose frame it was is a green another ECU can produce; frames the filter rejects are passed over and nothing is captured from them. The step reads again until a frame meets the claim or its timeout_s passes, and an unmet claim fails with the last frames the bus did carry, so a plan waiting on the wrong identifier and a node that never sent anything read differently. Without a comparator it is the version 2 step unchanged: one read, answered exactly as can_read answers it.

read_symbol claims a value in target memory, which is the first thing this format asserts about the board's own state rather than about what it sent. It reads one allowed symbol through the debug session the plan opened, or straight off the probe on a backend that serves the read with no session, exactly as the debug_symbol_value tool reads it: the same debug.allowed_symbols, the same debug.max_dump_size_bytes on the resolved size, and the same resolution, which asks the image's debug information first and falls back to the ELF's own symbol table where there is no type to work with, so an assembly-defined object is readable too. Its comparator: is a numeric one, because a word of memory has no text to match: equals: for one value, range: for inclusive bounds, and mask: beside equals: for (value & mask) == equals, which asserts one flag of a status word without stating the rest of it. signed: true picks the signed reading, since the tool returns both and above the top bit the two disagree, and a mask is a claim about bits, so signed is refused beside it. A claim that goes unmet fails with the value it did read, the bytes it decoded and, for a mask, the masked value it compared. Without a comparator the step is a plain read whose value the report keeps.

That decoded integer exists only at 1, 2, 4 and 8 bytes, so a plan claiming a number at any other width is refused rather than answered with an invented one. A step may state size_bytes:, which is both an assertion (a firmware that turned a uint32_t into a uint16_t fails the step instead of having its new value read as the old one) and what makes the width knowable before the run: a declared width that no integer reading exists at, and one above this bench's debug.max_dump_size_bytes, are refused at preflight, where nothing has been opened. A plan that declares no width is refused at runtime instead, with the reason stated and the bytes it did read. There is no timeout_s on this step and no waiting: a halted target's memory does not change under it, so the symbol is read once and judged once.

Which debug steps need a debug_start before them depends on the probe the step names, and the answer comes from what that probe's backend serves rather than from what it is called. run_until_breakpoint and the session lifecycle need a typed debug session, which openocd and pyocd open through their own GDB servers and stlink opens through ST-LINK_gdbserver wherever the configuration names one or finds it beside STM32_Programmer_CLI. dump_memory and read_symbol need one only where the backend has no standalone read of its own: on the stlink backend STM32CubeProgrammer's memory read attaches and reads by itself, so a plan may carry either step with no debug_start at all and dump a RAM-resident measurement without a debug stack. On OpenOCD the same two steps run inside the session the plan opened, exactly as they always have, and a read outside one is still refused there, because opening one is what would fix it. A bench with neither route refuses the plan before the run, naming the backend, the step and its number, which reads that backend does serve, and the configuration change that would run it.

Close steps are optional: end-of-run cleanup closes every session the run opened, so a plan that never closes is complete. Write an explicit close where the plan means to close, reconfigure and reopen a line or a bus mid-run.

A plan states a cycle once and says how often it runs it. repeat is a block step the reactor serves itself: it routes to no device, the steps nested under its own steps: are the subset that repeats, and it is bounded by count: (iterations), by duration_s: (wall time), or by both, where the first bound reached ends the loop. At least one bound is required, so a plan cannot loop unbounded. Both bounds are asked between iterations and nowhere else, so a cycle is never cut in half and a block always runs at least one whole iteration however small its duration_s is; reaching a bound is the green exit. The nested list has the same shape as the plan's own, so a repeat may contain a repeat.

Nothing is opened, closed or reset between iterations: a port opened before the loop is still open inside it, a port opened inside the loop is opened once and closed by end-of-run cleanup like any other, and the loop adds no cleanup of its own. There is no retry and no until: a nested step that fails ends the run exactly as a failed top-level step does, with the same cleanup, the same report and the same recovery action after it. The report says where it stopped: failed_step is the top-level number of the block, step_error_type is the inner step's own error, and the block's record carries the step records of every iteration beside exit_reason, iterations_run, elapsed_s, and, on a failure, failed_iteration and failed_nested_step. Validation, the version gate, preflight and the run's lock declaration all descend into the nested steps, so a nested step is refused at the path that names its nesting (steps[2].steps[1].timeout_s) and a device a plan touches only inside a loop is a device the run locks.

version: 4
steps:
  - {device: dut_uart, action: uart_open}
  - action: repeat
    count: 1000
    steps:
      - {device: dut_can, action: can_send, frame_id: "0x100", data_hex: "01"}
      - {device: dut, action: delay, duration_ms: 500}
      - {device: dut_uart, action: uart_read, comparator: {equals: "cycle done"}, timeout_s: 2}
  - {device: dut, action: reset}

Before the first hardware action, the reactor validates every device name, permission, session order, artifact, breakpoint symbol, and dump path. A plan that contradicts the bus it declared is refused there as well: can_send on a bus configured listen_only: true can never work, because that flag is the claim that observing the bus sends nothing. uart_write on a port whose entry does not grant permissions.allow_write is refused by name in the same pass, as is a range: whose pattern captures nothing to put in it, and a read_symbol whose declared width the bench's debug.max_dump_size_bytes will not read or the format cannot decode a number from. Execution is fail-fast, each reactor-created breakpoint is removed after use, and debug, UART and CAN sessions opened by the runner are closed even when a step raises an exception. Breakpoint, dump and read symbols must be present in debug.allowed_symbols unless allow_all_symbols: true is explicitly set.

The run pipeline is deliberately simple. It validates everything, then executes, then always cleans up:

Test reactor pipeline: plan → preflight (no hardware touched; any finding rejects the whole plan) → sequential fail-fast execution → guaranteed cleanup → structured report Test reactor pipeline: plan → preflight (no hardware touched; any finding rejects the whole plan) → sequential fail-fast execution → guaranteed cleanup → structured report

# .agentic-hil/testconfig.yaml
version: 6
name: capture-state
steps:
  - {device: dut, action: flash, image_path: build/app.elf}
  - {device: dut_uart, action: uart_open}
  - {device: dut_can, action: can_open}
  - {device: dut_uart, action: uart_write, text: "capture\n"}
  - {device: dut_uart, action: uart_read, comparator: {pattern: "temp=(\\d+)C", range: {min: 20, max: 30}}, timeout_s: 5}
  - {device: dut, action: debug_start, image_path: build/app.elf, mode: attach}
  - {device: dut, action: run_until_breakpoint, location: capture_done, timeout_s: 5}
  - {device: dut_can, action: can_read, comparator: {id: "0x201", pattern: "^02(..)$", range: {min: 20, max: 30}}, timeout_s: 5}
  - {device: dut, action: read_symbol, symbol: capture_count, size_bytes: 4, comparator: {equals: 64}}
  - {device: dut, action: read_symbol, symbol: status_word, size_bytes: 4, comparator: {mask: 0x0f, equals: 5}}
  - {device: dut, action: dump_memory, symbol: capture_buffer, output_path: build/capture.hex}
  - {device: dut, action: debug_stop}

dut_can above is the listen_only: true bus from the configuration example, so the plan only reads it. A can_send step belongs on a bus whose entry sets listen_only: false and grants allow_write. The plan closes neither the port nor the bus, because it does not have to. Cleanup does.

A plan already written against version: 2, version: 3 or version: 4 keeps loading and behaving exactly as it did; the older uart_expect action stays valid too. A plan is held to what its own version: contains, so a version: 3 plan reaching for the version 4 block step, or a version: 4 plan reaching for the version 5 read_symbol action, is refused by name rather than working on one install and failing on an older one for no stated reason.

.agentic-hil/testconfig.yaml and --test-config select only this test plan: ordered test steps and the device names they run on. They contain no hardware resources or permissions. The reactor gets all hardware settings from the discovered authoritative config or its AGENTIC_HIL_CONFIG override:

agentic-hil test-reactor --test-config .agentic-hil/testconfig.yaml

See examples/testconfig.example.yaml for the expanded form.

An agent writing a plan reads the format over the connection it already has, not out of this page and never out of the installed package. agentic-hil://reference/test-plan is the whole of it: where the file goes, how test_config_path and workspace_root resolve its path, which version admits which step, every step with the entry it routes to and its required and optional keys, the comparator families with their rules, and two plans that run. agentic-hil://reference/test-plan-schema serves the schema itself, and the document is generated from that schema at read time, so neither can drift from what the reactor validates against.

The two ways to run a plan

An agent runs a plan with the test_reactor_run MCP tool; an operator runs it with agentic-hil test-reactor at a shell. Both reach the same code, so a plan behaves identically whichever one started it: the same preflight before the first hardware action, the same devices locked for the same duration, the same permission judged per step, the same report at the end. The tool takes test_config_path where the command takes --test-config, holds it to workspace_root the same way, and detach: true where the command takes --detach. test_reactor_status and test_reactor_stop answer as test-reactor-status and test-reactor-stop do.

The tool is the route for an agent because it is the only one an operator can see and audit: a plan run through a shell is a shell command, judged by whatever the agent host makes of it, while a plan run through the tool is coordinated and reported like every other hardware action here. A plan is also a run in its own right, so it needs no bench_run_start around it, and a call made while the bench is held by something else is refused with the holder named rather than queued; made while this session's own bench_run_start is open, it is refused as run_already_active, and bench_run_stop is the way out.

A JUnit report for CI

A run's report is this project's own JSON document, and no test-report reader understands it: a suite driven through the pytest fixtures inherits --junitxml for free, and the plan the examples recommend produced no such artifact at all. --junit-xml writes the second document those readers do understand, beside the report and without replacing it:

agentic-hil test-reactor --test-config .agentic-hil/testconfig.yaml --junit-xml build/junit.xml

Everything else about the command is untouched: the same JSON result on standard output, the same report files under reports.directory, the same exit code. Without the flag nothing new is written at all.

The document is one <testsuite> named after the plan and one <testcase> per plan step, named <index>.<route>.<action> in plan order and grouped under the plan name as their classname. A step that failed carries a <failure> whose type is its own error_type, whose message is its own summary and whose body is the step's whole record, comparator detail included; the steps after it are <skipped> with the reason, not passes, because they never ran. A plan refused at preflight has every case <skipped> and one preflight case carrying the refusal as an <error>, so a reader cannot mistake a refused plan for a failed bench, and a run refused before it reached a plan at all writes that preflight case on its own. A cleanup failure adds a cleanup case and a failed audit an audit case, because a run whose sessions did not close or whose record could not be written is not a passing run whatever its steps did. Timings are only ever the ones the run measured: a step whose backend reported no duration carries no time.

The file is written wherever the command answers at all, which includes every way a run can be red: a plan that does not load, a bench with no configuration, a plan refused at preflight, a step that failed on the board. Parent directories are created, so a path under a folder the upload step will make is enough. A file that cannot be written after a run is named in the JSON result under junit_xml_error and pulls the exit code down with it, rather than leaving a CI job to discover the artifact is missing. The one exception is a run that was being refused anyway: there the refusal is what the operator has to read, and it is not replaced by a story about a file path.

The one combination it refuses is --detach, and it says so: a detached start returns before the run has a report, so there is nothing to map yet, and no later command owns the path the start command was given. Run the plan synchronously to get the file, or follow the detached run with test-reactor-status and read the JSON report it names. The mapping itself is specified in the GitHub action design, which is what the planned action will pass a path to instead of holding a copy of.

The evidence bundle for CI

The JUnit file is one of four things the action design says a hardware job leaves behind. The other three come from one command, which reads a run report and writes them beside it:

agentic-hil run-evidence --report .agentic-hil/reports/last-report.json --out artifacts/evidence

It writes run-summary.json in the shape that document gives (the outcome, the plan with its path, name and sha256, the firmware the CI environment named, the tool versions, the bench by configuration digest and logical device name, and the run's failed step, error type, cleanup, audit and recovery), job-summary.md for a reviewer to read, and logs/, a copy of the JSONL event logs the report names: one com-*.jsonl per serial session, one can-*.jsonl per bus session, and the debugger backend's own per-invocation logs. On GitHub the job summary is appended to the file GITHUB_STEP_SUMMARY names as well, so it renders on the job page; it is always written under --out too, which is what makes the same command serve a GitLab job and an operator at a shell.

The exclusions are the command's, not the caller's. A run summary is world-readable on a public repository, so config_in_force.path, executable, probe_id, any absolute path from outside the workspace and every lock key are absent from both documents and from the names beside them: a device another run was holding is named by the logical name its plan gave it, never by the probe:<serial> the mutex took. Where one of those values sat inside a sentence a backend wrote, it is replaced by [withheld] and the sentence is kept, because the failing step's own words are what the artifact exists for.

A credential is withheld too, and by the code every other output goes through. The report is passed through the same redact_sensitive that print_json and the rendered result take, before anything at all is mapped off it, so a token, password or api_key field a tool result carried is already masked when the mapping reaches it. That function masks by key name and deliberately does not read prose, and prose is what these two documents publish, so every string on its way into either of them takes the content pass those captured streams take as well: a URL's userinfo, an Authorization: Bearer or Basic line and a token= style assignment become [redacted] where they stood. Only the credential span goes, so https://ci-bot:[redacted]@packages.example.com/simple/ still says which account reached which host, and the failing step's sentence is still the sentence the backend wrote. And the rule fails closed exactly as the other two sinks do: when redacting the report does not hand back a document, run-summary.json and job-summary.md are written as a redaction_unavailable refusal built from the command's own name, no log is collected, no field of the report appears anywhere in the bundle, and the command exits nonzero, because a bundle nothing could vouch for is a bundle that was not produced.

The logs are the one thing that is copied rather than derived, and they are copied byte for byte. A backend writes its own command line into its own log, so a device path can stand there; editing it out would leave a mirror that can no longer be checked against the hash-chained copy under state_root, which is the only reason the mirror is worth uploading. A secret a backend printed into its own log therefore stays in the copy, and it is the log's problem: fix it where the backend prints it, not in a mirror that would then verify against nothing. It is also the reason the logs are not world-readable by default. The two summaries are the half of the bundle written to be read by anyone, and logs/ is not: it goes into the job's artifact and never onto the job page, so a bundle from a private bench is uploaded on the terms the repository is published under, and a repository whose artifacts anybody may download is a decision to take deliberately rather than one this command takes for you. The rule is therefore the design's own: the summaries carry no identity and no credential, and the collected evidence is what the run recorded.

Nothing in it is invented. A field the environment did not supply is absent rather than empty: no firmware block outside CI, no bench.runner without RUNNER_NAME, no debugger version line unless the report carries that debugger's debugger_info result. The command loads no configuration and touches no hardware, so a report that records a preflight refusal, a missing configuration or a bench held by another run produces the same three outputs with outcome: refused, and a workflow's evidence step never has to special-case it. Its own exit code is about the evidence and not about the run, which is what makes if: always() safe to write.

It reads two things: the report, and the workspace it is running in. A report written by a run in another workspace is served rather than refused, and says so: the logs it names are not here, so they are listed under logs_missing, and any it names by an absolute path from elsewhere are counted under logs_outside_workspace without being published, because those paths are identities.

Detached runs, status and a cooperative stop

A time-bounded endurance plan runs for hours, and a caller that has to sit in front of it for that long is a caller that cannot do anything else. --detach runs the plan in its own process and returns at once with a run handle and the path the report will be written to:

agentic-hil test-reactor --detach
agentic-hil test-reactor-status --run run-3f9c2a1b4e6d8071
agentic-hil test-reactor-stop   --run run-3f9c2a1b4e6d8071

The start command answers as soon as the run holds the devices it declared, so a handle it printed is a handle with the bench. The plan is loaded before the worker is started, so a plan that does not load is refused immediately rather than behind a second command. The synchronous invocation is the default and does exactly what it always did; it is registered under a handle too, so a plan running in one terminal can be stopped from another. Over MCP the same three are test_reactor_run with detach: true, test_reactor_status with run, and test_reactor_stop with run; the handle is the same handle, so a run an agent detached is one an operator can ask about and end from a shell, and the other way round.

test-reactor-status answers for one handle: starting, while it takes the devices its plan declares; running, with the step and, inside a repeat, the iteration it is on; finished or stopped, with the verdict and the run's own report (canonical_report_path, the per-run copy under the state root that later runs do not overwrite; report_path is the workspace mirror the next run replaces); or worker gone. Without --run it lists the runs the bench still has records of. Whether the process behind a handle is still there is asked of the operating system rather than of a process id: a run holds a lock for as long as it runs, so a lock that can be taken is a run whose process has ended, however it ended.

test-reactor-stop is cooperative and nothing else. It writes a request; the run reads it between its steps and inside a delay, which waits in slices for exactly this reason and still waits its full duration when nobody asks it to stop. The run then finishes the step it is in, closes its devices in the usual cleanup order and writes its report. A run still waiting for a device another run holds (--wait-s on the start) reads the request inside that wait too, and ends there: it takes nothing, runs no step, and its report says it was stopped before any step. A stopped run is not a passed run: ok: false with error_type: run_stopped, the step it stopped after, and the records of everything that did run, so the caller decides what a partial run is worth. It is not a failed run either: its devices were closed and confirmed by the same cleanup a passing run uses, so there is no incident and no recovery action.

Killing the worker instead is the case this exists to replace. A process that dies without releasing anything is the dead-owner case the bench already handles: the status command names it (worker_gone) rather than guessing what the run had reached, a stop is refused because there is nobody left to honour it, and agentic-hil lease-status is what reads and heals the bench. Locks are machine-wide and unchanged by any of this: a detached run holds what it declared, and every other caller is refused with the owner named.

pytest Plugin

Installing agentic_hil registers the agentic_hil pytest plugin, so CI regression suites can drive the same permission-gated tools without an MCP client.

The agentic_hil fixture uses the same discovered config or absolute-path override as every other entry point and verifies that its workspace_root matches the pytest rootdir. Tests using the fixture skip when no config exists and fail loudly when an available config is invalid. Pytest executes project code and is therefore not a sandbox or security boundary; real unattended hardware runners must still use OS isolation and host-managed invocation. COM and CAN sessions opened during a test are stopped afterwards so stimulus state cannot leak between tests. See examples/nucleo-f446re_demo/ for the complete loop on real hardware.

--agentic-hil-config is deprecated, and it fails rather than degrades

The plugin accepts a --agentic-hil-config option and a matching agentic_hil_config ini key. Both are deprecated, and neither is the way to select a configuration. They are accepted only while they resolve to the configuration discovery has already found; a path that resolves anywhere else fails the session before the first test runs, with cannot change policy authority and both paths printed. It does not warn and carry on, and it does not fall back to the discovered file.

That is deliberate. These options live in the repository, in a command line or a pytest.ini a pull request can edit, and the authoritative configuration deliberately does not. An option that could point the suite at a different policy would let repository-controlled data decide what this bench may be told to do; one that silently fell back to the discovered file would run the suite under a policy nobody in that pull request asked for, and report it as a pass. Failing is the only remaining answer, so it is the one you get.

There are two supported ways to say which configuration a run uses, and neither is an option on this command line:

  • Say nothing. The authoritative configuration is discovered from the pytest rootdir, which is the project root, exactly as doctor, mcp-stdio and test-reactor discover it from their own project working directory. This is what a CI job wants: remove the option and the ini key, and the plugin finds the file the rest of the tooling finds.
  • Set AGENTIC_HIL_CONFIG to an absolute path when an operator-controlled override is wanted, in the runner's own environment rather than in a repository-controlled file. See Where it is found.

A suite that passes one of the deprecated selectors and resolves it to the discovered file still runs, unchanged. It is the only case in which they do anything at all, which is why removing them costs that suite nothing.