An AI agent given a procedure as text decides how much of it to follow. It drops a step, calls the task done, and leaves no record of what it ran. In τ²-bench’s single-control domains, a claimed success the environment does not show accounts for 45% to 48% of failures, and no LLM judge in that study scored above 0.65 AUROC at catching it (Advani, 2026). Another study found that 27% to 78% of the successes benchmarks report hide a procedural violation (Cao, Driouich and Thomas, 2026).
Stepgate is the MCP server I built to stop that, and the name is the design: a gate at every step. The procedure is a YAML stepfile. When it runs, Stepgate shows the client’s agent only the current step, makes every API call itself, and moves on only when that step’s gates pass. The agent cannot skip ahead, because it is not told the next step exists until the current one passes, and it cannot declare a step done, because a step passes only when its gates pass over what it submitted.
A gate is a mechanical check (JSON Schema, a JSONLogic rule, an HTTP verifier, or a person’s approval) and never a model’s opinion. Because Stepgate makes the calls, a gate can compare what the model submitted with what the API returned, which is how a fabricated value gets caught. That makes the path through a procedure deterministic, within limits worth being exact about. Given the same inputs, the same submissions and the same API responses, a run takes the same path: the steps run in the same order, and every gate gives the same verdict. What the model submits still varies, and so does anything a person approves.
The second goal was portability. I wanted an agent I could hand to anyone: one file, with no packages, runtime or versions to install, that runs in whatever AI client the other person already uses, with the model they already pay for. MCP, the Model Context Protocol that clients such as Claude Code and Cursor use to call tools, makes that possible, because a server that speaks it plugs into all of them. So the stepfile is the whole agent. It names no model, provider or framework, calls only remote APIs, and declares the credentials it needs without saying where they live, so whoever runs it supplies their own keys.
That is also what separates a stepfile from a workflow engine such as n8n. An
n8n workflow runs inside an n8n server, with the model and keys configured
there. A stepfile travels to the agent: anyone runs it with npx -y stepgate in
the client they already use, and its gates come with it.
Two more properties come from Stepgate making every call. The model never holds
an API key, because Stepgate attaches credentials itself and sends requests only
to the hosts the stepfile declares. And every run writes a hash-chained ledger
of each step, tool call, gate verdict and retry, so editing a record afterwards
breaks the chain, which stepgate --verify detects.
Half of every stepfile was checking
Stepgate ships with a catalog of eighteen stepfiles: CVE triage, a 10-K ratio extraction, a regulatory monitor, a Zendesk-to-Jira escalation and so on. I wrote a script to count where their lines go, and the answer was uncomfortable:
#### THE CATALOG, AS FIRST WRITTEN ####
Lines, counted as block-style YAML: 16247
Share of lines: {"gates":"53%","produces":"15%","tools":"14%","instructions":"8%","other":"11%"}
Filters over calls by tool: 145
Predicates reading: {"calls":106,"steps_or_inputs":110,"output_only":45}
Gates were 53% of the catalog and instructions were 8%. The same filter (“the successful results of this operation”) was written out by hand 145 times, and 110 of the 261 predicate gates compared the output with the inputs or an earlier step. Reading through them, most were not checking judgement. They checked that the model had copied a value correctly, counted a list correctly, or done a sum correctly.
The clearest case checks the vehicle on an insurance claim against its VIN. It
decodes the VIN with NHTSA’s vPIC service and compares the claimed make and
model with the decode, ignoring letter case. The model did the comparison, and
a gate redid it by lowercasing both strings with a 26-branch if, one branch
per letter of the alphabet. Four agent steps and 372 lines, mostly proving that
the model could copy fields out of JSON and compare two strings.
That is backwards. A model is slow at copying and unreliable at arithmetic, and each of those steps cost tokens, sometimes a retry, and a gate few people could review. None of it needed judgement.
The model judges, the server computes
A stepfile has two kinds of step. An agent step has instructions, and the model
does it. A mechanical step has do instead: Stepgate makes the listed calls,
with arguments built from the inputs and earlier outputs, and computes the
step’s output from a template. The client never sees a mechanical step. The VIN
decode is one:
- id: decode
do:
calls:
- { id: decode, operation: decodeVin, arguments: { vin: { var: inputs.vin } } }
output:
make: { var: responses.decode.Results.0.Make }
model: { var: responses.decode.Results.0.Model }
model_year: { var: responses.decode.Results.0.ModelYear }
When only part of an agent step is a computation, such as a count, a derived field covers it: Stepgate fills the field in after the model submits, so the model is never asked for it and no gate has to check it.
The VIN stepfile is now four mechanical steps (decode, compare, recalls,
verdict) and one agent step, the summary for the claims handler, in 242 lines.
Against the live NHTSA APIs, the VIN decodes to a 2009 Toyota Prius with six
recalls, the claim says a 2010 Camry, and the verdict is refer, with no field
touched by the model. Its one job is the prose, and the gates on that prose
check it against those values. A summary citing a recall campaign that was never
returned is rejected by name.
The whole catalog works this way now:
#### THE CATALOG NOW ####
Lines, counted as block-style YAML: 12397
Share of lines: {"gates":"12%","mechanical":"30%","produces":"19%","tools":"18%","instructions":"4%","other":"17%"}
Filters over calls by tool: 2
This did more for determinism than any gate. A mechanical step returns the same output for the same API responses every time, so each step moved out of the model is one less place for a run to vary.
The files went from 8,547 lines to 6,321, and gates from 53% to 12%. I trust
that second number less than it looks, because much of the gate logic did not
disappear. It moved into the do templates, which are JSONLogic too, and gates
and templates together are still 42% of the catalog. The model does far less.
The stepfile is only somewhat easier to read.
Writes wait for a person
Six of the stepfiles write something: a Jira issue, a Gmail draft, a monday.com item, a note on a Zendesk ticket or a ServiceNow incident. Each follows the same shape. An agent step drafts exactly what will be written, a person approves it, and a mechanical step performs the write from the approved output. The write step can read only inputs and earlier outputs, never a fresh API response, so what reaches Jira is what the person saw.
The approval travels as an MCP elicitation, a form the client shows. Not every client shows it. Testing from the Claude Code extension in VS Code, every approval came back declined within a millisecond of the submission, which is not a human reaction time: the extension declares the capability and then declines without drawing the form (#79174). The same test in a terminal session came back approved 5.8 seconds after the submit. Stepgate cannot tell those two declines apart, so it reports only that the approval was declined, and the docs name the clients that work.
Where this stops being true
Running a stepfile still needs Node.js 22.18 or later wherever Stepgate itself runs, so “no dependencies” is true of the file and not of the machine. Every number here comes from one catalog of eighteen stepfiles that I wrote, so the 53% says as much about how I wrote gates as about gates in general. The eleven public-data stepfiles pass against their live APIs with qwen3.6-35b-a3b through OpenRouter. The six that write were checked only against recorded cases, because I do not have accounts on those services.
The bigger limit is JSONLogic. Inside a map or filter, an expression sees
only the current item, so anything that needs outer data turns into a reduce
carrying context or a join against a one-element list. Sorting and date
arithmetic are both awkward today. Those gaps are open as proposals in the
issue tracker, with the largest
being CEL expressions as strings beside JSONLogic. CEL’s specification guarantees
that an expression terminates, which matters when the file is untrusted.
Try it
Stepgate is on npm and the MCP Registry, and any MCP client can run it:
npx -y stepgate --list # the catalog
npx -y stepgate auto-claim-vin-validation
If you already have a procedure written down as a skill or a runbook,
stepgate_outline turns it into a skeleton stepfile, with each rule it states
(“never”, “must”) listed as a gate still to write. The
repository has the format reference,
the catalog, and the script behind both measurement blocks.
Moving work out of the model also made the agent more portable. A step Stepgate computes behaves the same whichever model the client runs, so the fewer steps a model does, the less a stepfile’s behaviour depends on where it lands. The path through a stepfile and every step that needs no judgement are deterministic, as long as the APIs return the same data. Only the judgement travels with the model.
The 53% was my own doing, but I doubt the pattern is mine alone. Ask a model to do mechanical work and you end up writing checks to catch it, and each of those checks needs a retry budget and a reviewer who can read it. If I were reviewing a team’s agent design, the first question I would ask is which values the model produces that code could compute. Every one of those is cheaper to move into code than to guard. Most of mine should never have been the model’s job.
Comments
No comments yet. Yours would be the first.
Comments need JavaScript.