Stopping AI agents from skipping steps, with one file that runs anywhere

Stepgate is an MCP server that shows an agent one step at a time and moves on only when mechanical checks pass, so it cannot skip a step or fake a result. A stepfile is one portable file that runs in any MCP client. Measuring its checks showed most of them were doing work the model should never have been doing.

An AI agent given a procedure as text decides how much of it to follow. It drops a step, calls the task done, and leaves no record of what it ran. In τ²-bench’s single-control domains, a claimed success the environment does not show accounts for 45% to 48% of failures, and no LLM judge in that study scored above 0.65 AUROC at catching it (Advani, 2026). Another study found that 27% to 78% of the successes benchmarks report hide a procedural violation (Cao, Driouich and Thomas, 2026).

Stepgate is the MCP server I built to stop that, and the name is the design: a gate at every step. The procedure is a YAML stepfile. When it runs, Stepgate shows the client’s agent only the current step, makes every API call itself, and moves on only when that step’s gates pass. The agent cannot skip ahead, because it is not told the next step exists until the current one passes, and it cannot declare a step done, because a step passes only when its gates pass over what it submitted.

A gate is a mechanical check (JSON Schema, a JSONLogic rule, an HTTP verifier, or a person’s approval) and never a model’s opinion. Because Stepgate makes the calls, a gate can compare what the model submitted with what the API returned, which is how a fabricated value gets caught. That makes the path through a procedure deterministic, within limits worth being exact about. Given the same inputs, the same submissions and the same API responses, a run takes the same path: the steps run in the same order, and every gate gives the same verdict. What the model submits still varies, and so does anything a person approves.

How a Stepgate run works Sequence diagram: the client's agent starts a run and is shown only step 1, calls an operation through Stepgate, which adds the credential and calls the API, then submits the step's output; Stepgate's gates either return a diagnosis for a retry or show step 2. ALT [a gate fails] [every gate passes] START RUN (INPUTS) STEP 1 ONLY STEPGATE_CALL REQUEST + KEY FULL RESPONSE RESULT STEPGATE_SUBMIT RUN GATES DIAGNOSIS, RETRY STEP 2 Agent IN YOUR MCP CLIENT Stepgate HOLDS KEYS, RUNS GATES API DECLARED HOSTS ONLY CALL REPLY
The agent never sees step 2 until step 1's gates pass, and never holds the API key.

The second goal was portability. I wanted an agent I could hand to anyone: one file, with no packages, runtime or versions to install, that runs in whatever AI client the other person already uses, with the model they already pay for. MCP, the Model Context Protocol that clients such as Claude Code and Cursor use to call tools, makes that possible, because a server that speaks it plugs into all of them. So the stepfile is the whole agent. It names no model, provider or framework, calls only remote APIs, and declares the credentials it needs without saying where they live, so whoever runs it supplies their own keys.

That is also what separates a stepfile from a workflow engine such as n8n. An n8n workflow runs inside an n8n server, with the model and keys configured there. A stepfile travels to the agent: anyone runs it with npx -y stepgate in the client they already use, and its gates come with it.

Two more properties come from Stepgate making every call. The model never holds an API key, because Stepgate attaches credentials itself and sends requests only to the hosts the stepfile declares. And every run writes a hash-chained ledger of each step, tool call, gate verdict and retry, so editing a record afterwards breaks the chain, which stepgate --verify detects.

Half of every stepfile was checking

Stepgate ships with a catalog of eighteen stepfiles: CVE triage, a 10-K ratio extraction, a regulatory monitor, a Zendesk-to-Jira escalation and so on. I wrote a script to count where their lines go, and the answer was uncomfortable:

#### THE CATALOG, AS FIRST WRITTEN ####
Lines, counted as block-style YAML: 16247
Share of lines: {"gates":"53%","produces":"15%","tools":"14%","instructions":"8%","other":"11%"}
Filters over calls by tool: 145
Predicates reading: {"calls":106,"steps_or_inputs":110,"output_only":45}

Gates were 53% of the catalog and instructions were 8%. The same filter (“the successful results of this operation”) was written out by hand 145 times, and 110 of the 261 predicate gates compared the output with the inputs or an earlier step. Reading through them, most were not checking judgement. They checked that the model had copied a value correctly, counted a list correctly, or done a sum correctly.

The clearest case checks the vehicle on an insurance claim against its VIN. It decodes the VIN with NHTSA’s vPIC service and compares the claimed make and model with the decode, ignoring letter case. The model did the comparison, and a gate redid it by lowercasing both strings with a 26-branch if, one branch per letter of the alphabet. Four agent steps and 372 lines, mostly proving that the model could copy fields out of JSON and compare two strings.

That is backwards. A model is slow at copying and unreliable at arithmetic, and each of those steps cost tokens, sometimes a retry, and a gate few people could review. None of it needed judgement.

The model judges, the server computes

A stepfile has two kinds of step. An agent step has instructions, and the model does it. A mechanical step has do instead: Stepgate makes the listed calls, with arguments built from the inputs and earlier outputs, and computes the step’s output from a template. The client never sees a mechanical step. The VIN decode is one:

- id: decode
  do:
    calls:
      - { id: decode, operation: decodeVin, arguments: { vin: { var: inputs.vin } } }
    output:
      make: { var: responses.decode.Results.0.Make }
      model: { var: responses.decode.Results.0.Model }
      model_year: { var: responses.decode.Results.0.ModelYear }

When only part of an agent step is a computation, such as a count, a derived field covers it: Stepgate fills the field in after the model submits, so the model is never asked for it and no gate has to check it.

The VIN stepfile is now four mechanical steps (decode, compare, recalls, verdict) and one agent step, the summary for the claims handler, in 242 lines. Against the live NHTSA APIs, the VIN decodes to a 2009 Toyota Prius with six recalls, the claim says a 2010 Camry, and the verdict is refer, with no field touched by the model. Its one job is the prose, and the gates on that prose check it against those values. A summary citing a recall campaign that was never returned is rejected by name.

The VIN stepfile before and after Before, the model did all four steps and 15 gates checked its copying and comparisons, in 372 lines; after, Stepgate does four steps itself and the model writes only the report, checked by 2 gates, in 242 lines. BEFORE · 372 LINES · 15 GATES decodeAGENT · 2 GATEScompareAGENT · 4 GATESrecallsAGENT · 4 GATESreportAGENT · 5 GATES AFTER · 242 LINES · 3 GATES decodeSTEPGATE · 1 GATEcompareSTEPGATErecallsSTEPGATEverdictSTEPGATEreportAGENT · 2 GATES AGENT STEP: THE MODEL WRITES THE OUTPUT MECHANICAL STEP: STEPGATE DOES IT
Twelve of the fifteen gates checked work the model no longer does.

The whole catalog works this way now:

#### THE CATALOG NOW ####
Lines, counted as block-style YAML: 12397
Share of lines: {"gates":"12%","mechanical":"30%","produces":"19%","tools":"18%","instructions":"4%","other":"17%"}
Filters over calls by tool: 2

This did more for determinism than any gate. A mechanical step returns the same output for the same API responses every time, so each step moved out of the model is one less place for a run to vary.

The files went from 8,547 lines to 6,321, and gates from 53% to 12%. I trust that second number less than it looks, because much of the gate logic did not disappear. It moved into the do templates, which are JSONLogic too, and gates and templates together are still 42% of the catalog. The model does far less. The stepfile is only somewhat easier to read.

Writes wait for a person

Six of the stepfiles write something: a Jira issue, a Gmail draft, a monday.com item, a note on a Zendesk ticket or a ServiceNow incident. Each follows the same shape. An agent step drafts exactly what will be written, a person approves it, and a mechanical step performs the write from the approved output. The write step can read only inputs and earlier outputs, never a fresh API response, so what reaches Jira is what the person saw.

The approval travels as an MCP elicitation, a form the client shows. Not every client shows it. Testing from the Claude Code extension in VS Code, every approval came back declined within a millisecond of the submission, which is not a human reaction time: the extension declares the capability and then declines without drawing the form (#79174). The same test in a terminal session came back approved 5.8 seconds after the submit. Stepgate cannot tell those two declines apart, so it reports only that the approval was declined, and the docs name the clients that work.

Where this stops being true

Running a stepfile still needs Node.js 22.18 or later wherever Stepgate itself runs, so “no dependencies” is true of the file and not of the machine. Every number here comes from one catalog of eighteen stepfiles that I wrote, so the 53% says as much about how I wrote gates as about gates in general. The eleven public-data stepfiles pass against their live APIs with qwen3.6-35b-a3b through OpenRouter. The six that write were checked only against recorded cases, because I do not have accounts on those services.

The bigger limit is JSONLogic. Inside a map or filter, an expression sees only the current item, so anything that needs outer data turns into a reduce carrying context or a join against a one-element list. Sorting and date arithmetic are both awkward today. Those gaps are open as proposals in the issue tracker, with the largest being CEL expressions as strings beside JSONLogic. CEL’s specification guarantees that an expression terminates, which matters when the file is untrusted.

Try it

Stepgate is on npm and the MCP Registry, and any MCP client can run it:

npx -y stepgate --list                 # the catalog
npx -y stepgate auto-claim-vin-validation

If you already have a procedure written down as a skill or a runbook, stepgate_outline turns it into a skeleton stepfile, with each rule it states (“never”, “must”) listed as a gate still to write. The repository has the format reference, the catalog, and the script behind both measurement blocks.

Moving work out of the model also made the agent more portable. A step Stepgate computes behaves the same whichever model the client runs, so the fewer steps a model does, the less a stepfile’s behaviour depends on where it lands. The path through a stepfile and every step that needs no judgement are deterministic, as long as the APIs return the same data. Only the judgement travels with the model.

The 53% was my own doing, but I doubt the pattern is mine alone. Ask a model to do mechanical work and you end up writing checks to catch it, and each of those checks needs a retry budget and a reviewer who can read it. If I were reviewing a team’s agent design, the first question I would ask is which values the model produces that code could compute. Every one of those is cheaper to move into code than to guard. Most of mine should never have been the model’s job.

no account needed

Comments

    Comments need JavaScript.

    © 2026 Charangan Vasantharajan