Explainer · 4 min read

Grex retires the wrapper around one task

A same-day comparison found no measurable gain from wrapping one bounded AI task in Grex's full execution machinery, and it surfaced a real secrets leak along the way.

Grex is a platform that gives AI agents bounded jobs on machines the operator owns. Some of those jobs are simple: read these inputs, produce these files, stop. Today the team asked whether that kind of single, bounded task needs all the machinery Grex normally wraps around it, ran a same-day comparison to find out, and decided it doesn’t.

The question

Inside a Grex mission, even a one-step task runs through several layers: a process that builds the AI agent’s environment, a job “pack” that decides how to call the model, and — when testing a pack against a benchmark — a mission simulator that stands in for the whole production pipeline. Each layer exists to support more complex, multi-step work. The question was whether a single bounded call gains anything from having them, or whether they are overhead with nothing underneath to justify it.

The comparison

The team ran five paired tasks. Each pair used the same AI model and the same frozen prompt once through the full wrapped path and once through a stripped-down direct path: one AI session, launched once, checked afterward by a step that runs outside anything the AI itself could touch. A separate AI model graded every answer against a scoring rubric, three times each, keeping the middle score.

This was a local scoring setup, not the official industry benchmark it borrowed tasks from, and it has known accuracy limits. Two of the five tasks were ones the team had already looked at while building the test harness. A sixth, stricter test used a task picked by a fixed rule before anyone read it — and on that one, both the wrapped and direct paths ran out of time and produced no usable final answer. That counts as a tie by exhaustion, not a tie by success, and it is reported as such.

On the five scored pairs, the two paths answered the same number of questions correctly overall — 17 of 41 graded items each — and no single pair differed by more than two items.

What the wrapping actually cost

With the outcome a wash, what the heavier path cost became the more useful finding.

  • A real secrets leak. Commands run inside the wrapped path could read credentials meant only for Grex’s own internal systems — administrative tokens and API keys — because nothing restricted what the AI’s own shell commands could see. That has been fixed with a default-deny list of what the AI session is allowed to read, and the fix has already shipped.
  • Mismatched environments. The two paths didn’t hand the AI model quite the same information about its surroundings, which is exactly the kind of inconsistency a fair comparison is supposed to rule out.
  • False failures. A labeling mismatch caused three of four wrapped-path attempts to be marked as failed even though the work they produced was fine.
  • A lot of code for one call. The direct path’s code came to about 3,400 lines. Reaching the same single model call through the wrapped path required roughly 53,000 lines of supporting code.

The decision

Going forward, a single bounded task inside a Grex mission runs as one solver call — one AI coding session, or one run of Grex’s own tool loop — with a receipt that independently records whether it ran, whether its output checks out, and whether it cleaned up afterward. That logic now lives in Grex’s own node software, not inside each individual job pack.

Nothing else about how Grex runs a mission changes: scheduling, spending limits, leasing work, and coordinating multiple AI agents on one job are untouched. The mission simulator also stays — it’s how Grex proves out test scenarios, including proving out this new direct path itself. Only the wrapper that existed solely to launch one call goes away.

The new receipt format, the launcher, and the pieces that prepare, run and verify one task call all merged into Grex’s codebase today, along with the secrets-leak fix. Grex’s node software advanced to a new version carrying all of it, and that version is now live on one of the two machines in the fleet; the other has not yet picked it up.

This post is AI-written. The comparison, the architecture decision and the code described here were the owner’s own direct engineering work; this account summarizes it afterward.

← Back to blog