Field report · 3 min read

The eighth mission finally produced usable drafts

A repaired mission allowance let Grex deliver four verified reply briefs after seven trials had returned no verified results.

Seven Grex missions ran without producing an independently verified result. The eighth delivered four verified reply briefs against a target of two.

Grex is a system for assigning AI work with limits and checking what comes back. Today’s short trials tested where that process broke before committing to longer business work.

A rehearsal’s result leaked into a new mission

The overnight social-research mission reported “target met” at zero of twenty-five. It had inherited a daily allowance record from an earlier rehearsal.

The record was shared across missions when it needed to belong to one assignment. A bounded trial reproduced the failure on demand.

After the correction, a three-hour rehearsal completed with four verified reply briefs and $0 external spend. That was the first live evidence that the repaired allowance let the new mission work.

Four market briefs remained unverified

Local-market trials examined roofing, dumpster rental, HVAC repair, and plumbing. None produced a verified result.

The dumpster-rental task crashed when it tried to publish a price signal its team had not declared as an allowed output. Other briefs recommended “hold” but could not be independently confirmed.

In the final trial, the supervising agent explicitly reported that its plan kept producing unverified briefs. The operator stopped it on that evidence.

Growth research could not see prior work

Another mission tried to measure changes already deployed on business websites. Its planning process could not attribute changes made outside its own execution history.

It repeatedly produced empty measurement instructions and ended with a failed checking stage. No drafting calls were made, and no spend was recorded.

Several recent defects shared this memory problem. Agents repeated deployed changes, invented asset facts, or confused one mission’s history with another’s.

Shared memory was introduced

The team introduced Doctrina, Grex’s curated shared memory. Its initial records held completed work and facts about website offers, prices, and pages. Agents received read access; they could not freely rewrite those facts.

The initial store was seeded with thirty facts across five sites and ten completed-work records. Its first populated read failed, then a repair restored it. Memory-aware research teams began rolling out, while a growth-team release remained blocked by a separate publication regression.

Creating the store did not yet prove that planners would use it correctly. That needed a behavioral trial.

A longer social mission began

The successful evening rehearsal allowed a seven-day mission to launch with a target of twenty-five briefs and a $2 cap. Its first hour had no plan because activation missed the scheduled planning cycle. The next cycle recovered it.

The rehearsal also exposed limited source diversity: its first three briefs came from one author’s feed. A later business draft exceeded Bluesky’s reply limit. Both issues needed follow-up.

The day’s eight trials produced four verified results, all from the final rehearsal. External spend totaled $1.55 against a $3 cap, with no revenue. The longer mission’s first verified brief arrived overnight and belonged to that separate assignment.

← Back to blog