Field report · 2 min read

The first comparison produced no usable briefs

A controlled agent experiment found drafting failures and a reviewer whose rejection did not change the accepted-result count.

What does Grex add beyond the AI model inside it? A new experiment held the model and task constant to investigate that question.

One arm used a plain agent loop on a laptop. The other used the same budget model inside Grex, with mission limits, a plan, evidence requirements, and a separate reviewer.

The objective and scoring rules were frozen before scored work began. Today’s results were calibration findings from the Grex arm, not a completed comparison between the two systems.

Eight candidates reached drafting; none became output

The Grex arm searched for social posts worth responding to. It found 53 candidates, of which eight passed the relevance filter.

All eight failed the drafting policy because the proposed responses were promotional. The mission published zero briefs and closed itself as failed.

Refusing unsuitable replies preserved the quality threshold. It still left the requested work undelivered.

The process had problems before model quality mattered

Calibration found five operational issues:

  • Search access had not followed a worker’s new identity after its rebuild.
  • The search path remained pinned to an exhausted free tier.
  • A scoring bug assigned zero relevance to one input class.
  • Quoted multiword searches returned too few candidates from the source.
  • Rejected candidates were offered for drafting again on later cycles.

The last problem wasted repeated attempts on decisions already made. The correction was to remember refusals, rather than lower the editorial bar.

The reviewer’s verdict did not control the count

A separate review run later examined the one brief available to it. The source resolved and the evidence references were structurally valid, but the reviewer rejected the brief as promotional and insubstantial.

The mission evaluation still counted that brief as verified. The reviewer’s verdict was advisory at this stage and could not retract the accepted count.

That was a material limitation. A separate reviewer helps only if its decision reaches the system that declares success. The gap was recorded for repair.

← Back to blog