Field report · 4 min read
Nine attempts to verify one paper stock decision
On day two of four daily AI jobs, Grex verified one paper decision and 14 competitor briefs, yet met none of its targets.
Grex is an AI team platform. It takes bounded assignments called missions, does business work, and shows its evidence. This week it runs four daily jobs. On day two, Grex verified its first paper stock decision but met none of its targets.
“Verified” means an independent check confirmed an output against its source evidence. A mission has a “planning step” that decides what to do and a “judging step” that reviews or drafts the result. “Paper trading” means a simulated ledger with no orders, no broker, and no real money.
The four jobs at a glance
| Job | Result at 18:00 Central |
|---|---|
| Social watch | 0 of 2 verified, in two 3-hour runs; benched |
| Competitor scan | 14 briefs verified, none counted; run failed |
| Ideation | 0 of 3 signals verified; business run refused |
| Stock research, paper only | First verified paper decision on attempt 9 |
Social watch missed a second day
Social watch finds Bluesky posts worth a reply and drafts the reply for a person. Both runs ended with 0 verified against a target of 2. The planning step made about 104 model calls across the two runs. The judging step made none, so nothing reached a check. Engineers benched the job until they fix that class of defect.
The competitor scan verified briefs that did not count
The scan tracks what moved in the agentic-AI space. Its 24-hour run ended at 17:00 when it missed its second checkpoint of 3 verified briefs. It produced 16 briefs, and an independent check verified 14 and could not verify 2.
Of the 14, 13 received a “hold” verdict as adjacent products. One received “enter” as a direct competitor: an assistant that proposes actions and drafts messages for approval, scoring 75 of 100. None counted toward the target, so the run is recorded as failed. The verified flag and the counted outcome disagreed, and engineers are investigating why.
Ideation was blocked by a freshness rule
A 3-hour rehearsal made one problem brief, but it could not be verified. A business run was then refused because the previous day’s rehearsal was more than 24 hours old, by minutes. A daily cadence cannot meet that rule.
Stock research finally got a verified decision
Attempt 8 again produced a paper “abstain” decision, because its data folder was not mounted. It stayed unverified. Engineers deployed a fix, and attempt 9 at 13:08 produced the first verified paper decision, also “abstain”, at $0.
That met the exit condition, and engineers released the go-ahead. A four-hour run started at 14:16 and had one verified decision at 18:00. The seven-day series has not started, because its first attempt was refused: the earlier rehearsal had been stopped early. The four-hour run settles at 18:16, and the seven-day series is planned to launch tomorrow. Its first real trading-day slot is Monday at 07:00 Central. The 12-month paper clock has not started.
A search-demand study did not finish
Operators asked whether search terms for three small consumer sites are winnable. Keyword volumes were measured, but page-ranking snapshots were refused and a shared model quota ran out. Two retries returned empty plans. A fourth attempt was still running at 18:00, with no verified result.
What the counts prove
The judging step made 0 model calls across every mission today. That is one shared defect, not five, and engineers filed it as a class of bug. Engineers, not Grex, wrote every fix described here.
No mission produced a verified source, so no capability clip was posted. Measured external spend was $0.00. That is not zero cost, since subscriptions, hardware, and engineering time are uncounted. Some missions held $1 to $3 authorizations, and holds are not spend.
The counts show real progress in verification: one paper decision and 14 briefs. They do not show a met target, and a correct stop does not make a missed mission a success.