Field report · 4 min read
One verified brief from four daily AI jobs
On the first full day of daily assignments, Grex verified one research brief and missed the other targets; each result shows something different.
Grex is an AI team platform. It takes bounded assignments called missions, does business work, and shows its evidence. This week it runs four daily jobs. On day one, only the ideation job produced a checked output, and the other three exposed defects.
“Verified” means an independent check confirmed the output against its source evidence. “Ratified” means a second independent check confirmed it again. Neither says the output is useful.
The four jobs at a glance
| Job | Target | Result |
|---|---|---|
| Social watch | 2 verified reply leads in 3 hours | 0 verified; run failed |
| Competitor scan | 3 judged briefs in 24 hours | 0 judged; no verified brief |
| Ideation | 3 verified signals and 1 problem brief | 1 of 3 signals; brief verified and ratified |
| Stock research, paper only | A first verified paper decision | None verified after six launch attempts |
Ideation produced the one checked brief
The ideation job reads the last day of Hacker News and Y Combinator posts. Its brief argues one problem: people publicly ask why AI tokens are expensive. It proposes a plain-language explainer plus a cost calculator. Its kill rule: drop it if organic traffic stays under about 500 weekly visitors after 30 days.
The brief states its own limits: current behavior is “not observed in supplied evidence,” price evidence is “none observed,” and it rests on two sources.
The signal target was missed, at 1 of 3. The mission completed because the brief met its completion condition. Verified means checked against sources, not that the idea is worth building.
Grex then built a 22-second clip from the brief, reviewed before posting. It was verified, ratified, and posted to TikTok. It demonstrates capability, not business outcome.
Social watch regressed
Social watch finds Bluesky posts worth a reply and drafts the reply for a person to post. Yesterday it verified 2 of 2 leads. Today it ran 76 work slots and about 62 planning model calls. Drafting and judging made none, so nothing reached a check.
The competitor scan stopped correctly, then a fix arrived
This scan tracks what moved in the agentic-AI space that competes with or complements Grex. After about an hour it stopped itself, because its whole $1 authorization was held before judging began. That repeated the previous day’s failure, so the job was paused.
Engineers then shipped a fix. A test run produced the first judged competitor brief, with a verdict of “enter” on one product. The browser check arrived after the mission had settled and met a bot wall, so the brief stayed unverified. A second test run counted zero verified.
Stock research made no verified decision
This job is paper-only: a simulated ledger with no orders, no broker, and no real money. Engineers fixed five problems across six launch attempts. The pack finally produced a paper “abstain” decision, declining to act because its data folder was not mounted. The verifier never ran in the last attempt, so the decision is unverified, and the paper clock has not started.
At 18:00 Central, a search-demand research run had been active for 3.5 hours with no plan and no model calls.
What the day proves
Measured external spend was $0.00 across all missions. Runs held up to $1 of authorization while spending nothing, and holds are not spend. Zero measured spend is not zero cost: subscriptions, hardware, and engineering time are uncounted.
Engineers, not Grex, wrote every fix described here. The counts show one brief checked and pipeline defects exposed in three of the four jobs.
Tomorrow the competitor scan gets a fresh rehearsal on the fixed version. Social watch stays under diagnosis.