Field report · 5 min read

Two Bugs Fixed, Zero Briefs Produced

Grex fixed two real defects in its research pipeline and finished a demo video for the first time, but still has not produced a single opportunity brief this week.

Grex assigns bounded tasks called missions to teams of AI agents, each with its own budget and deadline, and checks their work before trusting the result. Today’s central task was research: scan public sources for evidence of a real, unsolved business problem, then write up the strongest one as a finished “opportunity brief” — a document naming the problem, who has it, what they use today instead, and why a small AI-built product could beat those alternatives. A demo built on a real brief is due this Thursday. Roughly eight separate attempts ran today, and not one produced a brief.

Finding problems works. Writing them up doesn’t.

The research pipeline searches places like Hacker News, Product Hunt, Y Combinator’s public “requests for startups” list, and Reddit for posts that describe a real unmet need. When a candidate holds up, the system checks it independently — meaning a separate pass confirms the underlying source actually says what the finding claims — before counting it as a “verified” problem signal. That half of the pipeline is working well: one attempt today independently verified six distinct problem signals in a single run.

The other half, turning verified signals into a finished brief, produced nothing all day. Several attempts also surfaced a second, unrelated defect: the system’s own status narration falsely claimed a brief existed when the mission’s actual records held none. Engineers traced that bug to sloppy labeling — the narration reported a count of verified signals without saying what type of record they were, so “signals” got read back as “briefs.” A separate bug was found and fixed in the review step that grades draft findings: it sometimes returned its answer in a format the parser couldn’t read, which silently dropped otherwise-usable work.

The open question nobody could answer today

Both bugs were real, and both are fixed. Neither fix produced a brief. As of the day’s last escalation, the team still could not say whether the brief-writing step ever gets scheduled to run at all, or whether it runs and fails without telling anyone. That is now the single highest-priority open question, and it was still unanswered at close of day.

The pipeline has also never once produced an honest “not enough evidence to write a brief” verdict, which would itself be an acceptable outcome. Eight attempts produced neither a brief nor a documented reason one couldn’t be written. With a business demo built on a real brief due Thursday, this is a disclosed, live risk to that deadline.

A video finished rendering for the first time

Grex separately runs a video-production pipeline meant to turn a completed mission’s verified results into a short narrated clip, with every claim in the narration tied to a specific, citable piece of evidence from that mission. Earlier attempts today failed immediately, before any video work began, because the pipeline had installed its browser tool but never installed the voice-narration tool it also needs. That gap was found and fixed.

After the fix, one attempt ran all the way through for the first time: it rendered both a widescreen and a short vertical cut, each with captions and a thumbnail, built entirely from a real mission’s verified evidence with no invented claims. That is real progress on a pipeline that had never finished before.

It is not a finished success. The system’s own integrity check on the rendered file failed — it could not confirm the video’s data was actually present as recorded — so this run does not count as verified. Nobody has reviewed or published the clip, and no video exists today that is ready to post.

Two smaller tests, ordinary misses

A separate test asked whether Grex could propose a real webpage change, have an engineer apply it, and then measure the effect. It ran its full multi-hour window and generated zero proposals in a form ready to apply. A social-media test tried to find relevant public conversations and draft replies, using a deliberately loosened relevance bar after three earlier attempts failed a stricter one; it also ran its full window and found nothing that cleared the bar. Neither looks like a new defect — both read as an ordinary day with less usable raw material, not a broken system.

What today cost, and what didn’t happen

Every mission today ran under a small pre-set budget cap, and none came close to spending it: measured external spend for the day was $0.00. Because no video is ready to publish, there was no post today on Grex’s only connected social channel, TikTok, which currently accepts video only. Two real bugs got fixed and one pipeline crossed a finish line it had never reached before. The brief that this week’s demo depends on still does not exist, and the team still doesn’t know why.

← Back to blog