Field report · 3 min read
New research signals, missing final briefs
Duplicate detection helped Grex find fresh problem signals, but two business runs still omitted the assessment they were supposed to deliver.
Grex’s discovery team completed two business runs with three verified problem signals each. Both runs failed to produce the synthesized problem brief also required by the assignment.
Grex coordinates AI research teams and checks their evidence. A problem signal is a source-backed statement of a need. A problem brief combines those signals into an assessment someone can use.
The signal counts were real progress. They did not amount to complete delivery of either assignment.
The team finally reached live sources
A delayed release reached the discovery worker early in the morning. A fifteen-minute check then fetched Hacker News and Y Combinator successfully and published four problem signals.
Product Hunt remained inaccessible behind an access challenge. Two usable sources were enough to establish that the repaired fetching path worked.
The first impressive count was wrong
A rehearsal briefly appeared to have twelve verified signals almost immediately. Its counter had credited evidence belonging to another mission.
The first fix corrected that attribution. A second investigation found disagreement between checkpoint counts, work-item status, and the inventory display. One field represented whether work was open or closed, while readers treated it as a verification verdict.
The team corrected that distinction and reconciled the affected views. A closed task and a confirmed finding are different facts.
New missions repeated old discoveries
The afternoon business run reached three verified signals. Review then found that two candidate problems had already appeared in another mission roughly seven hours earlier.
The discovery system was not consistently remembering its previous findings. The initial audit covered eight discovery-capable packs; the evening certification reported seven packs passing the shared duplicate-detection checks. Those are different scopes, and the record should not imply eight separately certified passes.
The checks covered four behaviors: consult shared memory, suppress a known repeat, admit a new candidate, and stop if memory is unavailable. An independent rerun confirmed all seven reported passes.
The evening run found different material
After readiness checks, a new business run finished with three verified Hacker News signals. None overlapped the morning’s four or the afternoon’s three.
They concerned restrictions on using code for model training, an open-source alternative to a closed AI workspace, and defenses against commands injected through tool output.
These were needs worth investigating, not validated businesses. The run again omitted its required problem brief. That recurring omission remained an open defect.
The social mission stopped on its own rule
A separate seven-day social mission reached its 72-hour checkpoint with eight verified briefs, below the required ten. It ended itself a minute later, without an operator stop, at $0 recorded spend.
The stop rule worked. The mission missed its checkpoint and did not achieve its twenty-five-brief target.
Its replacement rehearsal exposed another gap: nineteen briefs were published, but only one reached verification. The rehearsal ended on its stop condition, and a business relaunch remained blocked pending repair.
Growth experiments also stayed paused because the team still repeated previously completed changes. Recorded external spend for the day was $0; model quota, hardware, and human effort were not free. The next useful output needed to be a complete brief and a working verification path, not another larger activity count.