Field report · 3 min read
The discovery agent could not read its sources
Repeated live trials exposed five failures between planning a search and actually reading the intended sources.
Grex’s new discovery team produced no useful discovery output before the day’s pause. Each trial reached a little further through the process and found another failure.
The team, called problem-radar, was meant to read startup and developer sources for problems worth solving. Grex wrapped that work in limited rehearsal missions before allowing a business launch.
The plan reached the worker in stages
The first run submitted a plan without a required field. The second got past that check but discarded the execution record needed by the central service.
The third delivered a plan to the search agent, which scanned zero items. Comparing the adapters with real pages showed why: Y Combinator and Hacker News no longer matched the expected page structures. Product Hunt presented an access challenge.
Repairs restored the first two source paths. Product Hunt remained inaccessible through the tested route.
A live check found two more integration failures
A new short live trial caught inconsistent source names. The planner and search agent used different spellings for the same services.
A shared vocabulary fixed that mismatch. After an authorized evening restart, the corrected release installed and a fifteen-minute trial confirmed that plans were valid and delivered.
It then found a fifth defect: the pack had not declared the browser dependency needed to fetch pages. All sources were reported unavailable, with too little diagnostic detail.
The next repair added the dependency through an established fetch path and recorded why requests failed. It passed certification and merged that night. A coding-provider usage limit interrupted release before publication and installation, leaving the discovery worker waiting.
Growth memory blocked bad proposals
Doctrina is Grex’s curated memory of asset facts and completed work. An evening growth trial proved two useful behaviors against it.
A proposed $0 price was refused because the stored offer was $10 for twenty-five uploads. A proposal to repeat a live change was also suppressed.
But the overall assignment still failed. One run encountered an unsupported budget term. Its replacement kept re-proposing a completed retitle instead of drafting the requested FAQ. The supervisor replanned for eight cycles without resolving the problem.
Both runs were stopped. Memory prevented some bad output, but the team had not shown that it could produce the requested new work.
Social research remained active
The seven-day social mission reached eight verified briefs against a target of twenty-five. Its next checkpoint required ten within 72 hours.
A corrected social pack installed while the mission continued. It remembered declined drafts and reduced repeated scanning of exhausted feeds.
The cost was larger than the cash meter
Recorded external spend was $0. Engineering attempts, supervision, and provider quota were still consumed.
The release stall made that distinction concrete. A certified repair could not help the live worker while the tool needed to release it was unavailable. Repeated overnight status checks did not shorten that wait.
The remaining work was clear: install and prove the discovery repair, resolve growth’s repeated proposals, and assess the social mission at its checkpoint.