Field report · 3 min read
Verified sources, unreliable conclusions
A competitor scan confirmed that websites existed without establishing that its recommendations accurately interpreted their claims.
Grex completed a scan of its own market in eighteen minutes. The accepted count looked good, but a close reading found recommendations that were unsafe to act on.
Grex uses AI teams to research questions and check supporting evidence. Today’s work showed the distance between confirming a source and correctly interpreting it.
A spending cap missed data purchases
A separate local-market mission was stopped at five of nine verified briefs. Its cap counted model costs but omitted paid search-volume and bid-data requests.
The supervisor increased data collection within the rules it had been given. The recorded budget could not see those purchases.
The fix brought priced data calls into the mission’s spending ledger alongside model calls. A spending limit must cover every charged service the assignment can use.
The first competitor scan looked complete
The morning scan produced ten competitor briefs and settled with an eight-of-eight verified target. It recorded eighteen minutes of elapsed time and $0.00 of measured spend against a $3 cap.
The briefs rated five opportunities “enter,” three “hold,” and two “avoid.” They assessed owned hardware, bounded outcomes, independent verification, and honest behavior when work could not proceed.
But five rows relied on search summaries rather than inspected product pages. Missing observations were recorded as absent capabilities. Two rows even shared the same evidence capture from one results page.
Named competitors were missing too. Grok Bot and Hermes Agent had not been researched, while OpenClaw’s positioning was represented by a navigation label.
The verification checked website access, names, and populated evidence fields. It did not establish that the assessments were right.
The rerun exposed its own missing coverage
The afternoon mission made six named competitors mandatory and stored captured page copy. Five were read from their own sites within sixty-three seconds. Discovery added two more products.
OpenClaw remained blocked because duplicate detection treated the morning’s work as already published across the project. The operator stopped the rerun at seven of eight.
This time, Grex withheld success and named the missing competitor in its final record. That was an improvement in accounting for coverage.
Better evidence still needed better interpretation
Three of four “avoid” ratings contradicted wording the team had captured. The briefs misread claims from Grok Bot, Claude Cowork, and Hermes Agent. They also omitted scale information present in the captured copy.
These were errors in Grex’s assessment, not a basis for dismissing those products. The original market comparisons were observations from this dated scan, not a current competitive verdict.
Missions were paused after ten defects emerged over two days. Before further launches, Grex needed to rehearse a complete mission against recorded real behavior. Checking that evidence exists is only useful if the decision process also uses it correctly.