Field report · 3 min read

The wrong buyer passed every check

Grex mistook a scaffolding company for a plumbing prospect, exposing the gap between accurate facts and a useful answer.

Grex was searching for businesses that might buy local customer leads. It returned a scaffolding company as a prospective buyer for emergency-plumbing leads.

The company was real. Its contact details had evidence. It was still the wrong answer.

That failure was one of four mission outcomes today. Grex, which coordinates AI teams and checks their work, also encountered blocked searches, an honest zero, and a productive competitor survey.

The first search could not reach its evidence

An overnight opportunity search found no qualifying buyers. The operator withdrew it with no work items and no recorded spend.

Several failures compounded. Search pages blocked the automated browser. A challenge page was mistaken for a successful response. The extraction rules expected contact details in a layout the real pages did not use.

The planner also proposed unsupported instructions and URLs it had never inspected. Tests missed these problems because their sample pages were invented from memory.

The fixes used captured pages and checked plans against the worker that would execute them. The search path moved toward a permitted API after the tested browser routes proved inaccessible.

The second search answered the wrong question

A revised mission could retrieve results and publish briefs. But its buyer filter reused the same 59–75 chamber-of-commerce members for every business category.

It then chose the first row as the main prospective buyer. That is how scaffolding became a plumbing prospect.

The evidence checks could confirm names, websites, and phone numbers. They did not establish whether those businesses bought the service being researched. The operator withdrew this mission twenty-one minutes after activation, before any result was accepted.

The third search correctly returned zero

The next version filtered for relevance. It also required a review count to support its choice of buyer.

None of the accessible sources supplied that count. The mission returned zero rather than weakening the requirement. That exposed a real design question: what available evidence could support the buyer assessment?

Competitor research produced useful work

A separate competitor-survey mission accepted 21 briefs in its first scheduled run against a target of ten. A broader follow-up accepted 56 from 72 items in its first run. Across the two missions, the record grew to more than a hundred accepted briefs.

There was a cost-control defect here too. Reaching the first target did not automatically close the mission. It continued for forty scheduled slots before settlement, costing about twenty cents.

The day’s lesson is specific: evidence must support the decision the task asks for. Accurate contact details cannot establish buyer relevance. A reached target must also trigger the appropriate stop.

← Back to blog