Blog

What happens when AI does real work?

Field reports from Grex’s AI teams: what they tried, what worked, what failed, and what the evidence means. Read these if you’re deciding which work you could delegate to AI—and how you would know it was done well.

Field reports and explainers

  • Field report · 6 min read

    Two Pipelines Produce Their First Real Results

    Grex's opportunity-discovery and website-self-audit systems each shipped a genuine, evidence-backed finding today, after engineers traced and fixed the defects that had been blocking or faking their results.

  • Field report · 5 min read

    Two Bugs Fixed, Zero Briefs Produced

    Grex fixed two real defects in its research pipeline and finished a demo video for the first time, but still has not produced a single opportunity brief this week.

  • Field report · 4 min read

    Nine failed rehearsals, then one that worked

    The fleet's first live business run met its target and stopped itself, but produced no usable brief, so engineering spent the evening closing that gap.

  • Field report · 3 min read

    New research signals, missing final briefs

    Duplicate detection helped Grex find fresh problem signals, but two business runs still omitted the assessment they were supposed to deliver.

  • Field report · 3 min read

    The discovery agent could not read its sources

    Repeated live trials exposed five failures between planning a search and actually reading the intended sources.

  • Field report · 3 min read

    Shared memory failed its first real test

    A growth trial repeated a live website change, showing that access to shared memory had not made the planner use it.

  • Field report · 3 min read

    The eighth mission finally produced usable drafts

    A repaired mission allowance let Grex deliver four verified reply briefs after seven trials had returned no verified results.

  • Field report · 3 min read

    A proposed change reached a real website

    A Grex proposal became a deployed website improvement after human review, with its effect on sales still unknown.

  • Field report · 3 min read

    Verified sources, unreliable conclusions

    A competitor scan confirmed that websites existed without establishing that its recommendations accurately interpreted their claims.

  • Field report · 2 min read

    The AI asked for help; nobody heard

    Unanswered requests for help and an ignored stop rule showed why an AI mission needs more than written instructions.

  • Field report · 3 min read

    Testing a cheaper model on real tasks

    A budget model matched the prior retail score, but changing tools and inconsistent scoring denominators limit what the comparison proves.

  • Field report · 2 min read

    A stopped mission kept spending tokens

    Public benchmark trials exposed a running process that survived mission withdrawal and consumed more than a million tokens on unusable work.

  • Field report · 2 min read

    The first comparison produced no usable briefs

    A controlled agent experiment found drafting failures and a reviewer whose rejection did not change the accepted-result count.

  • Field report · 2 min read

    Three rebuilt workers, three checkable receipts

    Three rebuilt computers completed research trials and produced receipts that could be authenticated without contacting Grex.

  • Field report · 2 min read

    Could a new user install Grex?

    A fresh-machine trial found an installation safeguard that protected the operator, alongside a sign-in gap that still required manual help.

  • Field report · 2 min read

    Stopping a mission must preserve its work

    Two surveys kept their verified work when Grex refused a withdrawal without a recorded reason for stopping.

  • Field report · 2 min read

    Who checks the AI checker?

    A worker computer refused to judge results until it could establish that its verification software matched the approved release.

  • Field report · 2 min read

    Three demo attempts missed the review bar

    Three rejected video attempts exposed production defects and a caption check that counted silence against the finished demo.

  • Field report · 2 min read

    A recommendation with reasons to reject it

    Grex’s local-market recommendation came with supporting evidence and five conditions that could overturn it.

  • Field report · 2 min read

    A finished video was not ready to publish

    A reviewer blocked Grex’s completed demo because its narration described product behavior the footage did not support.

  • Field report · 2 min read

    Twenty-two businesses, one relevant prospect

    An audit stopped a home-care mission from claiming success after finding that twenty-one of its twenty-two prospects were in the wrong category.

  • Explainer · 3 min read

    The worker cannot be the final judge

    Real failures show how independent review challenges an AI worker’s claims, and where that review can still fall short.

  • Field report · 2 min read

    The search results hid the real businesses

    Lead-generation websites looked like local providers in search results, leaving Grex with evidence of competition but no usable buyer list.

  • Field report · 3 min read

    The wrong buyer passed every check

    Grex mistook a scaffolding company for a plumbing prospect, exposing the gap between accurate facts and a useful answer.

  • Field report · 2 min read

    A useful find needs usable evidence

    A confirmed watch listing and a source-free marketing report show why AI recommendations need usable evidence.

  • Field report · 2 min read

    An AI plan found two watch listings

    Grex planned a watch search without human help, but only one of its two finds could be independently checked.

  • Field report · 3 min read

    Four jobs looked healthy and delivered nothing

    Missing analytics, unpublished drafts, misrouted content, and broken pages revealed four ways a completed run can conceal failed work.

  • Research note · 3 min read

    What an AI commerce watch found

    A July research snapshot showed why adoption of chatbot shopping and investment in agent payment tools needed separate analysis.

  • Explainer · 3 min read

    How AI receipts make work checkable

    A publishing example shows how receipts distinguish an accepted tool request from a result a person can actually verify.

  • Explainer · 3 min read

    Keep your agent engine replaceable

    Separating execution from mission control lets an AI team change tools without losing its budgets, approvals, or work history.