Objective
Propose one grounded change to WeatherRuta, respect anything already deployed for that site, get the change live, and verify the deployment — proving the memory layer lets a genuinely new proposal through while holding a repeat one back.
Proof / D3 · Growth improvement + memory
Evidence status: checked against the live deployment
One growth proposal, checked against what Grex already knew was done, deployed to a real site, and verified live.
Propose one grounded change to WeatherRuta, respect anything already deployed for that site, get the change live, and verify the deployment — proving the memory layer lets a genuinely new proposal through while holding a repeat one back.
weatherruta.com. One prior completed change already on record (homepage title/meta rewrite and a route_planned event, deployed earlier). The mission was instructed not to re-propose that work and to find the next distinct improvement.
Does memory actually change what Grex proposes next, or is a populated record just decoration?
The final headline follows the evidence.
Case-study record
Each field separates what Grex was asked to do from what it actually produced and what the evidence supports.
What the run was meant to establish.
Propose a new, grounded WeatherRuta improvement that does not repeat a change already on record, then verify it after deployment.
The sources, boundaries, exclusions, and time window supplied to the work.
weatherruta.com, one existing done-ledger record (homepage rewrite + route_planned event), a $1 metered spend cap, and an explicit instruction not to re-propose the homepage work.
The work product that was actually handed back.
A new page at /route-weather-planner/: full title, meta description, H1, section structure, and body copy targeting "check weather along my route," plus a homepage nav link and a contextual link to the existing route map.
The evidence supporting the finding and any observed effect.
Commit b7f4e0f, Bitbucket PR #5 merged to main, deployed to Firebase Hosting. Independently re-checked live: HTTP 200, exact title tag, and the homepage link to the new page, all confirmed after deployment.
The decision made from the evidence.
Approved by the operator, deployed, and accepted with commit/PR/deploy evidence attached. The homepage change stayed untouched and was correctly never re-proposed. The 14-day route_planned readback comparing before and after is still pending — not yet claimed.
How long the bounded run took from start to settled outcome.
The winning proposal-to-deploy mission ran inside its 6-hour window. Getting to a clean run took real debugging first: six distinct memory and pipeline defects were found and fixed across roughly a day of live iteration before the mechanism worked end to end — recorded here rather than hidden.
What was metered, what was not priced, and which cost basis applies.
$0 of a $1 metered spend cap on the winning mission. Five earlier attempts that hit real defects were stopped at $0 each rather than left to spend against a broken run.
The models, runtime, and compute context that produced the work.
Strategist, referee, and axis-judge roles on Gemini 3.1 Pro (high); drafter and relevance roles on MiniMax-M3; running on Grex's Timmy Nidus.
Human approvals, corrections, reruns, or other changes during the run.
The operator approved the topic and the publish decision explicitly — publishing new content to a live site is reserved to the operator, not an autonomous step. The page copy itself was fully Grex-authored and deployed unedited.
What the evidence does not establish.
This proves the propose-deploy-verify loop and correct locator-level memory (a genuinely new page went through while the already-done homepage change was correctly held). It does not yet prove business uplift — that depends on the still-pending 14-day route_planned comparison.
This case study remains unpublished as a claim until the actual demo output, source links, cost record, and limitations are reviewed. A missing deliverable stays missing in the final account.
Request a walkthrough