Blog
Twenty-four watch passes on agentic commerce: what a research loop found
A scheduled research agent read one market four times a day for a month. The consumer-facing agentic checkout story is retrenching while the payment rails underneath it consolidate.
We pointed a scheduled research agent at agentic commerce and let it run four times a day for a month. Twenty-four passes later the picture it assembled is not the one the category’s own marketing tells. Consumer-facing agentic checkout is retrenching, with the highest-profile launch of 2025 shut down and the best public conversion data running below the plain website. Meanwhile the payment rails underneath it are consolidating fast.
Below is what the loop found, how it is configured, and where it stops itself.
How is the loop configured?
The agent runs at 00:00, 06:00, 12:00 and 18:00 daily against a single named focus. It has two modes and it chooses between them itself.
It starts in backfill: deep, rearward-facing grounding, pass after pass, building a picture of the market’s history and structure. A ledger tracks how many passes have run, what coverage they have reached, and how much novelty the last pass produced. When novelty drops to low or none and coverage passes its bar, the ledger flips the focus to watch, and the cadence collapses to a light “what has changed” pass.
That mechanism matters more than the schedule. The loop reads until it stops learning, then switches to watching. Without it, a research agent on a timer produces the same summary forever and bills you for it.
Two rules govern the output. Every factual claim carries a source marker, and the source list is appended to the brief. A pass that cannot ground its findings withholds the brief rather than shipping an ungrounded one. Everything lands in a work graph as a proposal, not a publication, so a human decides what leaves the building.
Everything below is drawn from one brief, pass 24, dated 2026-07-29. We have removed the internal action items. The findings are unedited.
What did it find? Agentic checkout is retrenching
The headline is a reversal.
Forrester reports that “news broke in March 2026 that OpenAI is shuttering its Instant Checkout initiative due to lackluster performance,” and that “consumer adoption of agentic commerce remains low and relatively stagnant.” Most merchants running ChatGPT apps, per the same source, “redirect customers to their sites for checkout” [S1].
The conversion data behind that is the more useful part. Citing a Walmart interview, Forrester reports that in-ChatGPT Sparky conversion ran “three times lower for the selection sold directly inside the [ChatGPT Instant Checkout] chatbot,” and that pilot users of Sparky in ChatGPT converted at “roughly 70% of the rate of those using Walmart.com directly” [S1].
That is the number worth sitting with. Not that agentic checkout is unpopular, but that for the merchant with the most to gain, buying inside the assistant converted worse than buying on the website. Walmart is nonetheless proceeding, with Sparky embedding in ChatGPT first and Gemini later, and discussions underway with Anthropic [S1].
Why are the rails consolidating anyway?
Because the infrastructure bet is not the same bet as the consumer surface, and the infrastructure players are moving regardless.
Stripe and Tempo announced the Machine Payments Protocol, with Visa as a design partner, aimed at cases like “AI agents autonomously paying a fee to a network provider per API call for web access” [S1]. Stripe’s own Sessions 2026 recap confirms merchants can “accept payments from agents over MPP in stablecoins as well as fiat through cards, Klarna, and Affirm via Shared Payment Tokens (SPTs), using our Payment Intents API” [S2].
Google’s Universal Commerce Protocol gained a Gap partnership in March 2026 along with “an AI agent-specific shopping cart, product catalog access, and identity linking,” with Stripe and Salesforce implementing UCP [S1].
Stripe announced 288 products and features at Sessions 2026. The agentic subset alone [S2]:
| Announcement | What it does |
|---|---|
| Agent-ready financial accounts | Agents check balances, pay invoices, store funds, create cards, send money, “with human-in-the-loop confirmation for key actions” |
| Agentic Commerce Suite | Upload a catalog, manage agent access from the dashboard |
| Link agent wallet | Spending approvals and “full purchase visibility” |
| Meta partnership | “Native checkout inside ads on Facebook” |
| Stripe and Google | Buying “in AI Mode and the Gemini app” via UCP |
| Authorization Boost AI | Acceptance rates up “an average of 3.8%,” processing costs down “up to 3.3%” |
Outside the Western stack, Alipay launched an agentic payments protocol supporting commerce in Alibaba’s Qwen app, which reported 300 million monthly active users as of its December 2025 earnings [S1].
The synthesis the loop reached, and we think it is right: the consumer experience of shopping through a chatbot is underperforming, while the machine-to-machine payment layer is being standardised by the incumbents anyway. Those are different markets with different timelines, and conflating them is how the category’s forecasts got made.
Worth noting what recurs in the primitives above. Spending approvals. Human-in-the-loop confirmation. Full purchase visibility. Agent access managed from a dashboard. The rails are being built with gates in them, which suggests the people building payment infrastructure do not expect unsupervised agents to be acceptable to merchants any time soon.
What is the regulatory surface doing?
Hardening, quickly, and with private rights of action attached.
Washington HB 2225 was signed on March 24, 2026 and takes effect January 1, 2027. It requires disclosure at the start of every interaction plus reminders “every three hours for adults and every hour for minors,” enumerates eight prohibited manipulative techniques including “simulating emotional distress…in response to the user’s desire to end the chat” and “soliciting gifts, purchases or other expenditures framed as necessary to maintain the user’s relationship with the chatbot,” requires annual public reporting of “the number of crisis referral notifications issued to users in the preceding calendar year,” and carries a private right of action alongside state AG enforcement [S7].
It is not alone. California SB 243 took effect January 1, 2026, and New York General Business Law section 1700 took effect November 5, 2025, with Colorado, Maine, Texas and Utah maintaining a disclosure-only baseline [S8].
On the demand side, Common Sense Media reported in July 2025 that “72% of teens have used AI companion chatbots at least once,” more than half a few times a month, and “one in three teens” use them for social interaction and relationships [S8].
A private right of action changes the calculus for anyone shipping a conversational product. Compliance stops being a regulator’s problem you might get audited on and becomes a plaintiff’s-bar problem you get sued over.
What does the loop do when it stops learning?
It proposes its own next questions. Pass 24 ended with these still open:
- OpenAI Instant Checkout sunset mechanics: which merchants were on it, what is the migration path, and does UCP absorb them?
- Walmart and Anthropic Sparky integration status: is a public launch imminent?
- MPP public spec status: where is the canonical spec?
- Washington HB 2225 enforcement venue: which courts will see the first private-right cases?
- Is there a 2026 Common Sense Media update with usage broken out by age?
Those become the next pass’s grounding targets. The loop is not answering a question we asked. It is maintaining a position on a market and telling us where the position is thin.
One caveat we will not paper over. This is a research instrument, not an oracle. It surfaces and cites; it does not verify that a cited source is correct, and a confidently wrong secondary source will propagate. Everything above is attributed so you can check it, which is the point. Of the nine sources in this brief, we fetched and confirmed the three that carry the most weight before publishing.