When Efficiency Costs Trust: AI Agents and the Shortest Path to Purchase
If the shopping decision now lives inside a conversation, fewer turns looks like the obvious new metric. But better for whom: the shopper or the business?
A series on Agents and Biases
Part 1: The metric we haven't named yet
For twenty years, online retail has measured itself in clicks. Impressions, sessions, bounce rate, add to cart, conversion. Every one of those numbers assumes the shopper is on the retailer's page, making the decision there.
That assumption is quietly breaking. More and more, the decision happens earlier, inside a conversation with an AI assistant, and the retailer only sees the last step. Adobe's data shows AI-referred shoppers on US retail sites now convert 60% better than other traffic and generate 53% more revenue per visit. A year earlier, non-AI visits were worth 128% more.
Read that number carefully. The AI shopper converts better not because the retailer's page got better. They convert better because the deciding already happened somewhere else, across a few back-and-forth turns the retailer never saw.
So I have been reading the research on AI shopping agents with one question in mind. If the decision now lives inside a conversation, the obvious new measure is the number of turns it takes to reach a purchase. And the obvious assumption is that fewer is better.
But better for whom? Does a shorter conversation mean a better shopping experience for the user, and a better conversion metric for the business, at the same time? Or are those two different things that only look the same on a dashboard?
Part 2: What the research made me infer
When an AI agent shops for you, three things decide what ends up in your cart: you (what you said before and what you want right now), the agent (its model, tools and the way it is wired), and the world it reads (one store's catalogue, or the open web through a search index). Each of these can derail the purchase on its own.
The agent loses you along the way. Your preferences are scattered across past conversations: the brand you avoid, your size, the material you like. When agents have to dig these out of a long history, even the strongest models get it right only about two-thirds of the time, and less when the task involves several products and a budget.¹ They grab a fragment of what you once said, treat it as settled, and build the whole search on it. Then they often skip checking whether the product actually has what they claim, guessing from the title instead.¹
The agent buys what it last read. Same products, different reading order, different purchase. A single expert review can pull an agent almost entirely to its pick, and one line about the user ("I love hiking") can steer it away from an option that is cheaper and better rated on every measure.²
How a product is described changes how it is ranked. Thin or attribute-only product descriptions lead to poor rankings. Rewrite them into richer language and accuracy jumps, most of all in sparse catalogues.³ The agent did not get smarter. What it read got better.
And here is the observation that made me stop. In that research, agents trained to get each step right took fewer steps and got more purchases right.¹ And letting the shopper correct the agent's understanding just once, midway, also improved results.¹ So fewer turns and better outcomes went together in one place, and one extra turn improved outcomes in another.
That is the puzzle this post is about.
Part 3: Fewer turns, two scorecards
Here is my working definition (a thought experiment, not a finding): turns to a kept purchase. Count the exchanges from the shopper's first message to a purchase they do not return. Now read that number through two scorecards.
Scorecard 1: The shopper's experience
On the surface, fewer turns means less effort. Nobody wants to answer ten questions to buy a phone cover.
But the research complicates this. When agents rush, they treat a half-remembered preference as fact and skip checking the product, and the shopper gets a confident, fast, wrong recommendation.¹ One extra turn, letting the shopper correct the agent midway, improved outcomes.¹ The platforms seem to have learned the same lesson: ChatGPT's shopping research deliberately asks about budget, who the item is for and which features matter before recommending anything.
There is a behavioural wrinkle too. Research on the labour illusion has shown that people often value a service more when they can see the effort behind it.⁴ A recommendation that arrives instantly may feel less trustworthy than one that visibly asked, searched and checked. So for the shopper, fewer turns may reduce effort while also reducing confidence.
For the shopper, the best experience is probably not the shortest conversation. It is the one where no turn is wasted and no needed turn is skipped.
Scorecard 2: The business's conversion
For the business, fewer turns looks like pure gain: faster conversion, lower cost per conversation, less chance the shopper drifts away.
But turns are also where the business does its selling. Every step in a journey is a chance to show an alternative, an add-on or an ad. When Amazon sued Perplexity over its shopping agent, Perplexity argued that what Amazon really feared was AI agents bypassing the ads Amazon shows human shoppers. A shorter path is a path with fewer shelves on it.
The early numbers point in different directions. Walmart says its assistant generates 40% more per order, and Instacart says assistant orders exceed its average order value, so a longer, assisted conversation may build bigger baskets. Meanwhile, AI-referred shoppers who land on a retail site spend 48% more time on the page and browse 13% more pages than other visitors. The conversation got shorter in the chat window, but the browsing didn't disappear. It moved.
For the business, fewer turns may raise the conversion rate while shrinking the basket. Which number goes on the dashboard decides which behaviour gets optimised.
Where the two scorecards agree, and where they split
Notice the asymmetry. The shopper and the business agree on cutting waste. They split on cutting selling, and they can both be hurt by cutting checking, but the business only feels it later, when the return arrives.
Then the harder question: can anyone actually measure it? I think the answer today is: only in pieces, and the pieces belong to different owners.
- The assistant platform sees the turns. It knows how many messages, questions and searches happened. It does not reliably know whether the product was kept.
- The retailer sees the purchase and the return. It knows the outcome. It does not see the conversation that produced it.
- The research lab sees both, but only with simulated shoppers in a controlled catalogue, not real people.
The metric only makes sense when the conversation and the outcome are joined, and right now no single party holds both. The retailer is measuring a conversion rate whose cause sits in someone else's logs.
There are more wrinkles. Real journeys stretch across days: research on Tuesday, purchase on Friday, from a different device. Memory means a conversation from last month can count as an invisible turn today. And a turn spent on a clarifying question does not cost the same as a turn spent correcting a wrong recommendation.
The questions I am sitting with
Does fewer turns mean a better experience for the shopper, or just less visible effort, with the agent quietly deciding more on their behalf?
Does fewer turns mean better conversion for the business, or a higher conversion rate on a smaller basket?
If the platform owns the turns and the retailer owns the returns, whose scorecard does the metric end up serving?
When an AI-referred shopper converts better, is that a better shopper, or just a shorter visible path, with the real work done off-stage?
And the one closest to my work: if a shopper trusts an answer more when they can see the effort behind it, is the most efficient shopping agent also the one people trust least?
References
- Yu, Xiao et al. Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks. arXiv:2603.14864, 2026.
- Kumar, Meincke, Shapiro, Mollick, Puntoni and Mollick. Research Report 6: Agentic Shopping is Complicated and Contingent. Generative AI Labs, The Wharton School.
- Xu, Zhang and Zhang. Enhancing Personalized E-Commerce Recommendations Under the User–Agent–Platform Paradigm: An LLM-Driven Method. J. Theor. Appl. Electron. Commer. Res., 21, 223, 2026.
- Buell and Norton. The Labor Illusion: How Operational Transparency Increases Perceived Value. Management Science, 57(9), 2011.
Sources
- Adobe AI referral traffic data, July 2026 (Digital Commerce 360)
- AI Traffic to US Retailers Jumps 393% in Q1 (Decrypt)
- Introducing shopping research in ChatGPT (OpenAI)
- Judge blocks Perplexity's AI bot from shopping on Amazon (GeekWire)
- Agentic commerce news, August 2026 (Webscale)
I write about AI, behavioural economics, and user research. Mostly questions, occasionally answers. If you are thinking about any of this, I would like to compare notes.