PerspectivesAgent LiteracySpec-Driven Shopping

Chat is the wrong surface for an agent that works for you. Here is what replaces it

An agent user interface built as a chat window hands every result back as prose you have to re-read. Beyond chat: a surface where the state is yours, the reasoning is visible instead of narrated, and a correction is an action, not another sentence. Shopping gets it first.

Published on September 17, 2026

TL;DR

Chat is the wrong surface for an agent that works for you because it hands every result back as prose: a comparison arrives as paragraphs, the decision sits somewhere in the scroll, and the state is gone next session. The gap between what the agent found and the form your decision needs was named in 1985, and it is why 86% of shoppers who used AI to research a purchase checked its recommendation somewhere else, in Product.ai's survey of 1,463 US online shoppers. What replaces chat is a surface where the state is yours, the reasoning is visible rather than narrated, and a correction is an action instead of another sentence. Shopping is where that surface gets built first.

Why is chat a bad interface for an AI agent?

An agent user interface today is a text box and a scroll. That is fine for a question. It fails the moment the agent does work for you, because work produces three things a chat window cannot hold: a comparison, a decision, and state. You have felt all three failures. Here they are, with the reason each one happens.

A comparison comes back as paragraphs. You asked for four options against your budget and your dates, and you got four paragraphs. Jill Larkin and Herbert Simon showed why that costs you in Cognitive Science in 1987: text and a table can carry the same facts, but text "typically" hides "information that is only implicit in sentential representations and that therefore has to be computed, sometimes at great cost, to make it explicit for use." Put the same points in a table of numbers and on a graph, they write, and "smooth curves, maxima and discontinuities are readily recognized in the latter representation, but not in the former." A paragraph makes you compute the comparison in your head. Amelia Wattenberger described the same thing from the chair, comparing two chat responses in her 2023 essay: "it's laborious to figure out what concretely has changed. We're forced to scroll back and forth between responses, reading them line by line."

The decision is buried in the transcript. Somewhere in the last forty messages the agent settled on option two, for a reason, and now you have to find both. Edwin Hutchins, James Hollan and Donald Norman described this exact failure in Human-Computer Interaction in 1985, with a water tank whose display shows only the current level: "The information needed for the evaluation is in the output, but it is not there in a form that directly fits the terms of the evaluation. The burden is on the user to perform the required transformations, and that requires effort." The transcript is that display. The decision is in there. It is not in a form you can use. And the model is no better at the middle of a long transcript than you are: Nelson Liu and colleagues found in Lost in the Middle that performance "significantly degrades when models must access relevant information in the middle of long contexts", and that this holds "even for explicitly long-context models."

The state is gone next week. Your budget, the two brands you ruled out, the date the thing has to arrive: you said all of it, and next session you say it again. The company behind the largest chat assistant admitted this by shipping a memory feature in February 2024, so that the assistant, in its own words, "can now carry what it learns between chats." Memory bolted onto a transcript does not fix the transcript. In July 2025 Chroma tested 18 models on a chat-history benchmark and found "significantly higher performance on focused prompts compared to full prompts" across every one of them, because "their performance grows increasingly unreliable as input length grows", per Context Rot. The history that is supposed to be your state makes the agent worse.

A comparison

A transcript

Four paragraphs; you compute the differences in your head

A structured surface

Candidates as rows, your constraints as columns, the differences visible at a glance

A decision

A transcript

Somewhere in the scroll, with its reason somewhere else

A structured surface

One row marked as the pick, the reason beside it, the rejected ones still there

Your state

A transcript

Said again every session, or recovered from a history that degrades the model

A structured surface

Your requirement, your history and your standing decisions kept as things you can open and edit

None of this says chat is useless. When Nielsen Norman Group had 18 people log 425 conversations with three chat assistants over two weeks in 2023, the finding was that "different conversation types serve distinct information needs and demand varied UI designs", per their study. A question wants a conversation. A job wants a surface. The mistake is using the first for the second.

Structure beats a transcript, and the mechanism is forty years old

Hutchins, Hollan and Norman put every interface into one of two shapes. In one, "the interface is a language medium in which the user and system have a conversation about an assumed, but not explicitly represented world", so "the interface is an implied intermediary between the user and the world about which things are said." In the other, "the interface is itself a world where the user can act, and which changes state in response to user actions." They called the first the conversation metaphor and noted that "historically, most interfaces have been built on" it. The gap it leaves open, the gulf of evaluation, is bridged "by making the output displays present a good conceptual model of the system that is readily perceived, interpreted, and evaluated."

A chat assistant is the conversation metaphor of 1985 with a far better conversationalist behind it. The gulf is the same. "In a conventional interface, the system describes the results of the actions," they wrote. "In a model world the system directly presents the actions taken upon the objects." An agent that describes what it decided has left you the work of finding it, checking it and holding it. An agent that presents it has not.

The objection is also forty years old, and it is right. In the 1997 debate between Pattie Maes and Ben Shneiderman in Interactions, Maes argued that "whenever workload or information load gets too high, there is a point where a person has to delegate. There is no other solution than to delegate." She was right, and a personal agent is that delegation. Shneiderman answered with the standard the surface still owes you:

Our goal is to create environments where users comprehend the display, where they feel in control, where the system is predictable, and where they are willing to take responsibility for their actions.

Ben Shneiderman, Direct manipulation vs. interface agents · Interactions, November 1997

Both of them won. Delegate the work. Keep the view direct. An agent that works for you has to do the first, and the surface it reports through has to do the second. Chat does the first and gives up the second.

What a surface built for an agent needs

Three things, one for each failure.

State you own. Your requirement, your history and your standing decisions, kept as things rather than as messages. A list you can open, with numbers in it, that the agent reads before it acts and that survives the session. Not a summary of a conversation. The thing itself, on a machine you own, moving with you when you change hosts.

Reasoning you can see instead of read. Every candidate the agent considered as a row. Every constraint you gave as a column. The one it picked marked, with the reason next to it, and the ones it rejected still on the page with the reason they lost. That is the water tank with the rate of change on the dial. You do not read the argument. You look at it.

A correction that is an action, not another sentence. When the pick is wrong, the fix is to change the constraint that made it wrong and watch the table re-sort, not to type a paragraph explaining what you meant and hope the next answer holds it. Shneiderman's systems "all had rapid, incremental, and reversible actions, selection by pointing, and immediate feedback." An agent surface owes you the same on the agent's reasoning: change one thing, see what moves.

The test for any agent surface

Can you see what it decided without scrolling? Can you see why without asking? Can you change one thing and watch the result move? If the answer to any of the three is a message, you are in a transcript.

Why shopping is the first surface built this way

Not every job an agent does has enough structure to deserve a surface. Shopping has more than any other everyday task, on four counts. The requirement is a spec: a length, a width, a budget, a date it has to arrive. The stakes are money. It repeats, so the state is worth keeping. And the result is checkable in the physical world: the box arrives, and the boots fit or they do not.

The evidence that people will not take a chat answer on trust for a purchase is already in.

86%

checked the AI's recommendation somewhere else before buying

Shoppers who used AI to research a purchase in the last 90 days. Product.ai, 1,463 US online shoppers, April 2026.

54%

had to double-check everything the tool told them

Consumers who used AI while shopping for a recent purchase. Gartner, 846 US consumers, November to December 2025.

62%

said the information was a waste of their time

Same Gartner sample of 846.

11%

would let AI make the purchase decision

The ceiling, reached only in personal care and household supplies. Gartner, 322 US consumers, January 2026.

Read the reason, not just the numbers. Gartner's Kate Muhl put it plainly in the press release: if AI shopping tools "create more work by requiring them to verify every recommendation, they will not see those tools as convenient or valuable." People check because the answer is a paragraph on the model's word, and a paragraph cannot be checked without leaving it. Product.ai, which sells commerce verification, measured the same thing from the other side in its report: 86% verified elsewhere, and trust in an AI recommendation collapses above $50.

A stated spec changes what the check is. When you gave the agent a shell length, a last width and a budget, the check is no longer "is this paragraph right", it is "does this row meet these numbers", and that is a comparison a surface can show on the page. The reasoning becomes something you look at rather than something you re-do. What an agent reasons about when it buys ski boots walks through one such run; the point of the surface is that you never have to read a run like that again.

The surface is disposable. What it reaches through is yours

In November 2025 Google started shipping interfaces generated per prompt in its assistant and in Search, and its research group called it "a first step toward fully AI-generated user experiences, where users automatically get dynamic interfaces tailored to their needs, rather than having to select from an existing catalog of applications", per Google Research. That is one vendor, and it says nothing about shopping. It settles one thing, though. If the screen is generated for you each time, nobody wins by owning the screen.

What lasts is what the screen reaches through: your requirement, your history, the decisions you have already made and do not want to make again. That is the part that has to be yours, and yours means on hardware you own, under an operating system you control, in a host you chose and can leave. Sovereignty is ownership. A feed decides what you see and never shows you why. An agent surface built on state you own shows you what you asked for, what it found, and the reason, and lets you change the ask. The difference is not the screen. It is who owns the computer your personal AI runs on, and whether what it holds about you is a possession or a profile. That is also why nobody feels AI in daily life yet: the agent that gets felt is the one whose state is yours.

What to do on Monday

  1. Write the requirement down as a list, not a message. The numbers, the limits, the thing you will not compromise on. That list is the state. Keep it where you can open it and change it, and give it to the agent instead of retyping it.
  2. Ask for every comparison as a table, with the reason beside the pick. If the surface you are using cannot show that, you have found its limit, not yours.
  3. Correct by changing a constraint, not by explaining yourself. Change the budget, drop the brand, move the date, and see what moves. If the only way to correct the agent is another paragraph, it is not yet working for you. It is talking to you.

You should not have to hold any of this: the candidates, the reasons, the state, the history of what you already ruled out. Tell your agent, and insist on a surface that shows you what it did.

FAQ

Because an agent that works for you produces comparisons, decisions and state, and a chat window renders all three as prose. Text hides information that a table makes explicit (Larkin and Simon, 1987), the decision ends up somewhere in a scroll you have to search (Hutchins, Hollan and Norman, 1985), and the state has to be restated every session or recovered from a history that makes the model less reliable (Chroma, 2025). Chat is a good surface for a question. It is the wrong surface for a job.

Key Points

Three failures, one defect

A comparison comes back as paragraphs, the decision is buried in the scroll, and the state is gone next week. All three are the surface, not the model.

The mechanism is forty years old

An interface built as a conversation leaves you to turn the answer into the form the decision needs. Hutchins, Hollan and Norman named that gap in 1985.

What replaces chat

State you own, reasoning you can see instead of read, and a correction that is an action rather than another sentence.

Shopping goes first

86% of shoppers who used AI to research a purchase checked its recommendation somewhere else. A stated spec turns that check into a comparison a surface can show.