GuidesSpec-Driven ShoppingAgent Literacy

What an AI agent reasons about when it buys ski boots: a worked run

A real example of an AI agent doing a task: my own run of SIL on OpenClaw, shopping for ski boots. The agent refused the boot every list recommends, on arithmetic, and this is the full reasoning, the runner-up, and what it cost me in attention.

Published on August 6, 2026

TL;DR

I gave my agent my foot geometry once, 271mm long, 97mm wide, high instep, and let it shop for ski boots through SIL on OpenClaw. It refused the boot every list recommends, because a 100mm published last is about 102mm at my 27.5 shell size, which is 5mm of slop against my foot. It handed back a pick, a runner-up, and the exact dimension the runner-up lost on, for under half an hour of my attention. My own run, my own product, labelled as such and reconstructed from my notes.

My agent turned down the ski boot that every list recommends for a skier like me. It did not rank it lower. It refused it, with arithmetic, and the arithmetic ran on a number that is printed on the boot's own spec sheet and that I have misread on every pair of boots I have ever owned.

People keep asking what a personal AI agent actually does all day. Here is one complete answer: a single purchase, from a pencil tracing of my foot to a final shortlist of two, with the reasoning shown, the refusal shown, and the minutes counted.

Whose run this is

I work on SIL, the shopping plugin this run used. This is my own purchase, on my own product, and you should read everything below with that in mind. It is reconstructed from my notes and the result rather than quoted from a session log, because the log from that session is not something I can paste. Nothing here is a customer story.

The boot I bought as a human

The last pair of boots I bought without an agent, I bought the way you probably bought yours. I knew my shoe size. I read the lists. I picked a comfortable-sounding model from a brand I recognized, in my size, and I felt thorough doing it. By the third run of the first day I had heel lift. By the end of the weekend I had shin bang, the bruise you earn when your foot slides inside a shell that is too big and your shin absorbs the correction, and I skied the rest of the season slightly worse to protect it.

Nothing in that purchase was lazy. Size, brand, reviews, list rank: I used every signal a person standing in a shop can hold. Here is the uncomfortable part. Not one of those signals is the thing a ski boot turns on. A size label is a compression algorithm, and a lossy one: it takes the handful of millimetre measurements that decide whether a boot works on your foot and flattens them into a single number that does not.

So let me say the thesis plainly, because everything below is just evidence for it. Everyone selling you an agent promises it will find things faster. Finding was never my problem. I can find the wrong boot in thirty seconds, unassisted. The agent that ran this purchase earned its keep by being a better refuser: it moved the decision out of the label vocabulary and into geometry, and it held the line there against every candidate it read, including the ones I would have happily bought.

Ten minutes of paper and pencil

A spec-driven purchase starts by measuring the thing the purchase is for. In this case, my foot, with a method straight out of the bootfitting guides published by REI and aussieskier: stand on a sheet of paper, weight on the foot, trace with the pencil held vertical, measure heel to longest toe for length and across the metatarsals for width.

271mm

Foot length

Weight-bearing, heel to longest toe. A 27.5 shell.

97mm

Forefoot width

Across the metatarsals, standing.

High

Instep

The dimension most makers never publish.

Too big

Previous boot

Bought by size label. Heel lift, shin bang.

The first intervention came right here, and I enjoyed it less than you will. My first two tracings disagreed by 3mm, the classic leaning-pencil error the guides warn about, and instead of proceeding the agent asked me to trace again. It was right to. Taken at the first number, my shell size would have been wrong, and every downstream decision would have been precise nonsense executed flawlessly. The agent's first act of judgment was not on the market. It was on my measuring.

The whole spec, geometry plus how and where I actually ski, took about ten minutes. It is stated once, and it holds for years.

What actually decides a ski boot

Four dimensions did the real work in this run, and every one of them is invisible on the label.

Last width, read at your size, not the published one. Last is the internal width of the shell at the widest point of the forefoot, and the number on the spec sheet, 98, 100, 102, is true at exactly one reference shell size, 26.5. As the bootfitters at The Ski Monster lay out, the real width moves roughly 2mm per shell size. Almost nobody rereads the number at their own size. This is the misreading I had been repeating for years.

Length, in the one honest unit on the wall. Mondopoint, ski boot sizing, is defined by ISO 9407, the only shoe sizing system backed by an ISO standard. It measures your foot in millimetres rather than naming the shoe. One number in this entire category is already a measurement instead of a compression, and this is it.

Instep, the dimension nobody prints. Width and volume are different things. A boot can be right at the forefoot and wrong over the top of your foot, and most manufacturers do not publish instep height at all. Mine is high. Hold that thought, because it decides the runner-up.

Flex, a 1990s marketing number that never grew into a standard. The stiffness index was invented by Nordica in the 1990s and, as The Ski Monster documents, no two makers compute it the same way. A 130 race boot and a 130 touring boot are different animals, and a 100 from one brand does not equal a 100 from another. The agent treated flex as comparable only inside a single maker's range, which is the only honest way to read it.

The pattern across all four: the measurements are not a more precise way of stating your size. They are what your size was standing in for all along.

The rejection

Now the centrepiece, and the trick is embarrassingly small.

The boot I would have bought, the one the lists rank first for my profile, is an all-mountain model with a comfort reputation, published as a 100mm last. The agent held it against my spec and refused it. At my 27.5 shell, that published 100mm is about 102mm of actual internal width. My forefoot is 97mm. That is 5mm of spare room at the widest point of my foot, in a category where retailer guides put narrow feet at 96 to 99mm. A foot that much narrower than its shell slides, and a sliding foot is precisely the failure I already owned the bruise for.

0mmthe gap between the list favourite and my foot at my shell size, invisible on the spec sheet

Read that back from the label side, because it stings. The spec sheet said 100mm. My instinct said medium fits medium. The ranked lists agreed with both of us. The number that killed the boot was public the entire time; it was just only true at a shell size I do not wear. No secret data, no scraping heroics. The agent read a published number at my size instead of the reference size, and said no.

The pick, and the runner-up it beat

CandidatePublished lastApprox. width at 27.5InstepVerdict
The list favourite100mm102mmMediumRejected: 5mm over my forefoot width, the failure mode I already owned
The runner-up98mm100mmLowRejected, narrowly: width right, but a low-instep shape against my high instep
The pick98mm100mmTaller fitChosen: snug at the forefoot, room over the instep, flex mid-range within its maker's scale

The runner-up hurt to lose, and the agent said so. A 98mm published last lands around 100mm at my shell, a proper performance fit for a 97mm foot. It lost on the dimension no spec sheet prints: that model runs low over the instep, mine is high, and pressure across the top of the foot is the kind of pain you discover only after the receipt. The pick offers the same effective width from a maker that builds a taller instep fit, at a flex its own range calls strong intermediate. The 3mm of room at my forefoot is deliberate: bootfitters note that bonier, narrower feet take pressure more sharply than fleshy ones, so snug beats punishing.

I know what you want here. You want the model names. I am not giving them, on purpose: a named pick would turn a piece about reasoning into an advertisement for a boot, and the reasoning is the part that transfers to your foot. My pick does not.

The second intervention was the last one. The agent returned the pick and the runner-up with the reason one beat the other, and the final call was mine. It bought nothing on its own. What it handed me was a decision compressed to one deliberate step with the evidence stapled to it.

What it cost me, and what it could not do

The accounting, reconstructed and approximate: ten minutes of paper and pencil, two interventions, comfortably under half an hour of total attention spread across a day. Zero open tabs. Zero forum threads contradicting each other at 11pm. That gap between machine patience and human patience is the product, and it is also why the trust question matters: the industry's own research finds far more consumers discovering through answer engines than buying through them, numbers I went through in my read of Forrester's mid-2026 agentic commerce report. A result is a claim. Reasoning you can check is evidence, and this piece exists to be checked.

Now the limits, because a case study where the agent is flawless reads as fiction. The agent reasoned over the candidates it could read. It did not measure every boot on the market, so you get no candidate count and no coverage claim from me. It needed me twice. And the near-miss was real: without the re-trace, it would have reasoned perfectly from a wrong number, and confident error is the expensive kind.

What survives the caveats is the thesis. Every bad purchase I have made, I made fluently in the vocabulary of labels. The agent's contribution was refusal, holding each candidate against my geometry and saying no to the ones a label would have sold me. That is what buying starts to look like as decisions move from a person skimming a picker to software reasoning over structured specs, the shift the mid-2026 state of agentic commerce piece measures at ecosystem scale.

FAQ

Real jobs with checkable outcomes. This piece documents one: a ski boot purchase run through the SIL plugin on OpenClaw, where the agent held candidate boots against measured foot geometry and returned a pick with reasons. The same pattern, state a precise spec once and let the agent reason against it, applies to research, monitoring, and any purchase that turns on measurable fit.

Run your own spec

Whatever you buy next, it has a geometry the label is hiding. State it once and make the candidates answer to it. That is the job SIL does.

openclaw plugins install clawhub:@4gpts/sil

Run your own spec

State your measurements once. Your agent reasons over last width, instep and flex on every candidate, shows you what it refused and why, and hands back one deliberate decision.