Why your agent picks the wrong tool, and how to fix it in your plugin
Your agent never sees your code. It chooses between tool descriptions, which makes a tool surface a piece of writing. Here are three ways ours was written badly, what we changed, and the audit you can run on what you already have installed.
TL;DR
Your agent picks the wrong tool because it is not choosing between your tools. It is choosing between your tool descriptions, and it has never seen your code, your tests or your README. Two plugins whose descriptions overlap put the model in a coin flip it has no information to resolve, which is why a surface that worked last week breaks when you install something adjacent. The fix is editorial rather than technical: write each description as when to reach for this rather than what this does, collapse tools that differ by one noun into one tool with a parameter, and move anything the model should never choose behind another tool. This piece uses our own Apache-2.0 plugin tree as the worked example, including the parts of it that are still wrong.
Last week your agent reached for the right tool every time. This week it does not, and you did not touch that plugin. You did not have to. You installed something else, and the new arrival's description reads a little like one you already had.
Here is the argument, and you are welcome to disagree with it: a tool surface is a piece of writing, and most plugins are written badly in a way that has nothing to do with how well they are engineered. You can have exhaustive tests, clean types and a tight API, and still ship a surface the model cannot navigate, because none of those things are in front of it at the moment it chooses.
What the model actually sees
At the moment of selection, the model has three things per tool: the name, the description, and the parameter schema. That is the entire interface.
It does not have your repository. It does not have the README that explains which tool is for the setup path and which is for the hot path. It does not have the integration test that encodes the ordering you consider obvious. It has a flat list of short texts, several of which were written by different people, in different weeks, in different moods, and it is picking between them on similarity to the request in front of it.
So the question "why did it pick the wrong tool" almost always has the same answer. Read your descriptions in a flat list with no repo context, the way the model gets them, and the wrong pick usually becomes the obvious pick.
This is not a prompt-engineering trick
Nothing here is about phrasing the user's request better. The reader of your tool descriptions is the model, on every turn, before the user says anything. Descriptions are the part of your plugin that runs most often.
Three ways we got this wrong, in a repo you can read
The examples below are all from Context4GPTs/sil-openclaw, which is Apache-2.0, along with its sibling Context4GPTs/klodi-plugin. Every quoted description is verbatim, including the one that embarrasses us.
1. A tool that was never a decision
We shipped a tool called sil_specs, labelled "Canonicalize coined spec definitions". Its description contained this sentence:
Internal plumbing - never surface ns.key to the buyer.
Read that as the model would. A description that has to instruct the model to hide its own output is not describing a decision the model should be making. It is describing a function call that some other tool should have made on its way to doing something the user asked for.
We deleted it in PR #74, commit b1d0910, titled "Retire the dead sil_specs tool and its client". The change removed 2,575 lines against 126 added and took the surface from eleven tools to ten. The changelog entry says "No compat alias, no deprecation stub", which is the right call for a tool nobody should have been calling directly in the first place.
The general rule: if a tool's description needs to explain that its result is not for the user, it is plumbing. Put it behind the tool that has a reason to call it.
2. A name that described the store instead of the act
sil_remember became sil_learn. No alias, recorded in the changelog under Removed.
That looks like bikeshedding until you think about the flat list again. remember describes what the storage layer does. learn describes what the agent is doing on the user's behalf, which is the thing the model is trying to match against the request. Name the verb for the act, not for the substrate, because the substrate is not what anybody is asking for.
3. Two failures that read identically, so the agent guessed
This is the one that actually cost us. Our search tool could refuse a call for two very different reasons: the category the agent asked for does not exist in the registry, or a requirement it passed could not be applied. Those are opposite problems with opposite next moves. On the wire they looked the same.
So the agent guessed, and half the time it guessed wrong, and the pattern looked like flakiness rather than a design defect.
The fix was not a better error string. It was giving the model a way to ask, in the form of a fifth catalog tool, sil_domain_find, shipped in PR #76 as "the fifth catalog tool, so the read route is reachable". Then we wrote the disambiguation into the search tool's own description, where the model would meet it:
A refusal naming the domain and a refusal naming a predicate read alike on the wire, so settle which one it was with sil_domain_find - a
pathprobe states whether the domain stands.
That sentence is doing work no error code could do. It tells the model what the ambiguity is and what to call next, at the moment it is holding the ambiguous answer.
What we shipped
What we changed
sil_specs
A registry call exposed as a tool, told to hide its own output
sil_specs
Deleted. 2,575 lines removed, eleven tools to ten, no alias
sil_remember
Named for what the store does
sil_remember
Renamed sil_learn, for what the agent does. No alias
Two refusals
A missing category and a rejected requirement read alike, so the agent guessed
Two refusals
A probe tool to ask with, and the ambiguity written into the search description
sil_specs
What we shipped
A registry call exposed as a tool, told to hide its own output
What we changed
Deleted. 2,575 lines removed, eleven tools to ten, no alias
sil_remember
What we shipped
Named for what the store does
What we changed
Renamed sil_learn, for what the agent does. No alias
Two refusals
What we shipped
A missing category and a rejected requirement read alike, so the agent guessed
What we changed
A probe tool to ask with, and the ambiguity written into the search description
When to collapse several tools into one
The opposite failure is a surface that splits one decision across several near-identical tools. The model then has to tell them apart on a distinction you understand and it cannot see.
Our profile lifecycle could easily have been three or four tools: create a method, rewrite a method, create a PRD, attach an asset. It is one, described as:
The single target+change verb owning the sil shopper's method/PRD lifecycle.
target says where the change lands, kind says what the change is. The words the description then spends are not on the API shape, they are on the rule that actually trips an agent up: "To change anything, WRITE the whole reconciled doc - there is no append/amend, so a correction never stacks a contradicting line."
The test is simple. Write the descriptions of your two candidate tools side by side. If they are the same sentence except for one noun, they are one tool with a parameter, and keeping them apart is asking the model to resolve a difference you have not explained.
Tool or skill
A tool is a decision the model makes at one moment, so its description answers when to reach for this.
A skill is behaviour that spans several of those moments, so its description answers when this whole workflow applies. Ours opens exactly that way, in the frontmatter of sil-shopping/SKILL.md: "Use when the user explicitly asks to shop with sil or manage their sil shopper", followed by the intents it covers and the tools it drives.
The part worth stealing is where the detail lives. The skill body carries the rules that are not any single tool's business, such as "Act, don't narrate" and the rule that a stored price must be quoted with the date it was read. Everything below that sits in five files under references/, loaded on demand. The window holds the routing gate; the disk holds the manual.
That is the honest answer to the context question. You do not fix a full context window by trimming adjectives out of descriptions. You fix it by deciding what has to be resident for the model to choose correctly, and putting the rest somewhere it can be fetched.
Where we are still failing
Thirteen is a lot. It is fewer than it was, and each one currently earns its place, but a surface that size is a standing risk of exactly the overlap this piece is about, and the honest position is that it is a ceiling rather than a target.
The worse one is in the same file. openclaw.plugin.json carries a security.packagingNote that runs past four thousand characters in a single string, sitting in the manifest next to the thirteen tool names. It is written for a human reviewing what the plugin is allowed to touch, which is a real need. It is still the wrong place, because it puts a wall of prose in a file whose job is to declare a surface, and it is shipped in the version live today.
We know. It is on the list. We are telling you because a piece about writing good tool surfaces that only quoted our good ones would be an advertisement, and you would be able to tell.
One more, on the same principle. Our search tool reports per requirement whether it could apply it, and today the common answer is that it could not, because spec values are not ingested yet. Rather than hide that, the description makes the model say it: a requirement reported applied: false is "NOT VERIFIED - neither a match nor a miss". A description that forces an honest answer out of an incomplete backend is worth more than one that lets the model round up.
The audit, on what you already have installed
You can run this today, on any host, without changing a line of anyone's plugin.
- 1
Dump every loaded description
Print the tool name, description and parameter schema for every tool your host currently has loaded, in one flat list, with no repo or plugin grouping. This is what the model sees.
- 2
Find the overlapping when
Read for pairs whose descriptions answer the same when. Two tools that both plausibly serve one request are a coin flip, and the model resolves it on wording you did not intend as a signal.
- 3
Test each description for when, not what
Ask of each one: does this tell me the moment to reach for it, or only what it does? A description that is only a restatement of the function name is a tool the model will pick by luck.
- 4
Count what is resident before you ask anything
Total the descriptions and schemas that are in the window on turn one. That is your standing cost, paid on every request, whether or not any of it gets used.
Then act in order of cost, cheapest first:
- Uninstall. The fastest fix for an overlapping pair is usually that you do not need both. This costs nothing and is reversible.
- Rewrite the descriptions you own. Turn each into when to reach for this, and name the failure mode and the next move, the way the search tool above does.
- Collapse. Two tools that differ by one noun become one tool with a parameter.
- Push detail out of the window. Whatever the model does not need in order to choose correctly can be a reference the skill loads on demand.
- File the issue for the ones you do not own. A plugin whose descriptions do not say when is a bug report worth writing, and most maintainers have never looked at their surface in a flat list.
FAQ
Because it is choosing between descriptions, not between tools, and two of yours probably answer the same when. The model never sees your code, your tests or your README. It sees a flat list of names, descriptions and parameter schemas, and it picks on similarity to the request. If a surface that worked last week is failing now, look at what you installed since: the usual cause is a new tool whose description overlaps one you already had.
You should not have to hold any of this
If you are new to the pieces underneath this, the install order for a personal AI stack is the read that comes before this one: host, model, plugins, skills, memory.
But the point of writing a tool surface properly is that the person who installs it never thinks about any of the above. They ask for the thing they wanted, the right tool gets picked, and the mechanism stays invisible. Nobody should be keeping thirteen descriptions in their head, including us. That is the job of the plugin, and when it is done well you get to forget it exists.
Read the surface for yourself
Both plugin trees quoted here are public and Apache-2.0, tool definitions and all. The good descriptions and the four-thousand-character one are in the same repository.