A desk, not a memory
Everything your agent works with has to sit in one space at the same time: your rules, the conversation, every page it read, every plugin you installed. The space has an edge.
Your agent does not forget the way a person does. It works from a desk with an edge: your rules, the whole conversation, every page it read and every plugin you installed share the same space. When the desk fills, the older part becomes a summary and your rule falls out. Here is why it happens, and a ten-minute routine that stops it.
Your AI agent does not have a memory the way you do. It has a desk: your instructions, the whole conversation, every page it read for you and everything your plugins bring along have to fit on it at once. When the desk fills, the older part of the conversation is replaced by a short summary, and your rule can fall out of it. A crowded desk makes the agent less reliable well before that, so keep one job per conversation, put standing rules where the agent re-reads them, and when it drifts, start fresh with a handoff note instead of correcting it.
On Monday you told your agent two things: answer in short bullet points, and never book anything without asking you first. It did both, all day. On Wednesday it sent you four paragraphs. On Friday it booked a table.
You did not change a setting. You did not install anything new. It simply stopped doing what you said, and when you reminded it, it apologised and did it again an hour later.
You are not imagining it, and you are not alone. People ask this in developer forums, in Reddit threads about conversations kept open for two weeks, and often enough that a consumer tech site ran a how-to on it in June 2026 and a software vendor has a help-page entry titled "Why doesn't the agent remember what I said earlier in a conversation?". The answer is not that the agent is getting worse at the job. It is running out of room.
The model behind your agent does not carry anything from one moment to the next. Every time it answers, it reads everything in front of it from the top, and answers from that. What the industry calls the context window is the size of what can be in front of it at once. Think of it as a desk.
Here is what is on that desk every time it replies:
The desk is finite. One model developer's engineering guidance puts it plainly: context is "a finite resource with diminishing marginal returns", and every extra piece of text spends a little of the model's attention. Nothing on the desk is free, including the things you never see.
When the desk is nearly full, the program that runs your agent (its host) does something sensible. It takes the older part of the conversation, writes a short summary of it, and carries on from the summary plus your most recent messages. In the host most personal-AI readers run, this happens by default: older turns are "summarized into a compact entry", recent messages are kept intact, and it triggers when the session nears the limit. The full history is usually still saved. The agent just no longer sees it; the documentation's own words are that it "only changes what the model sees on the next turn".
That is the literal moment your Monday rule went missing. It is now one line in a summary somebody else wrote, or it is not in the summary at all. The same engineering guidance names the risk exactly: summarising too hard "can result in the loss of subtle but critical context whose importance only becomes apparent later." A rule like "never book without asking" is the definition of a detail whose importance only shows up later.
The summary explains the sudden drop. It does not explain the slow one, where the agent gets vaguer and less careful over a long conversation. Three independent research groups found the same thing from three directions.
Longer input, less consistent answers. Researchers at Chroma tested 18 models in July 2025 and found that answers grow less consistent as the input gets longer, even on tasks as simple as repeating a list of words back.
Clutter misleads. In the same study, a single piece of text that looked relevant but was not was enough to pull answers off course: "Even a single distractor reduces performance", and more of them made it worse. A crowded desk is not neutral; the wrong paper gets picked up.
Focused beats full. On a long-conversation memory test, every model did significantly better when it was given only the relevant part than when it was given the whole history of roughly 113,000 tokens with the irrelevant parts left in.
The middle is the blind spot. A Stanford-led study found that models recall what sits at the start or the end of what they are reading far better than what sits in the middle. Your instruction from the first message does not stay at the start. As the conversation grows past it, it drifts into the middle, which is exactly where it is easiest to miss.
What these studies are, and are not
The research measured tasks like writing, summarising and answering questions, in controlled tests and simulated conversations. None of it measured your exact chat, and none of it gives a number of messages after which forgetting starts. Read it as the direction things move in, not as a forecast.
The largest study of this, by researchers at Microsoft Research and Salesforce Research, ran more than 200,000 simulated conversations across 15 leading models. They gave each model the same request two ways: all in one message, or spread across a conversation the way people actually talk.
The interesting part is where the drop comes from. The models did not lose much skill. They lost consistency.
-16%
Best-case ability
what the model can do on a good run
+112%
Unreliability
the gap between its good runs and its bad ones
That matches what you saw. Your agent can still follow the rule; it has stopped following it reliably. The researchers also found why correcting it rarely helps: models make assumptions early and lean on them, and "when LLMs take a wrong turn in a conversation, they get lost and do not recover." Every correction you add is one more paper on a desk that is already the problem.
Their advice to users is two lines long, and it is the most useful thing in the paper: if you can, try again in a new conversation, and consolidate before you retry, so the new conversation starts with everything that matters in one message.
None of this needs a setting you do not understand. It needs a routine.
New topic, new conversation. A chat that has planned a trip, fixed your CV and discussed your health is a crowded desk.
Your agent has a place for instructions that apply every time, usually called its instructions, profile or memory. Rules that must always hold go there, not in the first message of a chat.
Ask it for a handoff note: what we are doing, what we decided, what it would get wrong by guessing, and the next step. Read the note, fix anything wrong, and start a new conversation with it.
When one instruction really counts, repeat it at the end of your message, where it is least likely to be missed.
Paste the paragraph that matters, not the whole document. Ask it to read one page, not ten.
Each one takes room on the desk before you say anything.
Ask the agent to keep a short list of what you decided as you go, and to read it back at the start of the next session.
The handoff note in step 3 is the same practice the June 2026 how-to recommends, and it is the researchers' "consolidate before retrying" in plain words.
What it fixes
What it costs you
A new conversation per job
Clutter and early wrong turns
A new conversation per job
Seconds
A handoff note when it drifts
Rules lost in a summary, a conversation gone off course
A handoff note when it drifts
About a minute
Standing instructions
Rules forgotten between conversations
Standing instructions
Ten minutes, once
Removing unused plugins
Room taken before you type
Removing unused plugins
Ten minutes
A decisions list the agent keeps
Long projects that span many sessions
A decisions list the agent keeps
A habit
A host that saves notes before it summarises
The moment of forgetting itself
A host that saves notes before it summarises
A setup choice; check your host's documentation
A new conversation per job
What it fixes
Clutter and early wrong turns
What it costs you
Seconds
A handoff note when it drifts
What it fixes
Rules lost in a summary, a conversation gone off course
What it costs you
About a minute
Standing instructions
What it fixes
Rules forgotten between conversations
What it costs you
Ten minutes, once
Removing unused plugins
What it fixes
Room taken before you type
What it costs you
Ten minutes
A decisions list the agent keeps
What it fixes
Long projects that span many sessions
What it costs you
A habit
A host that saves notes before it summarises
What it fixes
The moment of forgetting itself
What it costs you
A setup choice; check your host's documentation
Be clear about what none of these do. They do not make the desk bigger. They keep what matters on it, and keep what does not off it. The last row is worth knowing about: some hosts remind the agent to write important notes to a file before they summarise, which is the software doing step 7 for you.
Look at the checklist again and one idea runs through it: what is written down survives, and what is only said in passing does not. Standing instructions, the handoff note and the decisions list all move something out of the conversation and onto paper the agent can re-read.
Nobody should have to hold all of this in their head. The point of a well-built plugin is that it does this for you, so you stop thinking about it. SIL, the shopping plugin we make, is built on the same idea: you state what you want once, as a spec, and the spec is written down and stays with you. The agent holds every candidate against that spec, so what you asked for does not depend on the conversation remembering it. If you are still setting up, do you need to know how to code to run a personal AI is the place to start.
SIL is the shopping plugin your agent loads: state the spec, hold every candidate against it, judge the result against it. Your specs are written down and live with you. It runs on OpenClaw, free and open source.
Because everything it works with has to fit in one limited space at the same time: your instructions, the conversation, the pages it read and what your plugins bring along. When that space fills, the host summarises the older part of the conversation and carries on from the summary, and details can fall out. Before it fills, a crowded space already makes the agent less consistent. Put rules that must always hold in its standing instructions, and start a new conversation for a new job.
For the 39% drop, the 16% and 112% split, and the advice to start again. Philippe Laban, Hiroaki Hayashi, Yingbo Zhou and Jennifer Neville (Microsoft Research and Salesforce Research), LLMs Get Lost In Multi-Turn Conversation, May 2025, ICLR 2026. 15 models, more than 200,000 simulated conversations, six task types.
For longer input and distractors reducing reliability, and focused beating full. Kelly Hong, Anton Troynikov and Jeff Huber (Chroma), Context Rot: How Increasing Input Tokens Impacts LLM Performance, 14 July 2025. 18 models.
For the middle of a long input being the weakest. Nelson F. Liu and others (Stanford University and co-authors), Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, 2023.
For context as a finite budget and the risk of summarising too hard. Anthropic Applied AI team, Effective context engineering for AI agents, 29 September 2025.
For summarising being on by default in a personal-AI host, and notes saved before it. OpenClaw documentation, Compaction, retrieved 29 September 2026.
For the handoff-note practice, and evidence that people ask this. Ben Patterson, PCWorld, a how-to on assistants forgetting things in long threads, 18 June 2026; Dust documentation, Why doesn't the agent remember what I said earlier in a conversation?.
Everything your agent works with has to sit in one space at the same time: your rules, the conversation, every page it read, every plugin you installed. The space has an edge.
When the space fills, the host summarises the older part of the conversation and carries on from the summary. A rule you gave on Monday can fall out of that summary.
Long before it fills, more clutter means less consistent answers. In one large study the same request spread over a conversation came out 39% worse.
One job per conversation, standing rules where the agent re-reads them, and a fresh start with a handoff note when it drifts.