Imagine telling an agent that a service is disabled today. Six months later you ask about it, and the agent says the service is disabled. It remembered perfectly. It is also wrong, because at some point since then you switched the service back on and never mentioned it.
Technically, the memory worked. Operationally, it failed. Most of what I think about agent memory comes from that gap.
Remembering everything is the easy part
Storage is cheap. Saving every conversation costs almost nothing, and the usual recipe is well known: keep the conversations, embed them, retrieve whatever looks similar to the next question, paste it into the prompt. It produces demos that feel like memory.
My problem with it is what it leaves out. Everything becomes the same kind of thing: a chunk of text with a vector. “The CrowPanel shows monitor status” is a fact about a design. “The broker was unreachable earlier” is an event. “Ask me before touching the database” is a rule. “I’m working on the firmware this week” is state that will be false soon. Similarity search can’t tell which of those should still be steering a decision, because it only measures how close they are to the question.
Retrieval is not memory
Similarity tells you what is about the same topic. It doesn’t tell you what is still true, how sure anyone was, or whether one thing should override another. At minimum I want to keep these apart: conversation, event, evidence, current state, rule, preference and historical fact.
I ran into a version of this while testing how a planner chooses which sources Pixel should read. There are three sources that all concern the same service: its current state, a log of events, and sampled telemetry over a period. In the extended test the planner returned valid JSON in 30 of 30 cases and still picked the wrong evidence in 11 of them. It kept reaching for the event log when the question needed recorded runtime transitions. An event log can’t tell you whether MQTT has been stable this week, and a snapshot of now can’t tell you how the mini PC behaved today. Same topic, different kind of past. A similarity score doesn’t see that difference. The source-selection benchmark in AI.003 records the failures.
State changes, and remembered state goes stale
The homelab makes this concrete. A container can be configured but not enabled, enabled but unavailable, or deliberately disabled. Those are different states, and an observation of any of them has a shelf life. “Unavailable” from this morning is a statement about this morning. “Disabled on purpose” may stay true for months, until I change my mind. A rule like “ask before touching the database” isn’t an observation at all; it is supposed to outlive all of them.
If memory flattens these into the same kind of record, the most similar old record wins, and nothing in the system can say “that was true then.”
Provenance: why do I believe this?
The second thing I want from a memory is a reason. In Pixel this is the idea of Evidence: a belief should point at the observation that supports it and say how old that observation is. If Pixel tells me a device is unavailable, I want to be able to ask “since when, according to what?” and get an answer that isn’t a shrug.
In the code, Evidence is deliberately boring. Its whole shape fits in four fields:
EvidenceRef
evidence_class operator_explicit | operator_correction | ...
source_kind activity_event
source_id the recorded event it points at
occurred_at when it happened
That is a pointer to something that was recorded, nothing more. Separately, claims that rest on evidence carry a level of confidence, from “declared” through “tool observed” and “runtime verified” up to “operator confirmed”. There is no free text anywhere in it, and a model’s own inference is excluded by construction, so it can never count as evidence for anything.
It also changes how the model should phrase things. “The door node hasn’t reported recently” is a statement with a source. “The door node is offline” is a conclusion that may have outlived its evidence.
Different things deserve different lifetimes
So Pixel’s memory isn’t one pile. The pieces have different jobs, and therefore different lifetimes.
Evidence is the pointer I just described: what was observed, from where, and when. An OperatorRule is something I told Pixel to do or not to do. Its text can’t be edited or deleted. Changing a rule means writing a new one, and the old one is marked as superseded, so there is always a trail. Rules are matched to a situation by deterministic scope, never by embeddings or by the model, and only a few can be always-on.
The rest are easier to compare side by side.
| Piece | What it holds | What ends its life |
|---|---|---|
| ProjectState | A short snapshot I write about a project: focus, status, next action, blockers. The status carries an evidence level. | Nothing deletes it. After 30 days it is labelled stale, and current data wins whenever they disagree. |
| PixelSession | The working context of a conversation, as structured fields rather than transcripts. Memory of the Agent only. | Idle expiry, or a restart of the Agent. |
| Policy | Who may see what: room-audible surfaces such as the CrowPanel and the Cardputer neither read nor write session memory. | It doesn’t expire. It changes when I change it. |
Underneath all of these sits operational history: a record of what happened, kept separately from the current truth. The diagram summarises the intended roles and lifetimes. It is not a claim that every memory feature is enabled.

Deployment boundaries: automatic rule learning is switched off. Separately, the deterministic composition path for operational history is still awaiting its live checkpoint, as described in the AI.003 planner experiment. Defined memory roles and lifetimes should not be read as proof that every path has been validated in use.
Forgetting is a feature
Never forgetting is not intelligence. A system that keeps everything at equal weight can’t tell last week from last year.
Good memory needs ways to let go: expiry, supersession (a newer observation replaces an older one instead of sitting beside it), lower confidence over time, and plain deletion. A ProjectState older than 30 days is labelled stale. It is labelled, not hidden: it still goes to the model, marked as possibly out of date, and the prompt says that current data and recent history win over it whenever they disagree. A session expires on its own. A rule candidate that nobody approves expires after 30 days. The number is blunt, and I still don’t know whether a single window is even the right shape. A service’s state goes stale in minutes. A rule I wrote may stay true for years. Different lifetimes again.
There is a quieter version of the same principle in the CrowPanel firmware. It parses the status summaries it receives strictly, and it never shows stale data as if it were live. If you can’t vouch for how fresh something is, say so instead of rendering it confidently.
The model shouldn’t be the database
Asking the language model to be the memory, to carry everything in context and sort it out on the fly, gets the division of labour backwards. A model is good at reading a small amount of material and reasoning about it. It is a poor store with expiry rules. I would rather select state outside the model, with provenance and age attached, and hand over only what is relevant to this question.
The model receives selected context with its provenance and age attached. It does not get to rewrite a verified fact just because a different answer sounds more plausible.
What I’m actually building
Pixel’s memory isn’t solved, and it certainly doesn’t have perfect long-term recall. It has some structure and a few opinions. Keep different kinds of memory apart. Attach where something came from and when. Let things expire or be replaced. Keep the model out of the storage business.
There is also a part I haven’t earned yet: automatic rule learning is implemented but switched off. Pixel can notice when I correct an answer and propose a rule, but a proposal has no effect until I approve it. I would rather leave that path inert than trust it before I’ve watched it work.
What I notice is that this makes Pixel more useful by making it know less. An agent holding fewer things, each with a source and an age, can answer “is that service disabled?” with “as of this morning’s check, yes.” That sounds worse than a confident “yes”. It is a much better answer.
The surrounding system is described on the Afterlab project page, and the operational view that this memory feeds is the subject of Why I built a tiny mission control for my homelab.