Harness Engineering · Chapter 11

Memory, Compaction, and Continuity

How runtime memory preserves useful continuity without turning summaries, retrieved notes, or prior model output into authority.

Memory is governed continuity

Memory lets a later decision benefit from earlier work. It can preserve user-confirmed facts, outcomes, preferences, unresolved commitments, environment knowledge, or procedural lessons. The engineering challenge is not merely storing material; it is deciding what may be captured, believed, retrieved, corrected, forgotten, and allowed to influence.

CaptureQualifyStoreRetrieveCorrect or forget
Compaction produces a bounded continuity input; it does not turn remembered material into authority.

Keep canonical records separate from memory projections. A task event, user statement, or provider receipt may be authoritative within its contract. A summary or embedding-derived result is a lossy view over such material and should retain its provenance.

Capture does not create belief

Candidate memories can come from explicit user input, completed work, observations, or model proposals. Assign each a source, subject, timestamp, scope, confidence or review status, retention class, and correction path. User statements should outrank behavioral inference about the user.

Sensitive material needs an explicit purpose and permission boundary. A system should remain useful without coercing optional context, and deleting a canonical record should trigger deletion or rebuilding of its projections where policy requires it.

Retrieve for a decision

Similarity is only one retrieval signal. Select memory for the choice at hand: current objective, identity, time, environment, authority, recency, contradiction status, and expected utility all matter. Return a small attributed packet rather than an undifferentiated history.

Retrieved memory is input, not command. A note that once said “publish automatically” cannot expand the authority of a new run. Current admission and locked policy remain in control.

Context database versus memory engine

A context database and a memory engine can both make old material available to a later run, but they own different seams. At the inspected OpenViking commit, resources, memories, and skills live in an addressable viking:// hierarchy. Its L0 abstracts and L1 overviews are derived directory sidecars with explicit freshness and sampling metadata; L2 remains source detail. Retrieval plans typed queries and recursively explores directories, while a separate context assembler applies tier, category, token, peer, and served-entry deduplication policy. Account and user identity determine the visible namespace; a peer only narrows content inside a user boundary.

At the inspected AgentMemory commit, hooks and MCP turn host activity into sessions, observations, memories, summaries, and lessons. Internal recall can fuse BM25, optional embeddings, and graph matches; mem::context can assemble a bounded packet, while host injection is opt-in and disabled by default. Project is an explicit or basename-derived selector, and agent recall is shared by default unless isolated mode is enabled. Unscoped legacy records can remain visible across project filters. The common MCP memory_recall tool does not expose or forward project, so project-aware storage and internal search do not imply project-safe recall on every surface.

| Seam | Context-database question | Memory-engine question | |---|---|---| | Address | Can the source and its projections be browsed through stable paths? | Can an observation be traced to its session, project, agent, and origin? | | Capture | Who may add a resource, session, memory, skill, or compiled artifact? | Which host events become observations or durable memories, and under what consent? | | Scope | Which authenticated account, user, or peer may read this URI? | Which project or agent selector is applied, and can unscoped records still pass? | | Correction | Which source record or memory diff supersedes a derived sidecar or extracted memory? | Which version remains canonical, and were stale versions removed from every retrieval index? | | Deletion | Which source files, indexes, queues, and compiled projections were removed? | Which KV records, indexes, exports, snapshots, audits, and connected hosts were purged or retained? |

A scope filter is not identity proof. In either design, keep authenticated subject, admitted scope, storage filter, returned context, accepted belief, and current effect authority separate. OpenViking's first-party benchmarks and AgentMemory's maintained evaluation results were not rerun for this chapter; they are not evidence that either system closes those wider contracts.

Boundary exercise

An agent automatically records “the user prefers weekly email summaries” in a repository named client. The user later corrects it to “never email; draft only.” A second, unrelated repository is also named client. Identify the canonical correction, the selectors and authenticated scope required before recall, every stale projection that must stop winning retrieval, and the evidence needed before claiming deletion. Then decide whether either memory may authorize sending an email. The safe answer should not depend on retrieval score.

Compact without manufacturing certainty

Compaction reduces working history when the context window, latency, or attention budget becomes constrained. A useful compacted packet preserves objective, locked constraints, durable decisions, confirmed progress, unresolved ambiguity, open effects, evidence pointers, and the next safe action.

Record what was omitted and the source boundary of the summary. Contradictions should remain visible. If detail can be retrieved later, store a pointer instead of a confident paraphrase. The goal is continuity, not a polished story.

Correct and forget as first-class operations

New evidence may supersede a memory, narrow its scope, or show it was wrong. Append the correction, link the affected projections, and prevent superseded material from silently winning retrieval. Forgetting may mean expiry, user-directed deletion, policy retention, compaction, or removal of a derived view; name which operation occurred.

Evaluate the lifecycle, not just retrieval relevance. Test capture precision, attribution, contradiction handling, decision usefulness, deletion propagation, authority non-amplification, and behavior after compaction or restart.

Failure boundary

Memory fails when inference becomes identity, summaries erase uncertainty, retrieval leaks across users or tasks, old preferences override current instructions, deletion leaves searchable projections, or compacted context omits the one unresolved effect that made retry unsafe.

Retrieval check

A prior summary says an email was sent, but the durable run journal says delivery is unresolved and the user has since revoked messaging access. What may be retrieved, what outranks what, and what is the next safe transition? Identify the correction and forgetting operations required after reconciliation.

Sources and further reading