OPENABBY · TECHNICAL BRIEF
№ 01 — THE MEMORY SYSTEM
Agent Architecture

How Abby remembers.

OpenAbby agents keep two memories in sync — one written for humans, one built for machines — stitched together by a rescue hook that saves knowledge from the context window before it's compacted away, and topped with a learning loop that turns corrections into behavior.

markdown workspacevector engineconfidence decaysleep-like consolidationknowledge graphlearning loop
01

Two memories, one mind — the dual-representation design

Most agent frameworks pick a side: either flat files the model edits, or an opaque vector store. OpenAbby runs both, deliberately. The file layer is auditable — you can open it and read what your agent believes. The semantic layer is searchable by meaning. A compaction-time flush keeps them in agreement.

The memory architecture, at a glance
Layer 1 · Files

The Workspace

plain markdown · human-auditable
  • SOUL.md — persona & identity
  • MEMORY.md — curated long-term notes
  • memory/YYYY-MM-DD.md — daily journal
  • memory/user_profile.md — who the user is
  • hot-reloaded into the prompt on change
⇠ ⇢
compaction flush
context is never lost — it's demoted to disk
Layer 2 · Engine

Semantic Memory

SQLite + embeddings · machine-searchable
  • facts, preferences, skills — auto-extracted
  • confidence decay — with reinforcement
  • consolidation — merge, insight, promote
  • knowledge graph — entities & relations
  • optional shared tier for the whole fleet
02

The workspace — memory you can read

Every agent owns a directory of markdown. The agent edits it with three tools — memory_read, memory_write, memory_search — and a file watcher hot-reloads identity files straight into the system prompt the moment they change. Nothing about what the agent believes is hidden from its operator.

SOUL.md

The persona. Personalized per agent at provisioning time, so each deployment introduces itself with its own identity and address.

MEMORY.md

Curated long-term memory. Deliberately overwritten — not appended — so the agent must decide what deserves permanence.

memory/2026-08-07.md

The daily journal. Append-only, one file per day. This is where the compaction flush lands notes rescued from the context window.

memory/user_profile.md

Who the user is. Pre-seeded from a team roster at provisioning — a freshly-created agent already knows its person on first contact.

USER.md · IDENTITY.md · AGENTS.md

Supporting bootstrap files, all watched and hot-reloaded. Editing a file is editing the agent's mind — no restart required.

humility check

Before the agent claims it doesn't know something, a humility pass greps the workspace first. "I don't know" is only allowed after actually looking.

03

The semantic engine — memory that behaves like memory

Underneath the files sits a vector store of extracted knowledge — and this is where the biology-inspired mechanics live. Memories here aren't rows that sit still; they age, reinforce, merge, and get promoted.

◐

Confidence decay, with reinforcementdecay engine

Every memory carries a confidence score with an exponential 30-day half-life. Each time a memory is touched again, its half-life stretches by 1.5× — so three reinforcements make it over three times more durable. Unused memories fade below threshold, get archived, and are eventually pruned. Different types age at different speeds: preferences outlive trivia.

☾

Sleep-like consolidationmaintenance pass

Periodically the engine does what brains do at night: it merges near-duplicates by embedding similarity, asks a model to derive higher-level insights from clusters of related memories, prunes the dead, and promotes frequently-reinforced, high-confidence items into core knowledge. The agent doesn't just store — it digests.

◫

Hierarchical retrievalcategory-first search

Recall isn't a flat scan over every vector. Search descends category-first — find the relevant topics, then the best memories within them — mirroring how human recall moves from theme to detail.

✳

Standing instructions as first-class citizenspreferences store

"Always answer in bullet points" is not the same kind of fact as "the user's dog is named Piper." An intent classifier spots durable instructions, stores them deduplicated by content hash, and injects the active set every single turn — no retrieval roulette for the things that must always hold.

⌘

A self-building knowledge graphentity extraction

A background pass extracts entities — people, organizations, projects, technologies — and typed relationships (works_at, uses, depends_on) from conversation, populating a graph nobody has to maintain by hand.

◱

Two storage tierspersonal + shared

Tier one is a per-agent SQLite database — private, local, zero-dependency. Tier two is an optional shared Postgres + pgvector store, so a fleet of agents can pool what they learn, with a Redis cache in front.

04

The turn lifecycle — remembering is automatic

The agent never has to remember to remember. Two hooks wrap every single exchange.

Before the turn

Recall

  • Embed the incoming message
  • Retrieve relevant memories, category-first
  • Load the user's active standing instructions
  • Inject it all into context — invisibly
After the turn

Absorb

  • Extract facts, preferences & skills from the exchange
  • Embed and store them — reinforcing what's already known
  • Classify intent; capture new standing instructions
  • Feed entities & relations into the knowledge graph
05

The compaction flush — nothing important dies in the context window

When context runs out, memory catches what falls.

Every long-running agent eventually fills its context window and must summarize old messages away. In most systems, detail is simply lost. In OpenAbby, a before-compact hook fires first: a cheap sidecar model — never the expensive primary brain — reads the messages about to be destroyed and writes the durable facts into the daily journal. Context isn't lost; it's demoted from RAM to disk, where memory search can still find it. The main model pays nothing for its own bookkeeping.

01Context window approaches its limitcompaction engine schedules a summary pass
02BeforeCompact hook interceptsthe doomed message chunk is handed off
03Sidecar model extracts durable notesa small, cheap model — capped output, fast
04Notes append to today's journalsearchable forever via memory_search
05Compaction proceeds safelythe context shrinks; the knowledge doesn't
06

The Pistis loop — memory of lessons, not just facts

Above both memory layers sits a separate system for learning from mistakes. Facts tell the agent what's true; lessons tell it how to behave. When a user corrects the agent — or something fails — that signal becomes policy.

Detect

Corrections, frustration, and failures are spotted as learning signals in the conversation stream.

Reflect

A reflection pass asks: what's the general lesson here — not just what went wrong this once?

Distill

The lesson is inferred and written down as a candidate policy snippet, tied to its evidence.

Apply

Approved lessons shape future behavior — the same mistake shouldn't need correcting twice.

07

What makes it different

✓

Dual representation, kept honest. Human-auditable markdown and machine-searchable vectors describing the same mind — synced by the compaction flush rather than drifting apart.

✓

Memories age like memories. Exponential decay, reinforcement-extended half-lives, archival, and pruning — instead of an ever-growing pile of stale embeddings.

✓

Consolidation generates, it doesn't just compress. The maintenance pass produces new higher-level insights from clusters of old memories.

✓

The brain never does the bookkeeping. All extraction, flushing, and graph-building runs on a cheap sidecar model — the primary model's tokens go to the user.

✓

Agents are born knowing their person. Roster-driven pre-seeding of user_profile.md means the first conversation is never a cold start.

✓

Facts, instructions, and lessons are three different things — stored, retrieved, and applied through three different mechanisms, because they fail differently when conflated.