Architecture32 min read

The Complete Agent Memory Stack

ClawQL's five-layer memory architecture — OKF vault, vector recall, PageIndex, CodeGraph, Onyx — and how they compose into persistent, auditable, sovereign agent intelligence.

ClawQL’s five-layer memory architecture — OKF vault, vector recall, PageIndex, CodeGraph, Onyx — and how they compose into persistent, auditable, sovereign agent intelligence.

This pairs with Git-native agent memory (vault backend, auto-commit, R2/Arweave durability), the Enterprise Ontology (entity instances as vault knowledge), and twelve layers of LLM cost. OKF serialization: docs/memory/okf.md. Token-efficiency layers: docs.clawql.com/architecture/token-efficiency.

Format note: the vault ships OKF-compatible Markdown by default — memory_ingest writes .md with the required type field. .cqk is the ClawQL Knowledge extension for the same frontmatter contract (ADR 0010). Examples below use .cqk; treat .md with the same fields as the Day-1 path.


The Tax Nobody Measures

Every AI coding session opens with the same overhead. The developer types something like “continue working on the authentication refactor.” The agent asks what the authentication refactor is.

It has no idea. The decision to use JWT over sessions last week is gone. The three hours spent benchmarking argon2 against bcrypt — gone. The specific parseConfig function that keeps failing during token validation and the reason the team switched to a different approach last Tuesday — all of it gone.

Most coding agents spend 60–70% of their tokens rediscovering code structure on every task. Ask Claude Code to “refactor parseConfig to support YAML” and it reads 5–15 files, runs grep on a few symbols, traces imports, then writes the change. That whole exploration happens again the next time, and the time after that. There’s no persistent record of what it already understood about the codebase.

The token cost is real and measurable. The productivity cost of re-establishing context at the start of every session is harder to put a number on but compounds faster. Organizational knowledge — the reasoning behind decisions, the approaches that were tried and rejected, the constraints that ruled out certain directions — doesn’t survive the context window boundary at all.

This isn’t a model failure. Models demonstrate within-session memory just fine. The problem is architectural: nothing persists between sessions.


What Agents Actually Need to Remember

“Memory” is a loose word that covers at least five distinct categories, each requiring different storage and retrieval mechanisms.

Codebase structure — which files exist, how they relate, what functions are defined where, how routes connect to handlers. Pre-indexed structure that the agent should pull from a graph rather than re-derive by reading files.

Decisions — the reasoning behind choices from previous sessions. Why JWT instead of sessions. Why argon2 instead of bcrypt. Why the current architecture rather than the three alternatives the team considered. This is narrative knowledge that benefits from explicit capture and structured recall.

Domain entities — typed, relationship-aware facts about the business: which customers hold which contracts, which assets carry which maintenance histories. Queryable against a schema. See the Enterprise Ontology post for the typed entity layer that makes this tractable.

Session context — what was in progress at the end of the last session, what was blocked, what the next step was. Temporal knowledge that bridges session boundaries.

Team knowledge — architectural decisions made by other engineers, debugging patterns that worked on similar problems, domain expertise accumulated over years. Shared across the swarm, not siloed per developer.

Standard AI tooling handles zero of these persistently. ClawQL’s memory stack addresses all five.


The Five-Layer Stack

Layer 5: Onyx — Semantic enterprise knowledge (40+ connectors, hybrid recall)
Layer 4: CodeGraph — Pre-indexed codebase structure
Layer 3: PageIndex — Embedding-free hierarchical retrieval
Layer 2: Vector Recall — Semantic similarity search over vault entries
Layer 1: OKF Vault (.cqk / OKF .md) — Durable, auditable, structured knowledge

Each layer builds on the one below. When memory_recall fires, it queries across active layers simultaneously, reranks by relevance and confidence, and returns a unified list. The agent doesn’t need to know which layer produced a given answer.


Layer 1: The OKF Vault

Every memory entry ClawQL stores is OKF-compatible: Markdown body, YAML frontmatter, file path as concept identity, wikilinks between concepts. The .cqk extension is the ClawQL Knowledge promotion of the same format, carrying additional provenance fields.

The Open Knowledge Format establishes just enough convention to make different producers and consumers interoperable. Required type field. Optional resource, description, tags. Catalog in index.md. Changelog in log.md. ClawQL adds worm_ref, correlation_id, agent_id, verdict, and confidence_score — fields that tie every entry to the WORM audit trail.

---
type: decision
title: "Authentication: JWT over sessions"
description: "JWT chosen for stateless auth rather than server-side sessions"
tags: auth, jwt, security, performance
timestamp: 2026-07-14T14:32:00Z
confidence_score: 0.94
agent_id: agent-daniel-dev-01
correlation_id: sess-8821-tool-047
worm_ref: sha256:a1b2c3d4e5f6...
verdict: passed
session_id: sess-8821
project: clawql-auth-refactor
---

## Decision

JWT over server-side sessions for the authentication layer.

## Reasoning

Server-side sessions require a shared session store that becomes a single point of failure
and complicates horizontal scaling. JWT tokens are stateless — the server validates
cryptographically without any external lookup.

## Constraints

Mobile clients can't use cookies reliably, so the JWT lives in the Authorization header.
Token invalidation before expiry needs a blocklist. That's acceptable for logout (rare)
versus the complexity of full session management.

## Alternatives Considered

Server-side sessions: requires Redis cluster, adds latency per request.
OAuth2 delegation: too heavy for internal API auth.
API keys only: no expiry mechanism, rotation complexity.

## Related

[[Authentication: argon2 over bcrypt]]
[[Authentication: 15-minute access token TTL]]

The wikilinks create a traversable knowledge graph. memory_recall can follow them the way a human would navigate a wiki — start at one decision, follow links to related ones, build a coherent picture of a decision cluster rather than a single isolated answer.

The worm_ref field links the knowledge entry to its WORM audit record, which captures the session ID, agent ID, full prompt context, and Evaluator verdict at the moment of ingestion. The entry is auditable: you can prove when it was created, by which agent, during which session.

Vault Directory Structure

~/.ClawQL/
  vault/
    index.md                      # OKF catalog — first thing read on recall
    log.md                        # append-only changelog of every ingest

    decisions/
      auth-jwt-over-sessions.cqk
      auth-argon2-over-bcrypt.cqk

    context/                      # type: context — cross-session bridges
      session-2026-07-14-end.cqk

    entity_instances/             # type: entity_instance
      contracts/
        contract-abc123.cqk

    task_results/
      weekly-summary-2026-07-14.cqk

    errors/                       # failure patterns + solutions
      jwt-validation-timeout.cqk

    patterns/
      retry-with-exponential-backoff.cqk

  ontology/                       # schema definitions — Git-tracked separately
    entities/
      Contract.cqe

The vault lives in S3/R2. Schemas (.cqe, .cqm) live in Git. These have different lifecycles: schemas change through PRs and deployments, knowledge accumulates continuously and should never require a deployment to persist. clawql sync handles R2 push/pull across edge nodes.

The OKF Index

The index.md is what makes the vault scale past a few dozen entries:

# ClawQL Agent Memory Vault

## Decisions (47 entries)

- [[auth-jwt-over-sessions]] — Authentication token strategy
- [[arch-effect-ts-migration]] — TypeScript error handling architecture
  ...

## Errors (23 entries)

- [[jwt-validation-timeout]] — JWT validation latency spike — RESOLVED
  ...

memory_recall reads index.md first. Scanning the catalog costs about 500 tokens. Loading the index plus the five most relevant full entries costs about 3,000 tokens. Loading every entry in a large vault would cost 50,000–200,000 tokens. The index-first pattern is what keeps recall economical as the vault grows.

The log.md

An append-only record of every memory_ingest event:

## 2026-07-14T14:35:22Z — sess-8821

**Ingested:** decision/auth-jwt-over-sessions
**Agent:** agent-daniel-dev-01
**Summary:** JWT chosen over sessions. Reasoning: stateless, horizontal scaling.
**WORM:** sha256:a1b2c3d4...

At the start of a new session, memory_recall reads log.md filtered to the last N days and the current project. The agent sees what was worked on recently without loading any full entries — the fastest possible session continuity mechanism.


Layer 2: Vector Recall

When memory_ingest creates an entry, it also generates an embedding vector and stores it alongside the file:

vault/
  decisions/
    auth-jwt-over-sessions.cqk
    auth-jwt-over-sessions.vec    # Float32Array, 1536 dims
  indexes/
    memory.db                     # SQLite — FTS5 + pgvector-compatible index

On edge gateways, this is a local SQLite file with zero network latency. In the Virtual Gateway, pgvector in Postgres handles it across the team.

The recall pipeline runs in six steps:

  1. Index survey — read index.md, identify candidate categories from the query. ~200 tokens.
  2. FTS — SQLite FTS5 on entry titles, tags, and descriptions. Fast, no cost, catches exact matches.
  3. Vector similarity — embed the query, cosine search the index, return top-K by semantic similarity.
  4. Wikilink traversal — follow links from high-scoring entries to related entries. The JWT decision links to argon2, which links to the token TTL. The agent gets a cluster, not a single point.
  5. Reranking — cross-encoder over the merged candidate set.
  6. Confidence-filtered loading — load full content only for entries above threshold.

Total pipeline cost: 500–1,500 tokens. Loading every entry in a medium-sized vault: 20,000–100,000 tokens.

Two design decisions worth naming:

Embedding failures fail open. If the embedding service is unavailable, the request proceeds to live inference. The cache is an optimization, never a gate.

Cache hits are flagged with cache_hit: true. The --exclude-cache-hits flag on Flywheel export removes them from training data — training on cached responses instead of model outputs degrades training quality.


Layer 3: PageIndex

PageIndex is a tree-based retrieval approach using hierarchical categories and LLM reasoning rather than embedding similarity. No API key, no local embedding model, no vector database required.

Before embeddings, information retrieval was hierarchical. Libraries use Dewey Decimal. File systems use directories. A wiki uses categories. Hierarchy is cognitively natural and often more precise than similarity search for queries that are inherently categorical.

How It Works

ClawQL maintains a category tree over the vault:

knowledge/
  technical/
    authentication/
      jwt/
        auth-jwt-over-sessions.cqk
        auth-jwt-15-minute-ttl.cqk
      passwords/
        auth-argon2-over-bcrypt.cqk
    architecture/
      effect-ts/
        arch-effect-ts-migration.cqk
  operational/
    incidents/
    runbooks/
  domain/
    contracts/
    customers/

When memory_ingest runs, a Frugal-tier LLM classifies the entry into the appropriate path. The assignment is stored in frontmatter (page_index_path, page_index_categories). On recall with backend: page_index, the same LLM classifies the query, walks the tree to the matching path, and ranks the candidates.

Total cost: one cheap classification call plus one ranking call over a small candidate set.

PageIndex suits: precise categorical queries (“what decisions did we make about authentication?”), environments without embedding infrastructure, large vaults where vector similarity gets noisy.

Vector Recall suits: semantic queries that don’t map cleanly to a category, cross-domain queries spanning multiple branches, novel queries where the right category isn’t obvious.

ClawQL runs both simultaneously and reranks the merged results.

Lesson: embeddings are optional infrastructure. PageIndex + FTS still recall when the embedding service is down.


Layer 4: CodeGraph

CodeGraph hit GitHub #2 trending on May 23, 2026 with a local knowledge graph aimed at cutting agent token spend on code exploration. Reported compression on large codebases lands in the same order of magnitude regardless of the codebase: far fewer tool calls, far fewer tokens, faster edits, because structure is pre-indexed rather than rediscovered.

ClawQL treats CodeGraph as a first-class memory layer because codebase structure is the most frequently needed memory for software development agents, and re-discovering it on every session is the largest single source of avoidable token waste in development workflows.

What Gets Indexed

Language support: TypeScript, JavaScript, Python, Go, Rust, Java, C#, PHP, Ruby, C/C++, Swift, Kotlin, Dart, Lua, Svelte, and others. Route detection for Django, Flask, FastAPI, Express, NestJS, Laravel, Rails, Spring, Gin, Axum, ASP.NET, Vapor, React Router, SvelteKit, and others.

The index captures: symbol index (functions, classes, interfaces, types, variables with file paths and line numbers), import graph, call graph, route map, type hierarchy, and a change index weighted by recency.

Auto-sync uses native filesystem events — FSEvents on macOS, inotify on Linux — with debounced updates. The index stays current as code changes. A renamed function propagates in seconds.

For nested repositories, CodeGraph descends into any nested repo with a real working tree and indexes it as an embedded repository, recursively. vendor/ and node_modules/ stay excluded. For ClawQL’s monorepo, this means the import graph across packages is correct.

MCP Tools

// Instead of: read_file + grep + read_file...
cg_find_symbol('parseConfig');
// Returns: { file, line, signature, docstring }

// Instead of: ls + multiple file reads...
cg_explain_module('packages/clawql-inference');
// Returns: summary, exported symbols, dependencies, recent changes

// Instead of: grep across all files...
cg_find_callers('processDocument');
// Returns: every call site with file + line + context

// Instead of: manually tracing imports...
cg_dependency_path('clawql-payments', 'clawql-core');
// Returns: exact import chain between modules

// Instead of: reading route files...
cg_find_route('/v1/chat/completions');
// Returns: { handler, file, middleware }

Each call replaces 3–15 tool calls. The saving is structural.

Knowledge Entry Generation

When an agent makes a significant code change, ClawQL generates a type: code_change entry capturing what changed, the reasoning from the session context, and the CodeGraph diff:

---
type: code_change
title: 'parseConfig: added YAML support'
affected_symbols: [parseConfig, ConfigOptions, ConfigParser]
affected_files: [src/config/parser.ts, src/config/types.ts]
breaking_change: false
correlation_id: sess-8821-tool-089
worm_ref: sha256:c3d4e5f6...
---

Future agents querying “why does parseConfig support YAML?” find this entry through both vault recall and the CodeGraph symbol index.


Layer 5: Onyx

The first four layers handle agent-generated memory. Onyx handles the organizational knowledge that already exists in Confluence, Notion, Slack, Google Drive, GitHub, Jira, and dozens of other sources.

Onyx is an open-source enterprise search platform with 40+ pre-built connectors, streaming indexing, hybrid keyword+vector search, permission-aware retrieval (respects the source system’s access model), and citation-backed results that tell the agent not just what was found but where and why.

The distinction from the OKF vault: the vault captures what agents learned. Onyx surfaces what the organization already knew. An agent working on the authentication refactor might query both:

  • OKF vault: “what authentication decisions did we make in previous sessions?” → typed decision entries
  • Onyx: “what does the security documentation say about token handling?” → Confluence pages, Slack threads, GitHub issues
knowledge_search('JWT token expiry best practices', {
  sources: ['confluence', 'slack', 'github'],
  date_from: '2025-01-01',
  max_results: 10,
  permission_level: 'agent_atrclaims',
});
// Returns: citation-backed results respecting ATRClaims authorization

How the Layers Compose

memory_recall runs all active layers in parallel and merges:

const memoryRecall = async (query: string, context: RecallContext) => {
  const [indexSurvey, vectorResults, pageIndexResults, codeGraphResults, onyxResults] =
    await Promise.all([
      vault.surveyIndex(query),
      vectorRecall.search(query, { topK: 10 }),
      pageIndex.search(query, { topK: 10 }),
      isCodeQuery(query) ? codeGraph.search(query) : Promise.resolve([]),
      onyx.search(query, context.atrclaims),
    ]);

  const vaultEntries = await vault.loadEntries([
    ...vectorResults.filter((r) => r.score > 0.75),
    ...pageIndexResults.filter((r) => r.score > 0.7),
  ]);

  const reranked = await reranker.rank(query, [
    ...vaultEntries,
    ...codeGraphResults,
    ...onyxResults,
  ]);

  return {
    entries: reranked.slice(0, context.maxResults ?? 10),
    index_summary: indexSurvey,
    token_cost: estimateTokens(reranked),
  };
};

The agent receives a unified ranked list. Source layer is recorded per entry for attribution. The agent doesn’t need to know whether an answer came from the vault, CodeGraph, or Onyx.


Team Memory

Individual developer memory is valuable. When every developer’s agent shares knowledge through the Virtual Gateway, the swarm collectively knows more than any individual agent.

Namespace Model

s3://clawql-org-vault/
  personal/{developer_id}/vault/     # private to the developer
  team/{team_id}/vault/              # shared across the team, ATRClaims scoped
  org/vault/                         # shared across the organization

memory_recall queries all three namespaces the agent is authorized for. memory_ingest routes to the appropriate namespace based on knowledge type. Personal decisions stay personal. Architectural decisions shared with the team. Organizational standards go to the org namespace.

In the Agentic Fabric, the Virtual Gateway’s NATS JetStream propagates recall requests across the swarm when a local cache misses. Team vault answers don’t require every edge node to hold the full team vault locally.


Connection to the Token Efficiency Stack

The memory stack connects to multiple layers in the twelve-layer efficiency architecture:

Layer 1 (Context bloat): CodeGraph MCP tools replace file reads. The agent calls cg_find_symbol instead of reading 15 files.

Layer 3 (Redundant static context): The vault index.md is stable between sessions. The system prefix including the index survey is cacheable at the provider layer.

Layer 5 (Semantic cache): Recall queries are themselves cached. The same query in a different session returns the cached result.

Layer 6 (Undistilled history): When session history grows long, ClawQL distills it into a type: context vault entry. The transcript is gone; the semantic content persists.

Layer 7 (Cross-session rediscovery): The next session recalls the context entry rather than re-establishing project state from scratch.

Layer 12 (Flywheel): Every memory_ingest with verdict: passed is a training example. After enough Flywheel cycles, the Frugal-tier model gets better at predicting which vault entries are relevant for which queries.


WORM Audit

Every memory operation writes a WORM entry:

'MEMORY_INGEST'; // new entry created
'MEMORY_RECALL'; // query + entries returned
'MEMORY_CACHE_HIT'; // recall returned cached result
'MEMORY_UPDATE'; // existing entry modified
'MEMORY_ARCHIVE'; // entry moved to cold storage
'MEMORY_SYNC_PUSH'; // entry synced to R2
'MEMORY_SYNC_PULL'; // entry pulled to edge
'CODEGRAPH_INDEX_UPDATE'; // CodeGraph index updated
'ONYX_SOURCE_INDEXED'; // Onyx connector indexed new content

Compliance queries that no other memory system supports:

“Show me every decision created by agent X in the last 90 days” — WORM query on MEMORY_INGEST filtered by agent_id and type: decision.

“Prove that HIPAA-sensitive patient data was not accessed by the agent during this session” — WORM query on MEMORY_RECALL events filtered by pii_classification: PHI. If no PHI entries appear in recall events, the agent didn’t access them.

The negative proof is what regulated industries need. Not just what the agent remembered, but what it provably didn’t access.


What Alternatives Are Missing

Plain Obsidian vault: The knowledge structure is right. Agent access is manual. No memory_ingest automation, no vector recall, no WORM audit, no team sync, no CodeGraph integration.

MemoryGraph and similar MCP servers: Graph-based relational memory works conceptually. Missing: OKF compatibility, CodeGraph integration, Onyx enterprise connectors, WORM trail, team sharing, hot/cold tiering, Flywheel export.

Claude Projects: Persistent context within Claude’s platform. Missing: portability to other tools, independent auditability, data ownership. The memories belong to Anthropic’s platform.

Naive RAG over documentation: A retrieval system for static documentation. Missing: structured knowledge format, cross-session capture, WORM audit, team sharing. New knowledge never enters the index.

Vector databases directly: Storage infrastructure. They’re Layer 2 of the five-layer stack, not the stack itself. Missing: OKF catalog for token-efficient index survey, PageIndex fallback, CodeGraph integration, Onyx connectors, WORM trail, provenance fields.


Getting Started

Day 1: OKF Vault

clawql memory init
export CLAWQL_MEMORY_BACKEND=r2
export CLAWQL_R2_BUCKET=your-org-vault
clawql inference serve --port 8080
# Connect IDE: http://localhost:8080/mcp
# "Remember that we chose JWT over sessions — stateless, horizontal scaling"

Week 1: CodeGraph

npm install -g codegraph
cd your-project && codegraph init
# Auto-syncs on file changes
# MCP tools: cg_find_symbol, cg_find_callers, cg_explain_module, cg_find_route

Week 2: Vector Recall

export CLAWQL_EMBEDDING_MODEL=text-embedding-3-small
export OPENAI_API_KEY=sk-...
# Or local: export CLAWQL_EMBEDDING_MODEL=nomic-embed-text
export CLAWQL_MEMORY_VECTOR_RECALL=1
export CLAWQL_MEMORY_CACHE_THRESHOLD=0.92

Month 1: Onyx

# Deploy Onyx (Docker or Helm), configure connectors
export CLAWQL_ONYX_URL=http://onyx:3000
export CLAWQL_ONYX_API_KEY=...

Month 2: Team Sync

clawql sync init --provider r2 --bucket org-vault
clawql sync pull
clawql sync push

The Moat

The memory stack compounds. The longer it runs, the more the vault knows. The richer the vault, the better each session. The better each session, the more high-quality data enters the Flywheel. The Flywheel produces a fine-tuned model that’s better at recall for this organization’s specific domain.

A competitor building an identical stack starting today begins from zero. Your stack has months or years of organizational knowledge, your codebase’s structure captured through CodeGraph change entries, your domain’s patterns, your team’s accumulated decisions. None of that transfers.

The WORM audit trail adds a compliance dimension that no other memory system provides: not just “we remember” but provable records of what was remembered, when, by whom, and against what quality standard. For regulated industries, that proof is the difference between production deployment and indefinite evaluation.


OKF vault serialization and memory_ingest / memory_recall behavior: docs/memory/okf.md. Git backend: Git-native agent memory. Enterprise Ontology: docs.clawql.com/architecture/enterprise-ontology.

About the author

Daniel Smith builds ClawQL, an agent operating system for token-efficient discovery and execution over APIs — with observability, hardened tool boundaries, and production routing for LLM workloads. He writes here about the systems problems behind shipping agents.