FIG. 1 - R6 Updated 2026-06-20
Public scorecard / no fabricated scores

Memory leaderboard for inspectable agent memory.

A useful memory leaderboard should not guess hidden scores. Compare systems by published evidence, inspectability, forgetting, portability, and permission controls, then mark every missing measurement as unmeasured. This leaderboard is a public scorecard for evaluating Answer Engine, Mem0, Zep, Letta, and Cognee without inventing benchmark numbers.

Fig. 2 leaderboard / measured-vs-unmeasured
FIG. 2

Which memory system is best?

The best system depends on the control surface you need. Use this table to separate published evidence from unmeasured claims, then follow the source links before you shortlist a vendor.

Public memory leaderboard with unmeasured cells labeled
System Best fit Published eval evidenceAccuracy scoreInspectable retrievalForgetting / deletionPortabilityPermission-aware retrieval
Answer Engine Tenant-scoped content memory with source artifacts, library scope, MCP tools, CLI sync, and REST delivery. Teams that want source-aware memory through MCP, REST, and CLI with tenant-isolated retrieval and public roadmap labeling. Roadmap

The reproducible benchmark post is gated; this row does not publish a score.

Unmeasured

No public Answer Engine accuracy score is claimed here.

Source-aware now; trace viewer roadmap

Citations and source-aware retrieval ship; full live inspector is roadmap-labeled.

Delete handles and source lifecycle

Deletion posture is documented; deeper provable-forgetting audit is roadmap-labeled.

MCP, REST, and CLI

One memory layer can serve Claude Code, Cursor, Codex, Gemini CLI, Windsurf, and API clients.

Tenant and library scoped

Repository queries bind tenant_id and library scope before retrieval.

Mem0 Managed memory engine with add, search, update, delete, platform governance, and open-source options. Product teams that want a managed user-memory platform with SDK-first setup and hosted operational controls. Vendor-published memory evals

Public docs include a memory evaluation section; methodology is not normalized here.

Unmeasured

No same-method score is measured by Answer Engine in this leaderboard.

Audit and dashboard surfaces

Public docs describe platform audit and governance; retrieval-trace depth is not scored here.

Delete memory operation

Delete operations are documented; derived-artifact audit is not normalized here.

SDKs, API, MCP entry

Portable at integration level; export/migration depth is not scored here.

Workspace governance

Enterprise controls are documented; tenant retrieval SQL is not inspectable from docs.

Zep Temporal Context Graph and Context Lake for user, business, document, and JSON memory. Teams prioritizing temporal knowledge graph semantics and prompt-ready context from graph-derived memory. Vendor-published graph memory paper

Zep publishes graph-memory architecture and evaluation material; not normalized here.

Unmeasured

This leaderboard does not copy or compare a vendor-specific benchmark number.

Graph context visible by API/docs

Temporal graph structure is documented; answer-level trace inspection is not scored here.

Data management documented

Deletion and retention behavior should be verified against current account controls.

SDK/API, migration docs

Public docs include migration paths; MCP portability depends on integration choices.

Governed Context Lake

Governance is described; retrieval-layer tenant implementation is not scored here.

Letta Stateful agents with memory blocks, shared memory, archival memory, tools, and agent configuration files. Builders who want stateful agents with explicit memory blocks, archival memory, tools, and agent files. Evals docs exist

Letta documents eval tooling; this leaderboard does not run a comparable test.

Unmeasured

No same-method leaderboard score is published here.

Agent state and memory blocks

Memory state is explicit; retrieval-trace depth depends on implementation.

Mutable agent memory

Block and passage APIs exist; provable derived-artifact deletion is not scored.

API, tools, AgentFile

Portable through platform and files; client-neutral MCP memory surface varies by setup.

RBAC documented

Role-based access control is documented; vector-layer ACL implementation is not scored.

Cognee Open-source data processing and graph-oriented memory infrastructure. Teams exploring open-source graph and data pipelines for memory-like knowledge engineering. Architecture docs

Public architecture documentation exists; this page does not normalize evals.

Unmeasured

No same-method score is measured by this leaderboard.

Graph/data visibility

Knowledge graph artifacts can be inspected; answer-level memory traces are not scored.

Unmeasured

Specific provable-forgetting behavior is not scored from public docs.

Open-source stack

Open-source deployment supports portability; client integration surface depends on setup.

Unmeasured

Retrieval-layer tenant controls are not scored from public docs.

Fig. 3 honest claim ledger
FIG. 3

How should you read unmeasured cells?

Unmeasured is a feature of this leaderboard, not a gap to hide. A cell stays unmeasured until public evidence exists or Answer Engine runs a same-method evaluation that can be reproduced.

RULE 01

No inferred accuracy

Vendor-specific benchmark pages are not copied into a shared score unless the method, corpus, and axis match.

RULE 02

No hidden Answer Engine score

The gated benchmark and open harness remain unpublished here, so Answer Engine's score cell is explicitly unmeasured.

RULE 03

Public source or correction

Corrections should point to a public URL for the exact claim, then the cell can move from unmeasured to published or partial.

Fig. 4 public references
FIG. 4

Which sources ground the leaderboard?

These sources define the public comparison surface. They are not private demos, sales calls, or copied benchmark claims.

Zep overview

Public source for Zep Context Graph and Context Lake positioning.

Letta docs

Public source for Letta stateful agents, memory blocks, tools, and RBAC docs.

Cognee docs

Public source for Cognee architecture and open-source memory infrastructure.

Fig. 5 debugging and readiness links
FIG. 5

What should you inspect next?

The leaderboard is for shortlist discipline. Use the readiness checklist and failure-mode catalog to turn the same columns into hands-on review.

Fig. 6 FAQPage schema / visible answers
FIG. 6

FAQ

Why are many leaderboard cells marked unmeasured?

Because a public leaderboard should not invent scores. If Answer Engine has not run a same-method evaluation or the vendor has not published comparable evidence, the cell stays unmeasured.

Does this leaderboard publish an Answer Engine benchmark number?

No. The reproducible benchmark post and open harness are gated separately. This page compares control surfaces and marks accuracy scores as unmeasured until comparable evidence exists.

Which columns matter most for production agent memory?

Inspectability, forgetting, portability, and permission-aware retrieval matter because they decide whether teams can debug, delete, migrate, and safely scope remembered facts.

How should a vendor ask for a correction?

Point to a public source URL for the specific cell. The leaderboard should change when public evidence changes, not when a private claim is asserted.