No inferred accuracy
Vendor-specific benchmark pages are not copied into a shared score unless the method, corpus, and axis match.
A useful memory leaderboard should not guess hidden scores. Compare systems by published evidence, inspectability, forgetting, portability, and permission controls, then mark every missing measurement as unmeasured. This leaderboard is a public scorecard for evaluating Answer Engine, Mem0, Zep, Letta, and Cognee without inventing benchmark numbers.
The best system depends on the control surface you need. Use this table to separate published evidence from unmeasured claims, then follow the source links before you shortlist a vendor.
| System | Best fit | Published eval evidence | Accuracy score | Inspectable retrieval | Forgetting / deletion | Portability | Permission-aware retrieval |
|---|---|---|---|---|---|---|---|
| Answer Engine Tenant-scoped content memory with source artifacts, library scope, MCP tools, CLI sync, and REST delivery. | Teams that want source-aware memory through MCP, REST, and CLI with tenant-isolated retrieval and public roadmap labeling. | Roadmap The reproducible benchmark post is gated; this row does not publish a score. | Unmeasured No public Answer Engine accuracy score is claimed here. | Source-aware now; trace viewer roadmap Citations and source-aware retrieval ship; full live inspector is roadmap-labeled. | Delete handles and source lifecycle Deletion posture is documented; deeper provable-forgetting audit is roadmap-labeled. | MCP, REST, and CLI One memory layer can serve Claude Code, Cursor, Codex, Gemini CLI, Windsurf, and API clients. | Tenant and library scoped Repository queries bind tenant_id and library scope before retrieval. |
| Mem0 Managed memory engine with add, search, update, delete, platform governance, and open-source options. | Product teams that want a managed user-memory platform with SDK-first setup and hosted operational controls. | Vendor-published memory evals Public docs include a memory evaluation section; methodology is not normalized here. | Unmeasured No same-method score is measured by Answer Engine in this leaderboard. | Audit and dashboard surfaces Public docs describe platform audit and governance; retrieval-trace depth is not scored here. | Delete memory operation Delete operations are documented; derived-artifact audit is not normalized here. | SDKs, API, MCP entry Portable at integration level; export/migration depth is not scored here. | Workspace governance Enterprise controls are documented; tenant retrieval SQL is not inspectable from docs. |
| Zep Temporal Context Graph and Context Lake for user, business, document, and JSON memory. | Teams prioritizing temporal knowledge graph semantics and prompt-ready context from graph-derived memory. | Vendor-published graph memory paper Zep publishes graph-memory architecture and evaluation material; not normalized here. | Unmeasured This leaderboard does not copy or compare a vendor-specific benchmark number. | Graph context visible by API/docs Temporal graph structure is documented; answer-level trace inspection is not scored here. | Data management documented Deletion and retention behavior should be verified against current account controls. | SDK/API, migration docs Public docs include migration paths; MCP portability depends on integration choices. | Governed Context Lake Governance is described; retrieval-layer tenant implementation is not scored here. |
| Letta Stateful agents with memory blocks, shared memory, archival memory, tools, and agent configuration files. | Builders who want stateful agents with explicit memory blocks, archival memory, tools, and agent files. | Evals docs exist Letta documents eval tooling; this leaderboard does not run a comparable test. | Unmeasured No same-method leaderboard score is published here. | Agent state and memory blocks Memory state is explicit; retrieval-trace depth depends on implementation. | Mutable agent memory Block and passage APIs exist; provable derived-artifact deletion is not scored. | API, tools, AgentFile Portable through platform and files; client-neutral MCP memory surface varies by setup. | RBAC documented Role-based access control is documented; vector-layer ACL implementation is not scored. |
| Cognee Open-source data processing and graph-oriented memory infrastructure. | Teams exploring open-source graph and data pipelines for memory-like knowledge engineering. | Architecture docs Public architecture documentation exists; this page does not normalize evals. | Unmeasured No same-method score is measured by this leaderboard. | Graph/data visibility Knowledge graph artifacts can be inspected; answer-level memory traces are not scored. | Unmeasured Specific provable-forgetting behavior is not scored from public docs. | Open-source stack Open-source deployment supports portability; client integration surface depends on setup. | Unmeasured Retrieval-layer tenant controls are not scored from public docs. |
Unmeasured is a feature of this leaderboard, not a gap to hide. A cell stays unmeasured until public evidence exists or Answer Engine runs a same-method evaluation that can be reproduced.
Vendor-specific benchmark pages are not copied into a shared score unless the method, corpus, and axis match.
The gated benchmark and open harness remain unpublished here, so Answer Engine's score cell is explicitly unmeasured.
Corrections should point to a public URL for the exact claim, then the cell can move from unmeasured to published or partial.
These sources define the public comparison surface. They are not private demos, sales calls, or copied benchmark claims.
Defines MCP as an open standard for connecting AI applications to external systems.
Documents MCP clients, servers, tools, resources, transports, and capability discovery.
Public source for Mem0 platform positioning and feature surface.
Public source for Zep Context Graph and Context Lake positioning.
Public source for Letta stateful agents, memory blocks, tools, and RBAC docs.
Public source for Cognee architecture and open-source memory infrastructure.
The leaderboard is for shortlist discipline. Use the readiness checklist and failure-mode catalog to turn the same columns into hands-on review.
Because a public leaderboard should not invent scores. If Answer Engine has not run a same-method evaluation or the vendor has not published comparable evidence, the cell stays unmeasured.
No. The reproducible benchmark post and open harness are gated separately. This page compares control surfaces and marks accuracy scores as unmeasured until comparable evidence exists.
Inspectability, forgetting, portability, and permission-aware retrieval matter because they decide whether teams can debug, delete, migrate, and safely scope remembered facts.
Point to a public source URL for the specific cell. The leaderboard should change when public evidence changes, not when a private claim is asserted.