Carry decisions and their reasons into the next session

CPersona is a memory server for AI agents. It saves your project's decisions and the reasons behind them to a local SQLite database, where a later session can find them.

Works with Claude Code and Codex through MCP.

Session 1Monday
Save this design decision: keep the first version local so setup doesn't depend on a hosted memory service.
Saved the decision and its reason.
store
Next sessionThursday
Why did we decide to keep the first version local?
reconstruct
“keep the first version local so setup doesn't depend on a hosted memory service”quoted from the record saved on Monday
Illustration only, not output from a real session.
  • 81.6%of 500 LongMemEval-S questions answered correctly (2.6.4a1)
  • 1,787tokens per question, the fewest of the memory systems that reached 80% in the same benchmark
  • < 1 sto recall from 100,000 memories on a low-power PC, query embedding included

How these were measured

Set up CPersona with your coding agent

Copy the prompt and paste it into Claude Code or Codex. Your agent proposes the steps first; read them before you approve anything.

The prompt asks your agent to check and explain first. It tells the agent not to run commands, install anything, or change files until you approve the steps.

Setup prompt Copy
Help me set up CPersona for [Claude Code or Codex]. First read the official CPersona setup guide: https://cloto-dev.github.io/CPersona/getting-started/

Start by checking only this client's MCP configuration and tell me whether CPersona is already configured. Treat configuration as sensitive: do not reveal tokens, keys, or other secret values. Do not read unrelated project files.

Do not change or remove any existing memory database, memory file, client setting, or MCP server. Explain the exact setup steps you recommend, including commands, packages or model downloads, files to be edited, and where data will be stored. Wait for my explicit approval before running commands, installing or downloading anything, or changing files or settings. If a step would overwrite or alter existing data or configuration, stop and ask me first. If the plan must change, explain why and ask again.

After I approve the proposed steps, make only those changes. Help me verify the connection and test a saved design decision in a fresh session. Do not claim setup succeeded unless the documented checks pass.

Accurate answers at the lowest token cost

OmniMemEval runs the memory systems it compares through one LongMemEval-S pipeline: 500 questions, answered by gpt-4.1-mini and judged by gpt-4o-mini. CPersona 2.6.4a1 answered 81.6% correctly while sending 1,787 tokens a question to the answer model, the fewest of any system that reached 80% there.

MemOS answers more questions, at 2.3 times the tokens. EverOS is within one run's sampling error (about ±3.5 points) at about seven times the tokens. CPersona calls no generative model to store or recall.

Tokens sent to the answer model per question, for the memory systems that reached 80% on LongMemEval-S
SystemAccuracyTokens per question
CPersona 2.6.4a1 pre-release81.6%1,787
CPersona 2.6.3 current release80.8%2,355
MemOS89.2%4,151
EverOS80.4%12,379

Each CPersona version was run once, by this project, in October 2026; 2.6.3 was measured as its pre-release 2.6.3a1, which has the same code. The other rows are OmniMemEval's own reproductions. Vendors' self-reported results come from other pipelines and are not compared.

Read the results and the full table

Light enough for a low-power PC

On a low-power PC with an Intel N150 (four threads), recall over 100,000 memories took a median of 0.45 s, with the query embedded by jina-v5-nano through CEmbedding on the same machine. All 25 timed queries finished under 0.52 s.

CPersona keeps its records in a SQLite file on your machine and calls no generative model, so there is no hosted memory service to pay for or wait on.

Measured in September 2026 on a synthetic corpus, with the code of that time. The rule registered before the run allows this statement: recall over 100,000 memories completes in under one second on an N100-class PC, query embedding included.

Read the measurement

What is stored, and how recall works

Your memories live in a local SQLite database, and CPersona does not call a generative model. The reconstruct tool returns what it recalls with quotes from the stored records, so you can check the source behind an answer.

A compatible embedding server is recommended for vector search. Without one, CPersona falls back to keyword and full-text search and reports that recall is degraded. Different wording can make a saved memory harder to find. See embedding setup and fallback.

Finding a decision in a later session depends on it having been saved, and on both sessions using the same database and agent_id. A saved record is not guaranteed to appear for every question. The setup guide shows how to check recall across sessions.

Looking for a marketing partner

CPersona is looking for a business partner to lead its marketing. There is no fixed pay at present; we expect to share part of the revenue of a future business built around CPersona, on terms agreed individually.

About the partnership