by jeonjw85
Self-hosted memory server for AI agents. Share context across Claude Code, Codex, OpenCode, and ChatGPT over MCP.
# Add to your Claude Code skills
git clone https://github.com/jeonjw85/kenfoldSee how kenfold compares with popular alternatives.
kenfold is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by jeonjw85. Self-hosted memory server for AI agents. Share context across Claude Code, Codex, OpenCode, and ChatGPT over MCP. It has 61 GitHub stars.
kenfold's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/jeonjw85/kenfold" and add it to your Claude Code skills directory (see the Installation section above).
kenfold is primarily written in Go. It is open-source under jeonjw85 on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh kenfold against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
One memory for all your AI agents. Bring every agent into the fold.
Claude Code, Codex, OpenCode, ChatGPT, and local models each keep their own memory. Kenfold is a self-hosted memory server they all share over MCP: a decision made in one tool is known in the next.
Claude Code ─┐
Codex ───────┤
OpenCode ────┼── Kenfold ── PostgreSQL + pgvector
ChatGPT ─────┤
Local LLM ───┘
The main design choices (see ADR-0001):
Status: Phase 5 in progress. Shared storage, hybrid search, session hooks, memory extraction, code references, and OAuth remote access are implemented. Phase 5 adds export/import, a review dashboard, and consolidation. The LoCoMo and LongMemEval benchmark harness is ready (
make bench); results are still pending.
Requirements: Docker with Compose. Go 1.27+ only for development.
make up # Postgres 18 + pgvector and Kenfold on http://127.0.0.1:7077
make key AGENT=claude-code # prints an API key for Claude Code (shown once)
make key AGENT=codex # one key per agent, so every memory is attributed correctly
make up gives full-text search. For better search, which also finds memories phrased differently from the query, run the local search models instead:
make up-embed # adds Ollama with bge-m3 and a llama.cpp reranker (first run downloads ~1.9 GB)
Memories written without embeddings are embedded automatically once a model is available. The reranker reads the query and each candidate memory together; on the internal eval set it raised recall@5 from 0.90 to 0.96 (see Search quality), at the cost of 1–2 seconds per search on CPU. make up-embed RERANK=0 leaves it out.
Each agent uses its own key. The key decides the agent name recorded on every memory.
Claude Code
claude mcp add --scope user --transport http kenfold http://127.0.0.1:7077/mcp \
--header "Authorization: Bearer kf_..."
Codex (~/.codex/config.toml); set KENFOLD_CODEX_KEY in the environment Codex starts from:
[mcp_servers.kenfold]
url = "http://127.0.0.1:7077/mcp"
bearer_token_env_var = "KENFOLD_CODEX_KEY"
For the desktop app or IDE extension, which may not see your shell environment, use http_headers = { "Authorization" = "Bearer kf_..." } instead.
OpenCode (opencode.json in a project, or ~/.config/opencode/opencode.json):
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"kenfold": {
"type": "remote",
"url": "http://127.0.0.1:7077/mcp",
"oauth": false,
"headers": { "Authorization": "Bearer {env:KENFOLD_OPENCODE_KEY}" }
}
}
}
stdio (the client starts Kenfold as a subprocess): build with make build, then run bin/kenfold mcp with KENFOLD_AGENT set to the agent's name. The stdio server connects to the database directly (KENFOLD_DATABASE_URL, default: the compose database), so it does not use API keys and the agent name is whatever you configure.
claude mcp add --env KENFOLD_AGENT=claude-code --transport stdio --scope user kenfold -- /abs/path/to/bin/kenfold mcp
[mcp_servers.kenfold] # Codex
command = "/abs/path/to/bin/kenfold"
args = ["mcp"]
env = { KENFOLD_AGENT = "codex" }
With make up-embed, also set KENFOLD_EMBED_URL=http://127.0.0.1:11435/v1 for stdio servers.
ChatGPT, claude.ai, and other machines need a public HTTPS URL, behind a tunnel or a reverse proxy, with KENFOLD_PUBLIC_URL set:
KENFOLD_PUBLIC_URL=https://kenfold.example.com make up
docker compose exec kenfold /usr/local/bin/kenfold password # the owner password: approves clients, logs in to the dashboard
This turns on Kenfold's OAuth 2.1 server. ChatGPT and claude.ai find it on their own, and you approve each on a consent page, read-only or read and write. Agents with API keys use the public URL as before. See docs/deploy.md for Tailscale Funnel, Cloudflare Tunnel, and Caddy, and read its checklist before exposing your memory to the internet.
MCP tools only run when the model decides to call them. Hooks make memory automatic:
~/.local/state/kenfold, readable only by you, secrets redacted). This adds no network calls.episodic memory), so the next agent, in any tool, knows what happened. If Kenfold is down, the summary is kept and sent at the next session start.Sessions outside a git repository are not recorded. The hook never blocks the agent: problems are reported as a warning.
Set it up per agent. Store the agent's key in a file, then print the hooks configuration:
make build
mkdir -p ~/.config/kenfold && (umask 077; make -s key AGENT=claude-code > ~/.config/kenfold/claude-code.key)
bin/kenfold hook config claude-code --key-file ~/.config/kenfold/claude-code.key
Merge the printed hooks object into ~/.claude/settings.json. For Codex, use a codex key and bin/kenfold hook config codex, and save the output as ~/.codex/hooks.json. Codex asks you to review new hooks: run /hooks once to trust them.
--no-capture loads memory at session start without recording sessions. The hook also checks the code that memories refer to at session start (see Code references); --no-refs turns that off.
Session summaries record what happened. With a chat model configured, Kenfold also reads each summary and proposes the durable memories in it: project rules ("use pnpm, not npm"), facts about the code ("the webhook dedupes by event_id"), and your preferences ("answer in Korean"). It skips one-off tasks and anything from pasted documents.
make up-extract # up-embed plus qwen3.5:4b (first run downloads ~5.2 GB)
make review # go through proposed memories: approve, reject, or replace an old one
Extracted memories are proposed: nothing is served to agents until you approve it, with make review or in the dashboard. Each one shows the quote it came from and any existing memory it resembles; approving with "replace" retires the old one. A memory you reject is not proposed again. If an agent omits the memory type, the same model chooses it. This adds about 6 seconds to a remember call on CPU. If the model takes longer than 15 seconds (while it loads, or is busy extracting), the default type is used.
The model runs on CPU in Docker and takes 10–30 seconds per session, in the background. On CPUs with performance and efficiency cores, set THREADS to the number of performance cores (default 4; make up-embed THREADS=6): Ollama's default of one thread per core was 30–80× slower on an Apple M5, for embeddings as well as chat. Any OpenAI-compatible chat API works (KENFOLD_CHAT_URL, KENFOLD_CHAT_MODEL). Quality on the internal eval set is recorded in internal/extract/testdata/RESULTS.md (make eval).
KENFOLD_EXTRACT_POLICY=auto activates confident extractions that resemble no existing memory without review. Preferences are always reviewed. Keep the default unless you trust every source of your sessions: extraction is where instructions hidden in pasted content could become memory, and review is the main defense.
With a chat model configured (make up-extract), Kenfold proposes to retire duplicate or conflicting memories and summarize old sessions:
Nothing changes until you apply a proposal, on the dashboard's Consolidation page or with the CLI. Retired memories are kept as history and linked from the memory that replaces them. Nothing is rewritten: a duplicate or contradiction keeps one of the existing memories as it was written. A proposal you reject is not made again.
bin/kenfold consolidate list # pending proposals
bin/kenfold consolidate apply <id> # or reject <id>
bin/kenfold consolidate run # look now instead of every 15 minutes
Candidates are pairs of similar memories. The model judges each pair twice, with the memories in both orders, and a pair is proposed only if both answers agree. Digests are checked against their sessions: every file name, identifier, and number in a digest must appear in them. The model judges each pair in 7–16 seconds on CPU, and the worker yields to extraction. On the internal eval set it judged 12 of 12 pairs correctly (internal/consolidate/testdata/RESULTS.md). See ADR-0004 for the rules. KENFOLD_CONSOLIDATE=false turns consolidation off.