A self-learning skill layer for Claude Code — distills skills from your real sessions, updates them as you work, and prunes the ones that stop getting used. No daemon, no benchmark.
# Add to your Claude Code skills
git clone https://github.com/tigerless-labs/autoharnessLast scanned: 7/3/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-07-03T07:21:27.430Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}autoharness is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by tigerless-labs. A self-learning skill layer for Claude Code — distills skills from your real sessions, updates them as you work, and prunes the ones that stop getting used. No daemon, no benchmark. It has 127 GitHub stars.
Yes. autoharness passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/tigerless-labs/autoharness" and add it to your Claude Code skills directory (see the Installation section above).
autoharness is primarily written in Python. It is open-source under tigerless-labs on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh autoharness against similar tools.
No comments yet. Be the first to share your thoughts!
autoharness is a self-learning skill layer for Claude Code. It learns skills from your real sessions, merges same-scenario ones instead of stacking near-duplicates, updates them in use, and prunes any that stop getting used — so the layer stays clean on its own, touching only the skills it wrote itself.
Same model, different harness — 42% → 78% on CORE-Bench (HAL). The harness does much of the work (swyx's Big Model vs Big Harness), yet it's still rebuilt by hand every model generation. autoharness bets one slice of it — the skill layer — can maintain itself.
| Learns from real work | Each episode is distilled into a skill from the session you were already having — no separate data-collection or replay loop. |
| Groups, doesn't just pile up | A new episode doesn't always add a skill — the reflector compares it against what's there and folds same-scenario skills into one, so the layer consolidates by category instead of accreting near-duplicates. |
| Validated in use, not on a benchmark | A skill survives by being adhered to in later turns (invocation rate), not a held-out score. No oracle on the active path, and no tokens spent on a dedicated eval. |
| Only its own skills | Touches only the skills it generated through this plugin — everything else, whether you wrote it or installed it, is left completely alone. |
| Evidence kept for later | Every create/update logs its scenario and decision to a per-skill ledger — the raw material to build a benchmark from real usage if you ever want one. |
Requires python3 on your PATH — autoharness runs entirely as Python (zero third-party
dependencies); its hooks and MCP server won't fire without it.
/plugin marketplace add tigerless-labs/autoharness
/plugin install autoharness@autoharness
Run /reload-plugins (or restart Claude Code), then check it's live:
/plugin # autoharness@autoharness — enabled
/mcp # stage_skill — connected
Zero config. It now watches your sessions and lands learned skills into .claude/skills/ in the
background. Cadence and lifecycle thresholds are tunable — see Configuration.
/plugin marketplace update autoharness # pull the latest release
/plugin update autoharness # update the installed copy
/reload-plugins # apply in the running session
/reload-plugins re-arms the hooks and the MCP server in place; restarting Claude Code works too.
The installed copy is cached by the version in plugin.json; releases always bump it, which is
what makes the update take effect.
claude plugin uninstall autoharness@autoharness # stops the hooks + MCP server
claude plugin marketplace remove autoharness # optional — also drops the install source
Uninstalling only stops it from running — the skills it landed and its own state live outside the
plugin and stay on disk. To clear those too, delete its state dir (~/.claude/autoharness/ global,
<repo>/.claude/autoharness/ per project) and the self-authored skills under .claude/skills/ (each
carries a self-authored ledger marker, so they're easy to tell from yours). Your own skills are
never touched.
Every knob is an AUTOHARNESS_* environment variable with a built-in default — nothing to
configure unless you want to change the pace.
| Variable | Default | What it does |
|---|---|---|
AUTOHARNESS_REFLECT_EVERY_N |
3 |
Reflection cadence: a background reflection run fires every N host turns. Lower = learns faster, spawns more child sessions. |
AUTOHARNESS_DIGEST_EXCHANGES |
20 |
How many exchanges before the episode window are compressed into the reflector's prior-context digest (text + tool names only). |
AUTOHARNESS_MATURITY_PROJECT |
100 |
Probation gate, project layer: a skill only graduates — starts counting against capacity and becoming evictable — after this many requests have arrived in its layer since it landed. Until then it's recalled as usual but can't be archived. |
AUTOHARNESS_MATURITY_GLOBAL |
300 |
Same gate for the global layer — higher because a global skill loads in every project. |
AUTOHARNESS_CAPACITY_PROJECT |
50 |
Cap on mature skills in the project layer. Capacity contention is the only death: nothing is archived until the mature pool exceeds this, then the lowest invocation rates go first. |
AUTOHARNESS_CAPACITY_GLOBAL |
20 |
Same cap for the global layer — smaller because its blast radius is every project. |
Set them in the environment Claude Code launches with — either the shell
(export AUTOHARNESS_REFLECT_EVERY_N=10) or the env map in .claude/settings.json:
{ "env": { "AUTOHARNESS_REFLECT_EVERY_N": "10" } }
Hooks read the environment on every event, so a change applies from the next session. The defaults
are deliberate placeholders pending empirical calibration (tracked under experiments/); size caps
on captured windows and staged skill bodies are fixed constants, not env knobs.
A learning pipeline runs beside the host and stays off its recall path — symbols are plain native skills, recalled by the host's own name-and-description mechanism as if a human had written them.
Diagram source: docs/assets/pipeline.mmd — re-render to pipeline.svg after editing.
| Component | Role |
|---|---|
| CAP · capture | Hook-driven dumb pipe: grabs each turn (user input, agent output, tool I/O), redacts at egress, points back at the host log instead of copying it. |
| REF · reflect | At an episode boundary, reads the existing skill index and decides add / merge / patch / drop a support file / delete — emits an intent (body, delta, or path, plus reason and evidence). Proposes only; no write tools. |
| promoter · validate·store | The only writer. Lints the intent in memory (safety, structure, ledger, completeness, self-authored-only) and on pass does an atomic rename into the live skill directory. |
| MNG · lifecycle | Daemon-free: recomputed lazily at session start, once per session. Ranks symbols by invocation rate — calls over the requests that arrived since the symbol was created, so the measure is opportunity-relative and a closed laptop doesn't age anyone out (the wall-clock replacement). New symbols sit in probation until they've had a fair sample of requests: recalled as usual, but neither counted against the cap nor evictable. Capacity contention is the only death — nothing is archived until a layer's mature pool exceeds its cap, then the lowest rates go first. Archives, never deletes: an archived symbol is a directory moved out of recall, and moving it back revives it. |
| LED · ledger | Per-symbol append-only sidecar: why each symbol was born or changed, with evidence and a reflection watermark. Kept out of the skill body so recall stays clean. |
Everything autoharness does lands on disk as plain files — a demo is just opening them in the right order. For a fast-paced run, speed up the loop first (see Configuration):
{ "env": { "AUTOHARNESS_REFLECT_EVERY_N": "1",
"AUTOHARNESS_MATURITY_PROJECT": "5",
"AUTOHARNESS_CAPACITY_PROJECT": "2" } }
1 · The pipeline running. Work a few normal turns on anything non-trivial (debug something, figure out a workflow). Every Nth turn a background reflection fires — nothing blocks your session. Its bookkeeping is visible in the state dir:
ls .claude/autoharness/ # per project — ~/.claude/autoharness/ for the global layer
requests # layer request counter (MNG's denominator)
session-<id> # per-session turn count toward the next reflection
offset-<id> # byte watermark: where the last captured window ended
intents/ # queued skill proposals awaiting the promoter
2 · A skill is born. After a reflection lands, a new folder appears under .claude/skills/
(project) or ~/.claude/skills/ (global — for techniques that aren't repo-specific). Use ls -la:
the interesting files are hidden.
.claude/skills/<name>/
SKILL.md # the skill itself — plain native format, nothing proprietary
.ledger.jsonl # LED: why it was born / changed (append-only)
.sidecar.json # lifecycle counters MNG reads
references/evidence-*.md # the transcript slice that justified each ledger entry
scripts/ templates/ ... # optional support files the reflector attached
3 · LED — the paper trail. cat .ledger.jsonl — one JSON line per lifecycle event:
{"action": "create", "reason": "User asked about the correct command to update a plugin ...", "evidence": "references/evidence-21cd22cc.md"}
{"action": "patch", "reason": "User discovered /reload-plugins is required in-session ...", "evidence": "references/evidence-1a4ec51d.md"}
action + reason + evidence — and the evidence file is a real, redacted slice of the session
that taught it, materialized by the promoter (content-addressed, so the model never names files).
This is the "evidence kept for later" from