by ruvnet
🛠️ The meta-harness for AI agents — scaffold your own focused, branded agent harness with its own npx CLI, MCP server, memory, learning loop, and witness-signed releases. Works with Claude Code, Codex, pi.dev, Hermes, OpenClaw, and RVM (hardware-isolated sandbox).
# Add to your Claude Code skills
git clone https://github.com/ruvnet/metaharnessLast scanned: 6/30/2026
{
"issues": [
{
"type": "npm-audit",
"message": "@vitest/mocker: Vulnerability found",
"severity": "medium"
},
{
"type": "npm-audit",
"message": "esbuild: esbuild enables any website to send any requests to the development server and read the response",
"severity": "medium"
},
{
"type": "npm-audit",
"message": "vite: Vite Vulnerable to Path Traversal in Optimized Deps `.map` Handling",
"severity": "high"
},
{
"type": "npm-audit",
"message": "vite-node: Vulnerability found",
"severity": "medium"
},
{
"type": "npm-audit",
"message": "vitest: Vulnerability found",
"severity": "critical"
}
],
"status": "FAILED",
"scannedAt": "2026-06-30T07:53:31.534Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}See how metaharness compares with popular alternatives.
metaharness is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by ruvnet. 🛠️ The meta-harness for AI agents — scaffold your own focused, branded agent harness with its own npx CLI, MCP server, memory, learning loop, and witness-signed releases. Works with Claude Code, Codex, pi.dev, Hermes, OpenClaw, and RVM (hardware-isolated sandbox). It has 675 GitHub stars.
metaharness failed SkillsLLM's automated security scan, which flagged one or more high-severity issues. Review the Security Report section carefully before using it.
Clone the repository with "git clone https://github.com/ruvnet/metaharness" and add it to your Claude Code skills directory (see the Installation section above).
metaharness is primarily written in TypeScript. It is open-source under ruvnet on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh metaharness against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
npx metaharness · open the Studio →
(Repo: ruvnet/metaharness · CLI: metaharness · Library: @ruvnet/agent-harness-generator)
Every serious repo deserves its own agent. A repo-aware CLI, a repo-aware coding agent, a local MCP server, memory scoped to the project, skills generated from the actual file layout, governance policy, release verification, witness-signed provenance.
metaharness mints those, on demand, from a GitHub URL or a blank slate. It is not another agent framework. It is a factory for agent frameworks.
The model is replaceable. The harness is the product.
In under 60 seconds, in your browser, with nothing leaving your machine:
Output is an npm-publishable .zip with your name on it, your branding, your npx <your-name> CLI.
@metaharness/arc-agi-3 owns exact observations,
persistent evidence-backed memory, belief-state exploration, guarded actions,
supervisor interventions, checkpoints, and hash-chained receipts.
@metaharness/arc-agi-3-chatgpt exposes that
controller as a remote MCP server and MCP Apps canvas for ChatGPT Developer
Mode. ChatGPT is the OpenAI reasoning host: neither package uses the OpenAI
API or an OPENAI_API_KEY. An opt-in ARC-specific AVO loop adds governed
candidate selection, lineage, memory, blocking supervision, and an explicitly
separate retrodiction arm. The frozen synthetic mechanism gate passes, and
one actor-declared-clean single-game smoke favored AVO 3.2676 to 0.3968, but that
non-competition result is not claim-eligible. All packages remain private and
experimental; no ARC performance claim is made until an official closed
scorecard satisfies the frozen controlled-ablation gate in
ADR-254.@metaharness/avo lets an agent repeatedly inspect, edit,
execute real tools, evaluate, repair/revert, branch, consult structured RVF
memory, and commit—while MetaHarness retains immutable capabilities, budgets,
promotion, quarantine, rollback, and signed replay receipts. Simple work stays
on Darwin's fast path. The runtime has a deterministic 205-action RVF
interruption proof; the stronger “AVO-class” claim remains blocked on the
preregistered 100-task unseen SWE-bench gate in ADR-251.
ADR-253 now enforces
that boundary at publication against the exact tag SHA and npm tarball,
protected claim semantics, measured cost, lineage roots, and two independent
graders.@metaharness/host-prime-agent, --host prime-agent) emits your
tools as project-scoped, Python-backed Prime Agent skills (.prime/agent/skills/ — the host has
no MCP) plus an install runbook. Fail-closed: Prime Agent can't enforce a deny-list itself,
so a non-empty deny-list ships a prominent SANDBOX-REQUIRED.md instead of silently dropping
your posture. From its design we also shipped RefineMutator — an evidence-backed proposer in
@metaharness/darwin that must cite the
failing traces behind each edit or propose nothing — and an opt-in --sessions crash-recoverable
JSONL session log whose Rust and TS (wasm) replays produce identical state hashes. Its PTC
("kernel as the only tool") claim is honestly deferred behind a pre-registered A/B in
evals-toolcall. See ADR-246 /
ADR-247. ($0)npx metaharness score <repo> reads
the repo (never runs it) and prints a one-screen report card — how well a
harness fits, how likely it is to build, how safe the tools are, and the
rough cost per run — so you know what you'll get before scaffolding.@metaharness/router
routes each request to the right model from your own results — same quality,
far less spend. Works out of the box with zero native deps; train it on your
data for a sharper fit (npm i @metaharness/router). Add the optional
@ruvector/tiny-dancer
to train a fast native model instead — same training data, no API change.@metaharness/field-memory is an experimental,
single-process-by-default packed-attractor layer with verified outcome caps,
support quarantine, decay, cost-aware choice, and no raw-episode API. Its
RuVector adapter fails closed unless a corrected flat-index wrapper is used.@metaharness/darwin) wired in —
run npm run evolve and the harness mutates its own config, tests each change in a
sandbox, and keeps only what measurably improves. The model stays frozen; the harness
evolves. Safe by default (no network, no API key; pure refactor/tuning behind a safety
gate). Validated on real SWE-bench Lite bug-fixing. --no-darwin to skip.--field-memory to scaffold a small bootstrap around the forthcoming
@metaharness/field-memory@^0.1.0 package. The generated harness uses packed
attractors, three distinct verifier-derived principals for support, no
hysteresis, and a bounded drift window. Principal independence depends on the
verifier's identity system.
It will not open storage or accept shared updates until the deployment
supplies an absolute storage path, a compatible adapter, and a principal
verifier backed by deployment authentication. Package influence caps are
centroid-local; fleet-wide admission and rate limits belong in that verifier.
The RuVector adapter requires a configuration-verified, in-place FlatIndex
with cosine distance and rejects HNSW or unverified legacy stores.
Minimum support quarantines routing but does not make stored singleton
aggregates confidential; protect field snapshots and registry storage.
A deployment-secret identity hash key of at least 32 bytes is also required
and must survive state restore. Conflicting package declarations or an
existing field module fail before disk writes, and harness upgrade reapplies
the opted-in overlay only after validating its exact, versioned manifest
contract. Unsupported overlay schemas require an explicit migration.
storage.writerScope: process requires one writer service; multi-process or
fleet use requires distributed. The flag is off by default, so existing
scaffold output is unchanged.@metaharness/weight-eft, metaharness weight-eft) takes the
complementary lever to Darwin's gradient-free evolution: it exports the harness's gold-resolved
archive into standard SFT/DPO sets and LoRA-tunes the open cheap tier (GLM/Qwen), so the
cost-cascade escalates to Opus/GPT less often. It attacks cost (fewer $0.50 escalations),
not the frontier ceiling — and stays honest about it. Strict train/eval-disjointness +
reward-hacking filters keep the lift real; the tune is a gene Darwin can prune if it overfits.
See ADR-198. ($0 / GPU-gated.)A generated harness is a starting point you own, not a fixed framework. Open it and make it yours:
harness doctor / harness validate keep it healthy as you trim.