by alanhuangyoo
Taking the pi coding agent to Claude Code-level performance: 0.539 → 0.773 pass@1 on Terminal-Bench 2.1 with the same self-hosted 27B model, by re-engineering the agent.
# Add to your Claude Code skills
git clone https://github.com/alanhuangyoo/cruxLast scanned: 10/6/2026
{
"issues": [
{
"type": "npm-audit",
"message": "@earendil-works/gondolin: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "@vitest/coverage-v8: Vulnerability found",
"severity": "medium"
},
{
"type": "npm-audit",
"message": "@vitest/mocker: Vitest: Path Traversal / Arbitrary File Read via @vitest/mocker Redirect Mock",
"severity": "medium"
},
{
"type": "npm-audit",
"message": "brace-expansion: brace-expansion: Quadratic-time expansion of the `{a},b}` rewrite causes CPU denial of service",
"severity": "high"
},
{
"type": "npm-audit",
"message": "braces: braces vulnerable to stack-exhaustion denial of service through deeply nested patterns",
"severity": "high"
},
{
"type": "npm-audit",
"message": "fast-glob: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "micromatch: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "node-forge: node-forge RSA PKCS#1 v1.5 signature verification accepts extra nested DigestAlgorithm elements",
"severity": "high"
},
{
"type": "npm-audit",
"message": "shelljs: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "shx: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "source-map-js: source-map-js allows event-loop denial of service through indexed source-map section offsets",
"severity": "high"
},
{
"type": "npm-audit",
"message": "undici: undici vulnerable to Denial of Service via unhandled error in WebSocket permessage-deflate decompression",
"severity": "high"
},
{
"type": "npm-audit",
"message": "vitest: Vulnerability found",
"severity": "medium"
}
],
"status": "WARNING",
"scannedAt": "2026-10-06T11:02:28.628Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}See how crux compares with popular alternatives.
crux is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by alanhuangyoo. Taking the pi coding agent to Claude Code-level performance: 0.539 → 0.773 pass@1 on Terminal-Bench 2.1 with the same self-hosted 27B model, by re-engineering the agent. It has 186 GitHub stars.
crux returned warnings in SkillsLLM's automated security scan. It has no critical vulnerabilities, but review the flagged issues in the Security Report section before adding it to your workflow.
Clone the repository with "git clone https://github.com/alanhuangyoo/crux" and add it to your Claude Code skills directory (see the Installation section above).
crux is primarily written in TypeScript. It is open-source under alanhuangyoo on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh crux against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Crux takes pi, an open-source coding agent, and re-engineers its agent layer for long-horizon terminal work. On Terminal-Bench 2.1 — 89 real tasks spanning compilers, emulators, cryptanalysis, ML training and systems administration, each graded by a hidden test suite inside a container — it lifts pi from 0.539 to 0.773 pass@1, to the level of Claude Code running the same model (0.730).
The model weights never change. Crux's work is in the agent layer: the loop and its recovery paths, context management and compaction, tool contracts, and the runtime that keeps multi-hour runs alive. Every change ships with the measurement that justified it.
Self-hosted Qwen3.8-27B · the 89 Terminal-Bench 2.1 tasks that run without a GPU inside the container · 8× agent budget · pass@1 over whole runs.
| Configuration | pass@1 |
|---|---|
| Crux | 0.773 (0.793 on a second run) |
| Claude Code, same model | 0.730 (32K window) |
| Crux — agent loop and tool contracts only | 0.591 |
| pi, upstream | 0.539 |
Reproducibility. Two independent runs of the final configuration scored 0.773 and 0.793. The benchmark's run-to-run variance was measured directly, and every comparison in this repository is paired task by task to account for it.
flowchart TB
subgraph agent["Agent · packages/"]
direction LR
ctx["Context engine<br/>window-fitted compaction<br/>verbatim task"] <--> loop["Agent loop<br/>state machine · recovery<br/>budget pacing"] <--> tools["Tools<br/>bash · read · edit · write<br/>stream watchdog"]
end
subgraph evals["Evaluation system · benchmark/"]
direction LR
pre["Preflight"] --> run["Containerized runs"] --> attr["Failure attribution"] --> stats["Paired sign tests<br/>→ next change"]
end
agent ==> evals
Each capability below is tied to the trajectory evidence that motivated it.
| Capability | Evidence |
|---|---|
| Summarization requests sized to the window they live in | History and summary share one window and nothing checked their sum: 2,115 of 2,370 compactions had been rejected outright. Now 100% succeed |
| Adaptive retry on rejection | Characters per token ranges from 1.95 to 3.99 across real trial text, so a rejected request halves and retries rather than trusting a constant |
| Partial summaries kept | A summary cut at the output cap still carries the work; discarding it disabled compaction on smaller windows |
| Cut-point search that always makes progress | A large trailing tool result no longer leaves the search with nowhere to cut |
| Full summaries for single-task sessions | An agent run on one task is one conversational turn, and every compaction had taken a short turn-fragment path that kept 2,302–4,443 characters of ~250,000 tokens. Now the full structured summary |
| The task's exact words survive compaction | The original request is carried verbatim and re-read from the session each time, never paraphrased — the approach Codex (openai/codex#48115) and hermes-agent take |
| Capability | Evidence |
|---|---|
| Stream watchdog | SDK timeouts cover getting a response, not keeping one: 25 timed-out trials had been silent for a median 112 of their 121 minutes |
| Default command timeout | Ten minutes, chosen from data: 161 commands legitimately ran 2–10 minutes, against 60 runaway ones |
| Process-tree reaping | A detached child writing to an inherited pipe once held a container for eight hours after 48 seconds of work |
| Resume after OOM kill | A cgroup OOM kill ends every process in the group; the run now resumes instead of scoring zero |
| Capability | Evidence |
|---|---|
| Explicit loop state machine | Every recovery path — truncation, budget, escalation — lives in one immutable LoopState with explicit limits and a recorded transition reason per turn |
| Two-phase truncation recovery | Raise the output ceiling and retry silently first; speak to the model only if it truncates again — Claude Code's design |
| Reasoning carried across a cut | Reasoning is not replayed between turns, and 292 turns in 60 trials had ended mid-thought and restarted from scratch. The tail of the reasoning is now handed back |
| Stall detection | A turn that stops inside its reasoning with no answer is treated as a stall, not as completion |
| Escalation for truncated tool calls | All 42 cut tool calls had stopped at the initial 16K with the model's own 32K unused; they now get the raised ceiling, as in Claude Code and hermes-agent |
| Budget-aware pacing | The loop sees its deadline, warns before it, and questions a stop while most of the budget is unspent; each task's budget follows its own time limit (80–1,600 minutes) |
| Capability | Evidence |
|---|---|
| Edit failures that show the file | Instead of "must match exactly", a failing edit returns the nearest region of the file by bigram similarity, so the next attempt targets real text |
| Head-and-tail command output | Long output keeps its first lines as well as its last, so the first compiler error survives truncation |
| Visible working notes | The model is told its reasoning is not kept and writes short notes beside tool calls; the median visible text beside a tool call had been 0 characters |
| Calibrated output ceiling | A per-response default sized from this deployment's own p50/p95/p99 output distribution |
packages/ the agent: core loop, model layer, tools, CLI (built on pi)
benchmark/ the evaluation system
src/crux/ harbor agent, prompt sections, failure analysis
scripts/ launchers, preflight, paired comparison, health checks
tests/ 486 tests
cd benchmark && uv sync
harbor run --dataset terminal-bench/terminal-bench-2-1 \
--agent crux.pi_agent:CruxPiAgent --model openai/<model>
benchmark/scripts/preflight.sh <launcher> validates the endpoint, the launch
settings and the harness checksum before a run. See
benchmark/README.md for the evaluation system.
Crux is built on pi by Earendil Works. Upstream issues and pull requests belong at earendil-works/pi.
| Package | Description |
|---|---|
| @earendil-works/pi-coding-agent | Coding agent CLI (installs as crux and pi) |
| @earendil-works/pi-agent-core | Agent runtime with tool calling and state management |
| @earendil-works/pi-ai | Unified multi-provider LLM API |
| @earendil-works/pi-tui | Terminal UI components |
| @earendil-works/chord | Application-composition runtime for services |