by GhabiX
An enhanced Codex CLI for complex, long-running software engineering tasks—resolving 89% more tasks than Codex at 27% lower cost, with up to 10× the effective context length.
# Add to your Claude Code skills
git clone https://github.com/GhabiX/SpineCodexLast scanned: 8/11/2026
{
"issues": [
{
"file": "codex-rs/skills/src/assets/samples/plugin-creator/SKILL.md",
"line": 143,
"type": "prompt-injection",
"message": "Possible concealment directive: \"Do not tell the user\"",
"severity": "medium"
}
],
"status": "PASSED",
"scannedAt": "2026-08-11T05:06:43.831Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}SpineCodex is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by GhabiX. An enhanced Codex CLI for complex, long-running software engineering tasks—resolving 89% more tasks than Codex at 27% lower cost, with up to 10× the effective context length. It has 104 GitHub stars.
Yes. SpineCodex passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/GhabiX/SpineCodex" and add it to your Claude Code skills directory (see the Installation section above).
SpineCodex is primarily written in Rust. It is open-source under GhabiX on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh SpineCodex against similar tools.
No comments yet. Be the first to share your thoughts!
SpineCodex is an enhanced, independently maintained version of the OpenAI Codex CLI for complex, long-running software engineering tasks. It inherits your existing Codex configuration and works out of the box. Compared with Codex, it resolves 89% more tasks at 27% lower total cost on SWE-Milestone and extends the effective working context by up to 10×. It also improves the average score by 10.8 points on ProgramBench and the mean score by 9.2 points on FrontierSWE.
| Linear context | SpineCodex |
|---|---|
| ❌Run out of context? | ✅256K → 2.5M Effective Working ContextSpineJIT compiles completed branches into semantic Node Memory, extending effective working context beyond the native window. |
| ❌Drift after repeated compaction? | ✅Minimum Effective Context. Maximum Focus.Through the SpineTree, the agent manages tasks and context as one unified system, staying focused on the minimum context required by the current task. |
| ❌Lose patience and focus on long tasks? | ✅Recursive Subagent Scaling on Demand.SpineJIT lets the agent recursively unfold into specialized subagents on demand, bringing divide-and-conquer structure and greater reasoning depth to complex problems. |
Just install and run—SpineCodex automatically inherits your existing Codex configuration and works out of the box.
npm install -g @spinejit/spine-codex@latest
spine-codex
SpineJIT transparently manages context with no user intervention; after sending your first task, run /spine-tree to confirm it is working. Experimental Spine Spawn (spine_spawn) and Memory Projection (spinetree_memory_projection) can be enabled via /experimental.
To use SpineCodex with the official Codex Desktop app, first quit Codex Desktop, then download the launcher for your platform:
start-spinecodex-desktop.cmd
and start-spinecodex-desktop.ps1
in the same folder, then run the .cmd file.start-spinecodex-desktop.command.
If needed, make it executable with chmod +x start-spinecodex-desktop.command.The launchers use the native SpineCodex backend installed by npm and enable the Spine tree UI automatically.
| Feature | Purpose |
|---|---|
Spine Spawn (spine_spawn) |
At any node, concurrently spawn multiple differentiated branch agents that inherit its history, recursively collaborate, and converge through cache-friendly context reuse. |
Memory Projection (spinetree_memory_projection) |
Project compiled Node Memory into inspectable Markdown. |
Run /experimental to enable Spine Spawn or Memory Projection, then save and
start a new conversation.
Across three long-horizon coding benchmarks, SpineCodex delivers stronger outcomes: 1.89× resolved tasks at 27% lower total cost on SWE-Milestone, +10.80pp average score on ProgramBench, and +9.2pp mean score on FrontierSWE.
Long-horizon software development · 80 milestones · GPT-5.6 · sol high
| System | Resolved | Total cost |
|---|---|---|
| BaseCodex | 9 | $764.18 |
| SpineCodex | 17 | $556.46 |
1.89× resolved tasks at 27% lower total cost.
Whole-repo program reconstruction · Random sample: 50 of 200 tasks · GPT-5.6 · Sol high · conservative cost estimate
| System | Avg. score | Tasks scoring >95% | Cost |
|---|---|---|---|
| BaseCodex | 62.55% | 2/50 | $188.12 |
| SpineCodex | 73.35% | 7/50 | $475.10 |
+10.80pp average score and 3.5× high-scoring tasks.
Ultra-long-horizon coding · 9-task evaluation · GPT-5.6 · high · estimated API cost per trial
| System | Mean score | Best score | Cost |
|---|---|---|---|
| BaseCodex | 33.5 | 37.9 | $20.16 |
| SpineCodex | 42.7 | 46.8 | $37.29 |
+9.2pp mean score and +8.9pp best score.
Agent Morphogenesis: Each task shapes its own context and execution through just-in-time context-tree compilation and recursive subagent scaling.
TL;DR: SpineJIT replaces the live suffix of a context with shorter memory, while keeping the prefix unchanged so it can continue to hit the prompt cache.
To control this suffix replacement precisely, SpineJIT is implemented as a just-in-time compilation and context-mapping pipeline:
$$ \text{context messages} \rightarrow \text{Spine tokens} \rightarrow \text{SpineTree (ParseStack)} \rightarrow \text{new context} $$
The pipeline has two main stages.
SpineJIT treats a context $C$---a message list, or simply a sentence whose characters are messages---as a stream to compile.
At each sampling boundary, it turns newly appended messages and control events into Spine tokens and updates a live LR(0) ParseStack:
SpineJIT uses four token kinds:
$$ \Sigma_{\mathrm{Spine}} = {\mathrm{Message},\ \mathrm{Open},\ \mathrm{Close},\ \mathrm{SpineSpawnNode}} $$
Message represents a raw context item. Open, Close, and
SpineSpawnNode are special tokens emitted by SpineJIT at the corresponding
sampling boundaries.
$$ \begin{aligned} \mathrm{SpineTree} &\to \mathrm{Nodes}\ \mathrm{End} \ \mathrm{Nodes} &\to \mathrm{Node} \mid \mathrm{Nodes}\ \mathrm{Node} \ \mathrm{Node} &\to \mathrm{Message} \mid \mathrm{SpineTreeNode} \ \mathrm{SpineTreeNode} &\to \mathrm{Open}\ \mathrm{Nodes}\ \mathrm{Close} \mid \mathrm{SpineSpawnNode} \end{aligned} $$
End is only the logical end of a session; a live session never emits it.
Therefore, the ParseStack is the live SpineTree, and the reduction Open Nodes Close -> SpineTreeNode turns a closed subtree into one node.
In short, SpineJIT uses LR(0) JIT compilation to map context $C$ to a Spine Tree $PS$:
$$ PS = \mathrm{compile}(C) $$
The structured SpineTree can now be mapped into a shorter context while preserving its stable prefix. For ParseStack $PS$, define:
$$ C' = f(PS) = \prod_{i=0}^{n} h(PS[i]) $$
$$ h(X) = \begin{cases} \prod_{x \in X} h(x), & X = \mathrm{Nodes} \ \mathrm{raw}(X), & X = \mathrm{Message} \ \mathrm{memory}(X), & X = \mathrm{SpineTreeNode} \ \mathrm{spine\_node\_desc}(X), & X = \mathrm{Open} \end{cases} $$
Here, $\prod$ means ordered concatenation.
The mapping is deliberately small:
Message keeps its original content through $\mathrm{raw}(X)$.SpineTreeNode is replaced by its shorter $\mathrm{memory}(X)$.Open is represented by a concise
$\mathrm{spine\_node\_desc}(X)$, helping the LLM delimit the currently live
Spine node.As parsing progresses, completed work in the context suffix is reduced into a SpineTreeNode and then projected as memory. Earlier context remains unchanged:
$$ \mathrm{prefix} \cdot \mathrm{suffix} \longrightarrow \mathrm{prefix} \cdot \mathrm{memory} $$
This is th