by bladedevoff
Local proxy that learns your app's typed LLM decisions and answers them with a Laya head. Jev and OpenAI compatible.
# Add to your Claude Code skills
git clone https://github.com/bladedevoff/stuntdSee how stuntd compares with popular alternatives.
stuntd is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by bladedevoff. Local proxy that learns your app's typed LLM decisions and answers them with a Laya head. Jev and OpenAI compatible. It has 51 GitHub stars.
stuntd's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/bladedevoff/stuntd" and add it to your Claude Code skills directory (see the Installation section above).
stuntd is primarily written in Python. It is open-source under bladedevoff on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh stuntd against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.

stuntd is a local, self-hosted proxy that records the typed decisions your app already makes and
learns to answer them itself. It speaks the Jev System One protocol (POST /v1/systemone, with
choice, score and noul questions), the OpenAI Chat Completions API and the Anthropic
Messages API, distils each decision
site into a small head on a frozen Laya encoder,
and serves that answer locally with calibrated confidence, handing anything it is unsure about back
to the provider.
Try it in the browser: the Hugging Face Space puts zero-shot Laya and the heads stuntd trained side by side on three demos, no install.
Three ways to run it:
typesafe-sdk at stuntd and the base Laya
checkpoint answers your typed questions.Status: early release, v0.1. Every number below traces to a demo in this repository or to a source listed at the end.
typesafe-sdk client keeps working while you wait.PreToolUse gate answering "may this
command run" in about 50 ms, without sending every command you type to a provider.
examples/devtools/ ships the hook.pip install "stuntd[train]"
stuntd config init
stuntd serve
config init writes a commented stuntd.toml with everything at its default, which means no
OpenAI upstream and no Jev provider. serve says so and loads the base checkpoint:
stuntd listening on http://127.0.0.1:8787
no upstream: only the Jev routes are served
serving jev locally with convaiinnovations/laya
Then change the base URL in your client and nothing else:
from typesafe_sdk import Choice, Noul, TypeSafeClient
QUESTIONS = {
"category": Choice(
instructions="Which part of the product is this ticket about?",
criteria={"billing": "Payments and refunds.", "bug": "Something is broken."},
),
"needs_human": Noul(instructions="Does a support agent have to read this?"),
}
with TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8787") as client:
answer = client.system_one(
state={"subject": "Charged twice", "body": "Two identical charges on Friday."},
questions=QUESTIONS,
)
print(answer.choices["category"].choice, answer.nouls["needs_human"].noul)
Anything that speaks the protocol plugs in the same way, by changing one base URL wherever the
client exposes it: typesafe-sdk, JevRouter, the local-server request of fast-jev-compaction, and
the browser and game agents built on the same SDK.
The key is ignored unless you set [jev] require_key = true. To train a head on this daemon you
need labelled rows: either put the paid Jev API in front of it ([jev] upstream) and let it teach,
or write the rows yourself and stuntd import them.
pip install "stuntd[train]"
stuntd config init
stuntd serve --upstream https://api.openai.com
--upstream is the provider origin only, without /v1: your client's own path is appended to it
unchanged. In the client, point base_url at the daemon:
import os
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8787/v1", api_key=os.environ["OPENAI_API_KEY"])
Everything keeps working exactly as before. A request whose response_format (or single tool) asks
for an object made only of enum, boolean or number fields is a typed decision, and stuntd records
it with the provider's answer. Anything else is relayed untouched, streaming included.
An object with several fields, such as {category, urgency, needs_human}, learns one head per
field. Each field is its own site, named <site>.<field>, and a request is answered locally only
when every field's head is live and confident. If one is not, the provider answers and every field
is recorded.
Once a site has a few hundred captures:
stuntd status # what has been recorded, per site
stuntd train # one head per site with enough examples
stuntd report # holdout agreement, ECE, the operating point
stuntd enable <site> # let that site answer locally
pip install "stuntd[train]"
stuntd config init
stuntd serve --upstream https://api.anthropic.com
Point the SDK at the daemon; the key and the anthropic-version header are forwarded as they are:
import anthropic
client = anthropic.Anthropic(base_url="http://127.0.0.1:8787")
A non-streaming POST /v1/messages that asks for a JSON schema through output_config.format, or
for a single tool with an input_schema, is a typed decision on the same terms as the OpenAI path.
The same decision asked through either provider lands on the same site and shares one head. A
streaming request, or a request with tools and a free-text reply, is relayed untouched. Claude
Code and the Agent SDK reach the daemon through ANTHROPIC_BASE_URL.
The local answer is checked against the anthropic SDK's Message model in the tests, but no live
Anthropic key has been run through stuntd.
What a trained head does, measured on a laptop RTX 5060 with convaiinnovations/laya as the base
checkpoint, 3000 generated rows per site and 24 epochs (5890 rows for Snake), with
training.cache_encoder at its default. "Zero-shot" is the same daemon before training, answering
from the base checkpoint alone.
| demo | questions | zero-shot | trained heads | answered locally | p50 per request |
|---|---|---|---|---|---|
| Snake | 1 choice | 0.70 average score | 11.40 average score | 99.6% | 21.1 ms |
| support | choice, score, noul | 69.5 / 34.5 / 66.5% | 99.5 / 85.5 / 97.5% | 90% | 90.3 ms |
| banking | 2 choice | 89.5 / 30.0% | 100.0 / 72.5% | 78% | 75.7 ms |
| devtools | choice, noul | 45.0 / 84.5% | 97.0 / 100.0% | 98% | 58.0 ms |
Every decision site of the three text demos, measured against its teacher on held-out requests before and after the head was trained:
Snake is scored by what the answers do rather than by how many of them are right:
Five of the seven sites in the three text demos reach a holdout agreement of 0.95 or better. The
demos run at the default target_agreement of 0.99, which picks each head's operating point rather
than being a score it reached. The two that fall short, support's urgency at 0.935 and banking's
risk at 0.937, stay in the demos anyway: see limitations.
Latency, size and cost:
| measured | |
|---|---|
| one head answer on CUDA | p50 22 ms, 26 ms in process |
| the same head on CPU | 60 to 62 ms |
| one live answer through the daemon, over TCP | 28 ms |
| what the relay itself adds when the provider answers | 2.1 ms |
| Jev API round trip, through OpenRouter | p50 376 to 389 ms |
| Jev price | $0.042 per 1M input tokens, output free |
| training one head: 5890 rows, 24 epochs | 186 s |
| training three heads: 3000 rows each, 24 epochs | 362 s |
| a trained head on disk | about 50 MB, fp16 |
The first four rows were measured during development on an RTX 5060 laptop running Windows. They come from no published benchmark, and there is no link to give for them. The two training rows are the Snake and support demos of this repository, on the same laptop with the encoder cache on; the devtools demo measured the same 24-epoch run at 730 s with the cache off against 231 s with it on, 3.2x.
Why train at all, rather than run Laya zero-shot. From jevbench, n=500 per dataset, accuracy:
| dataset | labels | Laya zero-shot | Jev | fine-tuned DistilBERT |
|---|---|---|---|---|
| sst2 | 2 | 92.0 | 95.4 | 91.0 |
| agnews | 4 | 90.6 | 84.3 | 91.0 |
| banking77 | 77 | 38.2 | 76.4 | 88.0 |
Zero-shot Lay