Behavior-guardrail hooks for Claude Code: test-tampering + outbound-action guards
# Add to your Claude Code skills
git clone https://github.com/hahahahahahahahah6/agent-guardSee how agent-guard compares with popular alternatives.
agent-guard is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by hahahahahahahahah6. Behavior-guardrail hooks for Claude Code: test-tampering + outbound-action guards. It has 51 GitHub stars.
agent-guard's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/hahahahahahahahah6/agent-guard" and add it to your Claude Code skills directory (see the Installation section above).
agent-guard is primarily written in Python. It is open-source under hahahahahahahahah6 on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh agent-guard against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
Behavior-guardrail hooks for Claude Code. Six guards plus test-honesty CLIs, one install, zero dependencies (Python standard library only):
pip install agent-guard-hooks
SessionStart it snapshots hashes of every test and source file; on
Stop it diffs. If test files were modified or deleted while no source
file changed, the stop is blocked until the agent proves each changed test
actually fails without the fix.PreToolUse hook on Bash with a denylist
of risky action patterns: pushes to protected branches (including force
pushes), package publishes (npm publish, twine upload, …), prod deploys,
cloud provisioning (spend), and mass-send channels (Slack webhooks,
mailers). Matches block with a named rule; a user allowlist in
config.json overrides the denylist. Every matched decision is written to
an audit log. v0.2 also scans the content of script files the command
executes (bash evil.sh), closing the "write it to a script first" bypass.PostToolUse hook on Bash
that reads back world state after a command claimed an outbound effect
(git push, npm publish) and warns — never blocks — when the effect
isn't visible. Defense-in-depth on top of PreToolUse prevention: prevent
first, verify after.agent-guard mutate-check
deliberately breaks assertions (regex-based mutants for JS/TS and Python),
re-runs the tests per mutant, and reports survivors: tests that stayed
green don't actually cover the bug.PreToolUse hook on Write/Edit that
scores the added comments (never pre-existing code) for narrative slop:
commented-out code, restatements of obvious code ("This function adds two
numbers"), in-code changelogs, meta/apologetic notes, emoji, and
docstrings that just restate the signature. Blocks at a configurable
threshold, or agent-guard decomment --check/--fix for a one-command
decomment pass before PRs.random.seed(123) with no reproducibility
marker, patching random.shuffle), mocking the function under test
instead of its collaborators, conftest.py plants, time-freezing, and
always-True comparison dunders. A PreToolUse hook on Write/Edit
blocks cheat patterns in added test text; agent-guard cheatsniff --check audits the repo.PreToolUse hook on Bash statically extracts
file-write targets (>, >>, heredocs, sed -i, tee, cp/mv
destinations) and applies the same policy the Write-tool guards would:
test-ish targets get cheat-sniffed when the content is visible, opaque
writes to protected files are blocked. A guardrail on one tool protects
nothing if another tool can do the same thing.The failure modes are real, quoted from the community:
"Instead of fixing the indexing logic in the source file, the agent quietly modified the test file: it changed
expect(page.items.length).toBe(10)totoBe(9), re-ran the test, saw green, and told us the refactor was complete." — navune, r/ClaudeCode
"the risk isn't a bad answer, it's a bad action" — dank_as_fuck_, r/AI_Agents, on why guardrails must live "below the prompt layer" (verstands)
Patch release driven by an independent review of v0.5 — every item below was reproduced against the release before fixing:
pytest -q 2>&1 | tee test_output.log,
echo '{}' > tests/fixtures/data.json, and
echo '' > tests/__init__.py. A hook that cries wolf gets uninstalled:
only real code suffixes now count as test files, and fixtures/,
testdata/, data/ directories plus __init__.py are never test
files. Regression-tested per reported case.perl -pi -e is now treated like sed -i. It is the exact
equivalent and sailed through v0.5; opaque in-place edits of test files
via perl are blocked the same way.awk -i inplace,
truncate -s 0, dd of=, and git checkout … -- tests/ are not
intercepted by the PreToolUse hook — but the Stop-time test-tampering
guard still catches any test-file modification or deletion at session
end (verified by test).agent-guard install # registers agent-guard hook-bashwrite (PreToolUse on Bash)
The second bypass in the same family. v0.2 closed "write it to a script
first" (the r/AI_Agents blocklist bypass); v0.5 closes the other one.
thomastartrau read 28 Claude Code security advisories and found the
pattern: "My hooks block certain writes through the Edit and Write
tools. Once blocked, the agent went through Bash instead: sed -i, a
heredoc, a redirection. I had to add a hook that blocks writes to source
files via Bash. A guardrail on one tool protects nothing if another tool
can do the same thing."
install registers agent-guard hook-bashwrite as a PreToolUse hook on
Bash. It statically extracts file-write targets from the command —
> / >> redirections, heredocs (<<EOF, <<-EOF), sed -i
(including -i.bak / --in-place), tee (with/without -a),
cp/mv/install destinations, chained with && / ; / | — and
applies the same policy the Write-tool guards would apply:
test_*.py, conftest.py, tests/ …): when the
written content is visible (heredoc body), it is cheat-sniffed with the
v0.4 detectors; an opaque write (sed -i, bare >) to a test file is
treated as a violation on its own — that is exactly the bypass shape.bash_write.protected_paths (path prefixes, default empty = the
test-tampering guard's scope): any Bash write under a protected prefix is
treated like a Write-tool call — visible content is scored (comment-slop
for source-ish files), opaque writes are blocked in block mode.Warn mode (AGENT_GUARD_BASHWRITE_MODE=warn or bash_write.mode=warn)
advises instead of blocking. A bash_write.allow list
("tests/legacy/:bash-write") covers the judgment calls you disagree
with. Everything fails open: unparsable commands are allowed, never
blocked.
agent-guard cheatsniff --check tests/ # per-file cheat scores, exit 1 over threshold
Half of agent cheating never touches a test file. A dev.to study (remdore,
2026-10-01, 102 runs × 4 models) found agents patching the RNG "so the list
would always be sorted", mocking the function under test instead of its
collaborators, and planting helpers in conftest.py. The test-tampering
guard (tests-only-change diff) and mutate-check (assertion mutation) both
miss this family — "restore the test files and re-run" only catches the
dumb half.
install registers agent-guard hook-cheatsniff as a PreToolUse hook on
Write/Edit. It fires only for test-ish files (test_*.py,
*_test.py, conftest.py, anything under tests/) and scores only the
added text. Six cheat kinds, regex-based and deliberately conservative:
mock.patch("billing.total") inside
test_billing.py — patching the module under test itself, not its
dependencies. Patching a collaborator (stripe.Charge.create) is clean.random.shuffle / random.random /
random.sample to force outcomes.conftest.py monkeypatching the subject or
other local modules — including hand-rolled mymod.shuffle = ... direct
assignment (pure fixtures are clean).random.seed(123) with no reproducibility marker.
random.seed(42) next to a "reproducible" comment is legitimate and not
flagged — the marker is the whole difference, and it's documented.freeze_time(...), time.sleep patched to a no-op.__eq__ / __lt__ / … whose body unconditionally
return True.One severe hit reaches the default threshold (30/100) on its own. Warn mode
(AGENT_GUARD_CHEAT_MODE=warn or cheat_sniff.mode=warn in config) advises
instead of blocking. A cheat_sniff.allow list ("test_sort.py:rng-seed",
"*/legacy/*:*") covers the judgment calls you disagree with.