by qybaihe
mu (μ): a coding agent that thinks before it acts. A small, fast judge makes the routine calls, the big model does the work. Built on pi and AionUi.
# Add to your Claude Code skills
git clone https://github.com/qybaihe/muSee how mu compares with popular alternatives.
mu is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by qybaihe. mu (μ): a coding agent that thinks before it acts. A small, fast judge makes the routine calls, the big model does the work. Built on pi and AionUi. It has 374 GitHub stars.
mu's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/qybaihe/mu" and add it to your Claude Code skills directory (see the Installation section above).
mu is primarily written in TypeScript. It is open-source under qybaihe on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh mu against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
A coding agent makes hundreds of decisions per session that are not about the code: what stays in the context, whether a command is safe, whether a finding is worth telling another agent, when the work is done. Left to the big model, they cost tokens, latency and attention. Left to fixed rules, they are wrong too often. mu gives them to a judge: a small, fast model that answers one bounded question at a time, at 38 decision points in every turn. The big model keeps its attention for the work.
Early development. Pre-releases (0.1.x) are on npm and under Releases; its authors use it every day. Names, settings and formats may still change.
Contents: Quick start · A turn · Decision points · Judges · Measured · The hive · The board · Desktop app · Command line · Documentation · Contributing
Desktop app. Download the installer for macOS, Windows or Linux from Releases, open it, and paste an API key or sign in with a ChatGPT, Claude, Grok or Google subscription. Nothing else to install.
Command line.
npm i -g mu-agent # Node.js 22.19 or newer
mu setup # connect a model
mu # a session in the current directory
No judge key is needed to start: until you set one, the free Jev on OpenCode Zen answers. Every decision point acts on its verdicts from the start, and each verdict is logged (/status, mu ledger, the app's judgments tab). To only record what the judge would do, put {"modes": {"default": "shadow"}} in ~/.mu/agent/mu.json. More in Getting started.
you ──▶ input.preflight · task.frame · input.interjection
│
▼
model ──▶ tool call ──▶ tool.risk · tool.constraint · tool.approval ──▶ runs
▲ │
│ tool.injection web pages and MCP output: instructions aimed at the AI are withheld
│ tool.admission chunk by chunk: into the context, or archived behind a pointer
│ context.forget · context.compact when the context grows │
└─────────────────────────────────────────────────────────────────────┘
turn ends ──▶ turn.completion · turn.continue · turn.drift · turn.rewind · memory.applied · board.read · cache.warming
Every name is a decision point. Each one is asked as a short question about a small state; the answer changes what the model does next, never whether it asks you. Rules are the floor: a dangerous-looking command is caught by rules first, and the judge only vouches that you asked for it.
Each decision point is active (the default), shadow (asked and logged, changes nothing: for comparing judges before switching one on) or off, and each can name its own judge: jev, laya (local), classifier:<provider>/<model>, llm:<provider>/<model>, or a cascade such as laya,jev.
Input
| Decision point | Question | Effect |
|---|---|---|
input.preflight |
What kind of message is this, and how much thinking does it need? | A one-line hint to the model; optionally the turn's thinking level |
task.frame |
A new task, a hard constraint, a correction, a subgoal, or no change? | Only a change rewrites the task frame: goal, your constraints word for word with their source, acceptance criteria |
input.interjection |
A message arrives while the agent works: interrupt now, or after this step? | The turn is cut, or the message waits |
Context
| Decision point | Question | Effect |
|---|---|---|
skills.disclosure |
Which skills are relevant to this task? | Only those enter the prompt; the rest stay findable |
capability.disclosure |
Does this task need an installed pack or MCP server? | It is opened, and its process started, only then |
tool.admission |
Per chunk of a long tool output: does this matter now? | What matters enters the context; the rest is archived behind a pointer |
tool.admission.test-log |
In a test log, what is repetition? | Exact repeats are folded once, losslessly; optionally the judge selects from the rest |
context.forget |
Above a context threshold, which tool results are stale? | Each becomes a one-line tombstone in outgoing requests |
context.compact |
Keep or prune this passage? | Compaction by judgment; no summary is written |
memory.recall |
Which lessons apply to this task? | They are brought into the turn |
memory.capture |
Does this message correct the agent or set a rule? | It becomes a lesson |
memory.outcome |
After going in circles, did the way out deserve a lesson? | A lesson from the run, not from you |
memory.worth |
A lesson the model or a sub-agent proposes: useful again, a one-off, or known already? | Kept or dropped |
memory.merge |
The same as an existing lesson, more precise, or contradicting it? | No duplicates; the more precise one replaces the older |
memory.applied |
Were the recalled lessons followed this turn? | A lesson recalled often and never followed retires |
cache.warming |
Will you be back before the prompt cache expires? | The cache is refreshed, or left to expire |
Tools and safety
| Decision point | Question | Effect |
|---|---|---|
tool.risk |
A command the rules flag: did you ask for it? | Unsure means asking you |
tool.approval |
In the Jev approves mode: does the task clearly need this command, this change outside the project, this outside action, this sub-agent? | Only what it is sure of runs; the rest asks you |
tool.constraint |
Before a call that changes something: does it cross a constraint you stated? | The call is stopped |
tool.injection |
A web page, a search result or an MCP server's output, passage by passage: does it carry instructions aimed at the AI? | Those passages never reach the model; a note marks where each was |
files.locate |
Which files match what you describe? | Candidates ranked, instead of a string of greps |
judge.items |
The model's own yes/no question, about each of many items: files, log lines, findings | A probability per item, through the judge_items tool, instead of reading them all |
browser.step |
Observe, one judgment, act: what is the next operation, on which element? | The built-in browser moves one step |
review.triage |
For each finding of /review: does it change behaviour, and is it about this change? |
Findings ranked P0 to P3 |
diagnostics.delivery |
New language-server diagnostics after an edit: tell now, at the next pause, or never? | Errors reach the model; style warnings do not |
Turn
| Decision point | Question | Effect |
|---|---|---|
turn.drift |
Every few steps: does the work still serve the goal? | Rules catch circles; the judge catches drift |
turn.rewind |
The same failure again and again: is this approach a dead end? | Back to a checkpoint |
turn.completion |
The model says it is done: did anything verify that? | One nudge if not |
turn.continue |
The run ends on "Let me run the tests next", or on asking for a go-ahead on work you asked for: did it stop short? | Sent back to it, at most twice per message, never toward a step that is hard to undo |
output.drift |
While the model writes: does the tail of its output cross your constraints? | Experimental; corrected mid-stream |
goal.met |
In goal mode, when the big model gives no answer: is the condition met? | The fallback for /goal |
board.read |
Where do things stand, in multiple choice? | Feeds the plain-language board |
notify.routing |
An event such as the context budget: tell the model now, later, or never? | The model is told at the right time |
Teamwork
| Decision point | Question | Effect |
|---|---|---|
swarm.routing |
Which role, model tier and thinking level for this delegated task? | The sub- |