by Makson179
Have an agent setup in mind? Bello already supports it. Mix models, run agents in parallel, have them check each other’s work, and reduce unnecessary model calls and noisy tool output. Turn on only what you need.
# Add to your Claude Code skills
git clone https://github.com/Makson179/BelloLast scanned: 10/4/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-10-04T10:26:36.833Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}See how Bello compares with popular alternatives.
Bello is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by Makson179. Have an agent setup in mind? Bello already supports it. Mix models, run agents in parallel, have them check each other’s work, and reduce unnecessary model calls and noisy tool output. Turn on only what you need. It has 110 GitHub stars.
Yes. Bello passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/Makson179/Bello" and add it to your Claude Code skills directory (see the Installation section above).
Bello is primarily written in Python. It is open-source under Makson179 on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh Bello against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Bello runs a coding task through a coder and four optional parts: a runtime supervisor, a completion reviewer, an adversary, and a local log distiller. The coder, the supervisor, the reviewer, and the adversary each have their own model and reasoning effort, from a Codex or Claude Code subscription or from an API provider. The four components can be enabled or disabled independently.
In our tests, the Efficient Budget configuration used 66.3% less of the weekly Codex limit than Raw GPT-5.6 Sol XHigh without lowering average quality, and its mean score was 1.45% higher. Sol Ultra C+A raised the mean score over Raw Codex by 26.4%. The log distiller cut API-equivalent cost by 22% on Astra and by 14.8% on Luna.
Python 3.11 or newer, git, and macOS, Linux with the bubblewrap package, or
native 64-bit Windows 11 or Server 2022/2025 with the one-time sandbox
preparation from docs/windows.md.
pipx install bello
bello doctor
Then add the model sources you want, any one is enough:
codex login.pipx install 'bello[claude]' --force, then
bello runtime login claude-code. The extra bundles the official Agent SDK
and its CLI.bello runtime install once, then
bello runtime login <provider>.pipx install 'bello[log-distiller]' --force. The model
(599 MB) downloads on the first run with the distiller on. With subscription
Codex it also downloads a compatible Codex helper on Apple Silicon,
Windows x86_64, and Linux x86_64 (glibc 2.35+); see
docs/native-codex-selection.md.bello doctor shows what is ready, and bello update updates Bello. To run
Bello from inside your coding agent, add the plugin, which includes the
configuration advisor. For Codex:
codex plugin marketplace add AlexeyKulaev/Bello-codex-marketplace --ref main
codex plugin add bello@bello-marketplace
The same plugin has a Claude Code manifest in plugins/bello.
After installing the plugin, open your coding agent in the project that
contains task.md. The full start can be a short conversation:
You: Do you see the Bello plugin?
Agent: Yes. I can inspect the task, recommend a configuration, and run it with Bello.
You: Please recommend the best balance of price and quality for completing
task.md.Agent: I recommend Configuration X for
task.md. It offers the best balance of price, quality, and time for this task.You: Thanks. Please run
task.mdwith Configuration X and keep me updated on what is happening.
The agent shows the resolved configuration before launch. Bello then runs the
task and writes .supervisor/FINAL_REPORT.md with the result, changed files,
checks, and remaining risks.
You can also ask a stronger model to prepare an advisory PLAN.md, then have a
less expensive Bello configuration execute it. The coder receives the plan as
guidance, while completion review and adversarial testing remain independent.
https://github.com/user-attachments/assets/1a61ba61-444f-4347-b12b-e411d12c55c6
The coder implements the task in a disposable workspace and runs its own checks. Bello then hands the final patch back to your project. Around the coder, four parts can be switched on or off independently.
Confirmed problems go back to the coder, and the number of review and adversary rounds is a setting. Roles can use different providers, including API providers such as OpenAI, Anthropic, and OpenRouter, and the coder, the reviewer, and the adversary can also delegate work to subagents with their own models. For example, when these models are available in your accounts:
| Role | Model |
|---|---|
| Coder | GPT-5.6 Sol |
| Coder subagents, three in parallel | Claude Sonnet, GLM, GPT-5.6 Luna |
| Runtime supervisor | GPT-5.6 Terra |
| Completion reviewer | Claude Fable |
| Adversary | GPT-6 Astra |
| Log distiller | ModernBERT on your machine |
You do not have to pick all of this by hand. The advisor in the Codex and
Claude Code plugins reads the task and the repository and recommends one
complete setup for the priority you name. Every setting is also in
bello config.
Scores are ProgramBench completion scores unless a task has its own evaluator. Raw denotes the baseline without the Bello features being compared.
Three ProgramBench tasks (Solar, Samtools, and Rumdl), four models, each run
raw and through Bello: 24 runs in total. Sol ran at xhigh here.
Bello scored higher with every model. The gap is 13.6 points on Luna, 9.5 on Terra, 6.5 on Sol, and 6.0 on Astra. Solutions for all 24 runs.
Smart Execution runs independent tool calls in parallel and delivers results as they become ready. The agent can continue other necessary work while commands run; when it needs to wait, Bello handles the wait without repeated model calls to check progress. Unlike the log distiller, Smart Execution does not itself compress tool output.
We compared RAW with Smart Execution on Sonnet 5 and Astra, with three runs per model and setup, twelve in total. The table shows means; times cover the solver.
| Model | API-equivalent cost, RAW → SE | Solution time, RAW → SE | Score, RAW → SE |
|---|---|---|---|
| Sonnet 5 | $5.45 → $3.81 (−30.19%) | 42:26 → 21:13 (−50.00%) | 43.98% → 42.37% (−3.68%) |
| Astra | $9.37 → $6.41 (−31.61%) | 31:20 → 26:37 (−15.06%) | 58.88% → 58.44% (−0.74%) |
Both models cost less and finished sooner on average, with a small decrease in score. Every Smart Execution run cost less than every RAW run of the same model in this comparison.
Efficient Budget uses GPT-5.6 Luna at xhigh for the coder, the completion
reviewer, and the adversary, and Luna at high for runtime supervision with
cheap triage on. It allows one completion return before one adversary pass.
The baseline is Raw GPT-5.6 Sol XHigh.
Across four tasks with three runs per task and system, Budget used 5.27% of a weekly Codex limit against 15.63% for Raw, which is 66.3% less. Its mean score was 48.80% against 48.10%, 1.45% higher. Runs took longer: 1:49:44 on average against 28:27.
| Task | Raw score | Budget score | Change | Raw weekly limit | Budget weekly limit | Raw time | Budget time |
|---|---|---|---|---|---|---|---|
| Revive | 40.990% | 45.530% | +11.08% | 3.2757% | 1.0555% | 24:52 | 1:23:22 |
| JSONSchema | 56.821% | 54.673% | -3.78% | 3.0069% | 0.8138% | 23:34 | 1:20:43 |
| LightningCSS | 60.750% | 60.302% | -0.74% | 6.0281% | 2.2569% | 40:39 | 2:58:52 |
| Miller | 33.839% | 34.684% | +2.50% | 3.3190% | 1.1428% | 24:44 | 1:36:00 |
| All 12 + 12 | 48.100% | 48.797% | +1.45% | 15.6297% | 5.2690% | 28:27 | 1:49:44 |
Scores and times are means over three runs; weekly limit is the sum.