by phuryn
Bug Hunt Bench: 105 real bugs in two production repos, frontier coding models (GPT-6, Claude, Grok, Gemini, DeepSeek...) find and fix them in their own CLI, graded blind. Live leaderboard + every receipt.
# Add to your Claude Code skills
git clone https://github.com/phuryn/bug-hunt-benchGuides for using ai agents skills like bug-hunt-bench.
See how bug-hunt-bench compares with popular alternatives.
bug-hunt-bench is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by phuryn. Bug Hunt Bench: 105 real bugs in two production repos, frontier coding models (GPT-6, Claude, Grok, Gemini, DeepSeek...) find and fix them in their own CLI, graded blind. Live leaderboard + every receipt. It has 55 GitHub stars.
bug-hunt-bench's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/phuryn/bug-hunt-bench" and add it to your Claude Code skills directory (see the Installation section above).
bug-hunt-bench is primarily written in JavaScript. It is open-source under phuryn on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh bug-hunt-bench against similar tools.
No comments yet. Be the first to share your thoughts!
Unlocks once the catalog security scan passes (runs nightly).
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
105 real bugs, hidden in two production codebases. Frontier coding models — GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, GLM and more — get one round per repo in their own agentic CLI (Codex CLI, Claude Code, Grok CLI, Antigravity CLI) to find and fix what they can. Every diff is graded blind against a withheld answer key. The score counts planted bugs only.
Live board: bughunt.productcompass.pm · Method, caveats and definitions · Raw data · Findings by wave
If the numbers save you a benchmark run of your own, star this repo — that is what keeps the bench findable, and new models are added as they ship.

Updated Sep 13, 2026 · 77 scored runs · 28 models · 38 of 105 bugs have never been fixed by any model.
Current leader: GPT-6 Astra at max effort — 48 / 105 (24/45 on repo 1, 24/60 on repo 2).
Best run per lab: OpenAI: GPT-6 Astra (max) 48 · Anthropic: Fable 5.1 (max) 43 · Meta: Muse Spark 1.3 (max) 33 · Alibaba: Qwen3.8-Max (max) 28 · xAI: Grok 4.6 (xhigh) 27 · DeepSeek: DeepSeek V4.1 Flash (max) 24 · Google: Gemini 3.7 Flash (high) 22 · Moonshot AI: Kimi K3 (default) 21 · Z.ai: GLM-5.3 Flash (max) 19 · Tencent: Hy4 Preview (default) 18 · OpenAI (open weights): gpt-oss-20b (default) 0
| # | Model | Harness | Effort | Fixed /105 | Repo 1 /45 | Repo 2 /60 | Extras | Wall | Cost | Date |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Codex CLI | max | 48 | 24 | 24 | 45 | 79 min | $31.21 | 2026-09-04 |
| 2 | GPT-6 Astra | Codex CLI | xhigh | 43 | 23 | 20 | 53 | 59 min | $24.22 | 2026-09-04 |
| 3 | Fable 5.1 | Claude Code | max | 43 | 19 | 24 | 11 | 73 min | $77.55 | 2026-09-01 |
| 4 | GPT-5.6 Sol | Codex CLI | max | 42 | 19 | 23 | 40 | 164 min | $69.61 | 2026-08-01 |
| 5 | GPT-5.6 Sol | Codex CLI | xhigh | 39 | 18 | 21 | 59 | 126 min | $52.75 | 2026-08-28 |
| 6 | GPT-6 Astra | Codex CLI | high | 35 | 19 | 16 | 40 | 40 min | $20.60 | 2026-09-04 |
| 7 | GPT-6 Astra | Codex CLI | medium | 34 | 19 | 15 | 33 | 28 min | $15.78 | 2026-09-05 |
| 8 | GPT-5.6 Sol | Codex CLI | high | 34 | 13 | 21 | 28 | 67 min | $33.92 | 2026-07-31 |
| 9 | GPT-5.6 Luna | Codex CLI | max | 33 | 17 | 16 | 31 | 86 min | $1.80 | 2026-07-31 |
| 10 | Muse Spark 1.3 | Muse Code / Meta API | max | 33 | 14 | 19 | 16 | 102 min | $19.83 | 2026-09-13 |
| 11 | Fable 5.1 | Claude Code | high | 33 | 15 | 18 | 7 | 36 min | $41.52 | 2026-09-01 |
| 12 | GPT-5.6 Terra | Codex CLI | max | 32 | 16 | 16 | 45 | 160 min | $27.98 | 2026-08-27 |
| 13 | GPT-5.6 Sol | Codex CLI | medium | 29 | 13 | 16 | 24 | 48 min | $15.77 | 2026-09-06 |
| 14 | Fable 5.1 | Claude Code | low | 29 | 13 | 16 | 6 | 33 min | $27.27 | 2026-09-02 |
| 15 | Fable 5.1 | Claude Code | xhigh | 29 | 13 | 16 | 4 | 60 min | $40.42 | 2026-09-10 |
| 16 | Fable 5 | Claude Code | max | 29 | 12 | 17 | 5 | 57 min | $104.49 | 2026-08-01 |
| 17 | Qwen3.8-Max | Claude Code / Alibaba API | max | 28 | 13 | 15 | 6 | 103 min | $26.29 | 2026-09-11 |
| 18 | GPT-6 Astra | Codex CLI | low | 27 | 18 | 9 | 25 | 32 min | $11.69 | 2026-09-05 |
| 19 | Grok 4.6 | Grok Build CLI (ACP) | xhigh | 27 | 8 | 19 | 16 | 41 min | $16.96 floor | 2026-08-28 |
| 20 | Opus 5 | Claude Code | max | 27 | 13 | 14 | 2 | 60 min | $51.33 | 2026-08-01 |
| 21 | Qwen3.8-Flash | Claude Code / Alibaba API | max | 26 | 13 | 13 | 7 | 97 min | $1.81 | 2026-09-11 |
| 22 | Opus 5 | Claude Code | xhigh | 26 | 14 | 12 | 3 | 50 min | $59.59 | 2026-08-28 |
| 23 | DeepSeek V4.1 Flash | Claude Code / DeepSeek API | max | 24 | 14 | 10 | 7 | 43 min | $1.08 bill | 2026-09-10 |
| 24 | Opus 5 | Claude Code | medium | 24 | 11 | 13 | 4 | 30 min | $34.77 | 2026-08-27 |
| 25 | Fable 5 | Claude Code | high | 24 | 9 | 15 | 3 | 31 min | $68.07 | 2026-07-26 |
| 26 | Qwen3.8-Flash | Claude Code / Alibaba API | low | 23 | 11 | 12 | 7 | 98 min | $1.37 | 2026-09-11 |
| 27 | GPT-5.6 Luna | Codex CLI | xhigh | 23 | 10 | 13 | 53 | 136 min | $2.50 | 2026-08-28 |
| 28 | Grok 4.6 | Grok Build CLI (ACP) | high | 23 | 7 | 16 | 16 | 33 min | $15.73 floor | 2026-08-28 |
| 29 | Grok 4.6 | Grok Build CLI (ACP) | medium | 22 | 9 | 13 | 11 | 26 min | $5.80 floor | 2026-08-28 |
| 30 | Gemini 3.7 Flash | Gemini CLI (retired) + model-pinning gateway | high | 22 | 8 | 14 | 4 | 97 min | $8.43 | 2026-08-24 |
| 31 | Fable 5.1 | Claude Code | medium | 21 | 8 | 13 | 3 | 23 min | $17.46 | 2026-09-10 |
| 32 | Kimi K3 | Claude Code / OpenRouter | default | 21 | 4 | 17 | 6 | 108 min | $25.27 | 2026-07-26 |
| 33 | Opus 5 | Claude Code | high | 21 | 11 | 10 | 6 | 37 min | $38.77 | 2026-07-26 |
| 34 | GPT-5.6 Terra | Codex CLI | xhigh | 20 | 9 | 11 | 29 | 57 min | $8.98 | 2026-08-28 |
| 35 | Gemini 3.8 Flash | Antigravity CLI | high | 20 | 7 | 13 | 6 | 30 min | $9.78 | 2026-09-02 |
| 36 | Muse Spark 1.3 | Muse Code / Meta API | xhigh | 20 | 8 | 12 | 7 | 73 min | $15.58 | 2026-09-13 |
| 37 | DeepSeek V4.1 Flash | Claude Code / DeepSeek API | high | 19 | 9 | 10 | 4 | 26 min | $0.31 bill | 2026-09-10 |
| 38 | GLM-5.3 Flash | Claude Code / Z.ai API | max | 19 | 10 | 9 | 4 | 58 min | $0.94 | 2026-09-12 |
| 39 | Muse Spark 1.3 | Muse Code / Meta API | high | 19 | 8 | 11 | 10 | 64 min | $10.81 | 2026-09-13 |
| 40 | GLM-5.3 | Claude Code / Z.ai API | max | 19 | 9 | 10 | 2 | 40 min | $15.93 | 2026-09-10 |
| 41 | Hy4 Preview | Claude Code / OpenRouter | default | 18 | 9 | 9 | 2 | 68 min | $3.13 bill | 2026-08-28 |
| 42 | GPT-5.6 Terra | Codex CLI | high | 18 | 7 | 11 | 17 | 28 min | $5.89 | 2026-08-27 |
| 43 | Grok 4.5 | Grok Build CLI (ACP) | high | 17 | 5 | 12 | 10 | 28 min | $8.50 floor | 2026-08-06 |
| 44 | Muse Spark 1.2 | Claude Code / Meta API | xhigh | 17 | 6 | 11 | 12 | 36 min | $13.99 | 2026-08-06 |
| 45 | Ox Alpha (stealth) | Claude Code / OpenRouter | default | 16 | 8 | 8 | 3 | 59 min | free | 2026-08-25 |
| 46 | GLM-5.3 Flash | Claude Code / Z.ai API | high | 16 | 9 | 7 | 4 | 62 min | $0.94 | 2026-09-13 |
| 47 | DeepSeek V4-Pro | Claude Code / DeepSeek API | max | 16 | 8 | 8 | 4 | 37 min | $1.89 bill | 2026-09-10 |
| 48 | Gemini 3.7 Flash | Antigravity CLI | high | 16 | 4 | 12 | 2 | 23 min | $6.36 | 2026-08-28 |
| 49 | GPT-5.6 Terra | Codex CLI | medium | 15 | 4 | 11 | 6 | 20 min | $3.87 | 2026-08-27 |
| 50 | Grok 4.6 | Grok Build CLI (ACP) | low | 15 | 6 | 9 | 9 | 18 min | $4.14 floor | 2026-08-27 |
| 51 | Qwen3.8-27B | Claude Code / Alibaba API | xhigh | 15 | 6 | 9 | 2 | 47 min | $6.09 | 2026-09-12 |
| 52 | Opus 4.8 | Claude Code | max | 15 | 6 | 9 | 3 | 109 min | $52.07 | 2026-09-10 |
| 53 | DeepSeek V4-Flash | Claude Code / OpenRouter | default | 14 | 6 | 8 | 0 | 48 min | $1.52 bill | 2026-08-01 |
| 54 | GPT-5.6 Luna | Codex CLI | high | 13 | 5 | 8 | 22 | 64 min | $0.57 | 2026-07-31 |
| 55 | GLM-5.3 Flash | Claude Code / OpenRouter | default | 13 | 6 | 7 | 4 | 57 min | $0.79 bill | 2026-08-27 |
| 56 | DeepSeek V4-Pro | Claude Code / DeepSeek API | high | 13 | 4 | 9 | 1 | 27 min | $1.08 bill | 2026-09-10 |
| 57 | GPT-5.6 Luna | Codex CLI | medium | 9 | 5 | 4 | 5 | 15 min | $0.33 | 2026-08-27 |
| 58 | GLM-5.3 Flash | Claude Code / Z.ai API | low | 9 | 7 | 2 | 4 | 29 min | $0.42 | 2026-09-13 |
| 59 | Muse Spark 1.3 | Muse Code / Meta API | medium | 9 | 4 | 5 | 8 | 62 min | $10.02 | 2026-09-13 |
| 60 | Sonnet 5 | Claude Code | high | 9 | 1 | 8 | 4 | 33 min | $15.12 | 2026-07-26 |
| 61 | Opus 4.8 | Claude Code | high | 9 | 2 | 7 | 1 | 35 min | $19.35 | 2026-07-26 |
| 62 | Sonnet 5 | Claude Code | max | 9 | 3 | 6 | 3 | 61 min | $24.04 | 2026-09-10 |
| 63 | GPT-5.6 Luna | Codex CLI | low | 4 | 0 | 4 | 1 | 6 min | $0.10 | 2026-08-27 |
| 64 | gpt-oss-20b | Claude Code / OpenRouter | default | 0 | 0 | 0 | 0 | 4 min | $0.11 bill | 2026-09-12 |
| 65 | gpt-oss-120b | Claude Code / OpenRouter | default | 0 | 0 | 0 | 1 | 4 min | $0.13 bill | 2026-09-12 |
Extras are real, unplanted defects a model fixed on the way; they are counted and never added to the score. Costs are token estimates at published list rates unless tagged bill (an actual invoice or credits delta) or floor (a reconstructed lower bound). default effort means the serving path had no working effort dial; ran lower marks a run whose CLI quietly replaced the requested tier. Wall clock is repo 1 plus repo 2 agent time, dependency install excluded.
12 superseded re-runs stay in the CSVs and on the live board but are left off this table.
A benchmark of AI coding agents on the job they are actually sold for: reading an unfamiliar, real codebase and fixing what is wrong with it. Not a puzzle set, not a single-file task, not a synthetic repo.