by Shiyao-Huang
Open survey and evidence map for AI agent evolution, self-evolving agents, memory, skills, harnesses, benchmarks, and agent-swarm systems.
# Add to your Claude Code skills
git clone https://github.com/Shiyao-Huang/awesome-agent-evolutionGuides for using ai agents skills like awesome-agent-evolution.
Last scanned: 6/5/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-06-05T08:07:58.101Z",
"npmAuditRan": true,
"pipAuditRan": true
}一份面向 AI Agent 自进化研究与实践的开放 Survey:帮你判断一个系统是真能从反馈中改进,还是只是在 demo 里看起来聪明。
中文主入口 | 英文版 | 在线网站 | 论文 PDF | Evolve-AGI Index | 项目报告
GitHub Topics: agent-evolution, self-evolving-agents, self-evolution, self-improvement, ai-agent, llm-agent, agent-swarm, memory-system, skill-library, harness-engineering, benchmark.
GitHub topic 收录证据(2026-06-05):GitHub Topic Indexing Readiness 已验证远端 topics、GitHub Search 和 topic 页面均返回本仓库;如果网页 topic 页短暂滞后,按 GitHub search/API 作为更新鲜证据。

想判断一个 AI Agent 是不是“真自进化”,先问五件事:改了什么、为什么改、谁验证、是否保留、能否回滚。
| 读者 | 你会得到什么 |
|---|---|
| 研究者 | 一套从分类、方法、系统、评估到未来路线图的 Survey 主线。 |
| 工程师 | 判断一个 agent 项目是否具备可验证反馈、可审计记忆、评估框架和回滚能力。 |
| 产品/投资/行业读者 | 区分真实能力积累、刷榜、演示热度和治理成熟度。 |
| 内容/教育读者 | 获得带证据入口的选题地图:项目、论文、趋势、痛点、图谱和长尾主题页面。 |
| 你是谁 | 先读什么 | 你能带走什么 |
|---|---|---|
| 第一次来 | 什么才算自进化 AI Agent | 一张判断表:改了什么、谁验证、如何保留、能否回滚。 |
| 想理解机制 | 五类进化回路 | 把规范到执行、搜索、评估器、反思/记忆、种群/归档分开看。 |
| 想比较项目 | 代码自我改进 Benchmark Matrix 和 项目报告 | 不被 star 或 demo 带偏,先看 evaluator、archive、lineage 和限制。 |
| 想查趋势 | 2026 Star 抓取试点 和 Value LSH 证据分诊 | 区分历史热度、当前动量、启发式分诊和证据修复队列。 |
英文读者现在可以从 /en/ 进入定义、五类回路、代码 benchmark、项目证据、报告状态、Value LSH、资料库覆盖、Survey 快照、研究图谱、证据图、增长试点、Evolve-AGI worksheet、论文和博客导读。长尾文章正文和许多 report 页仍是中文优先或 source-tracing 页面,因此不宣称完整翻译 parity。
flowchart LR
RAW["原始证据<br/>GitHub / 论文 / 博客 / 社交"] --> PROC["加工证据<br/>分析 / 研究 / 项目"]
PROC --> SURVEY["Survey 综合<br/>五类进化回路 + 痛点 + benchmark"]
SURVEY --> SPARK["核心洞察<br/>受控自进化"]
SPARK --> EAI["Evolve-AGI Index<br/>证据加权估计"]
EAI --> PAPER["论文核心<br/>论点 + 贡献 + 路线图"]
SURVEY --> SITE["网站 + 图谱 + 报告"]
本轮是新的 hourly public metadata 修复包,不再沿用 2026-07-05 01:38 +0800 的上一个 authenticated packet 作为唯一前台口径。抓取链路本身是可用的,但仍按“先重试 live GitHub API,失败时明确回退到上一个 authenticated packet”的规则更新,避免伪造 freshness。
| 仓库 | 这轮状态 | 为什么重要 | 证据状态 |
|---|---|---|---|
| china-qijizhifeng/agentic-Harness-engineering | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 684 -> 685; updatedAt 2026-07-04T16:09:34Z -> 2026-07-04T18:39:02Z. | 它是“harness 本身可进化”的最直接锚点。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| NousResearch/hermes-agent | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 209,108 -> 209,194; forks 38,140 -> 38,173; issues 8,568 -> 8,600; PRs 17,135 -> 17,193; commits 14,393 -> 14,427; pushedAt 2026-07-04T16:07:28Z -> 2026-07-04T23:37:23Z; updatedAt 2026-07-04T17:34:42Z -> 2026-07-04T23:34:25Z. | 它回答“可用产品型 agent 长什么样”这个核心问题。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| stanford-iris-lab/meta-harness | No public metadata delta was observed relative to the previous authenticated packet at 2026-07-05 01:38 +0800. | 它是 outer-loop harness search 的最干净参考样本。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| rohitg00/agentmemory | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 24,550 -> 24,554; PRs 187 -> 188; updatedAt 2026-07-04T17:38:26Z -> 2026-07-04T22:56:40Z. | 它回答“长期记忆如何跨 Codex / Claude Code / Hermes / OpenClaw 持续积累”。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| pinchbench/skill | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 1,260 -> 1,261; updatedAt 2026-07-03T08:54:33Z -> 2026-07-04T22:18:22Z. | 它是 skill、memory、benchmark 三条线交叉的 evaluator substrate。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| lsdefine/GenericAgent | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 13,283 -> 13,284; PRs 71 -> 72; updatedAt 2026-07-04T15:35:56Z -> 2026-07-04T18:30:53Z. | 它是“不要预装技能,而是让技能树生长”的 self-evolving 极简路线。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| openclaw/openclaw | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 381,702 -> 381,730; forks 80,024 -> 80,030; issues 3,619 -> 3,604; PRs 3,489 -> 3,469; commits 63,965 -> 64,040; pushedAt 2026-07-04T17:39:05Z -> 2026-07-04T23:38:38Z; updatedAt 2026-07-04T17:39:18Z -> 2026-07-04T23:38:49Z. | 它是“agent 是否真的能给人用”的产品运行时锚点。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| obra/superpowers | Relative to the previous authenticated packet at 2026-07-05 01:38 +0800: stars 246,063 -> 246,190; forks 21,819 -> 21,830; issues 149 -> 148; PRs 173 -> 174; updatedAt 2026-07-04T17:35:24Z -> 2026-07-04T23:39:35Z. | 它把可复用技能和工程方法论这条线接进了自进化公开证据链。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| EvoMap/awesome-agent-evolution | No public metadata delta was observed relative to the previous authenticated packet at 2026-07-05 01:38 +0800. | 它帮助我们检查公开叙事是否比普通 awesome list 更有证据密度。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| uid4oe/insight-swarm | No public metadata delta was observed relative to the previous authenticated packet at 2026-07-05 01:38 +0800. | 它是“shared knowledge graph 替代中心 orchestrator”的 swarm 概念锚点。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
| desplega-ai/agent-swarm | No public metadata delta was observed relative to the previous authenticated packet at 2026-07-05 01:38 +0800. | 它把用户要求的 agent-swarm 主线补进了公开证据链。 | [KNOWN] Authenticated GitHub API;未做本地运行/benchmark 复核。 |
GitNexus 证据链本轮受限:node .gitnexus/run.cjs status 可读但显示索引 stale;query -r awesome-evolution-workspace-cleanup 当前被 LadybugDB storage-version mismatch 阻断(file:///Users/copizzah/.local/lib/node_modules/gitnexus/dist/core/lbug/pool-adapter.js:325),所以本轮不声明 GitNexus 关系证据已刷新。
一句话:本项目的核心洞察,是把 Self-Evolving AI Agents 从“自我改进的故事”变成“可审计的改进系统”。
三句话:一个系统只有在反馈中改变自己的 prompt、memory、tool policy、workflow、code、weights 或 population,并且保留可验证证据时,才进入自进化范围。Survey 背后的全部资源现在按同一个问题重排:哪个对象在变,什么信号驱动它变,谁阻止它变坏。Evolve-AGI Index 是这次重排后的工作型证据表,用来暴露 benchmark、闭环、迁移和治理证据是否足够,而不是给领域下最终分数。
五句话展开:
[UNVERIFIED]。| 序号 | Survey 结论 | 对读者的意义 | 证据入口 |
|---|---|---|---|
| 1 | 自进化是受控系统过程,不是 demo 标签。 | 读任何项目先问“改了什么、谁验证、怎么回滚”。 | paper abstract, ch1 intro |
| 2 | Benchmark 是选择压力,也是风险源。 | 分数提高不等于能力积累;要看隐藏测试、迁移、成本、失败候选。 | ch5 evaluation, survey ch5 |
| 3 | 记忆、技能、评估框架是核心基础设施。 | 不要只看模型层;可审计记忆、可安装技能和评估器才决定长期可用性。 | ch7 painpoints, agent-swarm evolve |
| 4 | 五类进化回路比项目名更稳定。 | 新项目可以按机制归类,而不是被营销词牵着走。 | survey methods, method taxonomy |
| 5 | Evolve-AGI Index 只能作为工作型证据表。 | 它把 benchmark、闭环、证据、迁移、可运行、动量、治理七个信号拆开看,不能当领域标准。 | Evolve-AGI Index, trend snapshot |
| 6 | 用户真正关心信任边界。 | 产品价值来自可靠、透明、可控、低成本,不来自“更自主”的口号。 | survey ch7, site survey |
| 7 | 失败候选和负结果是资产。 | 没有被拒补丁、回归记录和 lineage,无法判断系统是否真的会进化。 | ch8 future, survey spark analysis |
一句话:Evolve-AGI Index 是本 Survey 的工作型证据指数原型,用来检查这个领域的证据成熟度,不是 AGI 终局能力评分,也不是单个项目的最终排名。
EAI = Σ(signal_score × signal_weight)
| 信号 | 权重 | 为什么进入核心 |
|---|---|---|
| Benchmark 表现 | 18% | 自进化必须接受实测;但 benchmark 不能单独决定成熟度。 |
| 闭环强度 | 20% | 没有可变对象、反馈、选择和保留机制,就没有自进化。 |
| 证据链可信度 | 18% | 原始材料、分析、model card 和论文附录必须互相能追溯。 |
| 迁移与验证 | 14% | 只在一个公开测试上涨分,不能证明能力积累。 |
| 实现可获得性 | 12% | 能运行、能复用、能审计,才有工程价值。 |
| 领域动量 | 10% | 新项目和社区动量是趋势信号,但不能覆盖证据质量。 |
| 治理准备度 | 8% | 自修改系统必须有安全边界、日志、回滚和时间戳信心。 |
权重是当前 Survey 的 editorial/proposed weights,用来把不同证据放在同一张可讨论的表里;它们还不是经同行验证的领域标准,也没有完成敏感性分析或置信区间估计。
**Data Snapshot / 数据快照:**Evolve-AGI trend 使用的是 2026-06-01 趋势输入快照:93 个 strict evolution repos、200 个 broad evolution repos、239 条 trend public-report records。仓库治理和网站覆盖使用 docs/indexes/master-index.md 的最新生成口径:684 个 classified GitHub repositories、292 个 analyzed
awesome-agent-evolution is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by Shiyao-Huang. Open survey and evidence map for AI agent evolution, self-evolving agents, memory, skills, harnesses, benchmarks, and agent-swarm systems. It has 177 GitHub stars.
Yes. awesome-agent-evolution passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/Shiyao-Huang/awesome-agent-evolution" and add it to your Claude Code skills directory (see the Installation section above).
awesome-agent-evolution is primarily written in JavaScript. It is open-source under Shiyao-Huang on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh awesome-agent-evolution against similar tools.
No comments yet. Be the first to share your thoughts!