Deepdive skill for Claude Code — 12-phase research pipeline: plan-review gate, parallel sub-agent search, claims-ledger triangulation with dissent protection, relevance × authority evidence filter, multi-angle red team, four-layer citation verification. 105 blocks, 29 channels, 460+ stat sources, 47 APIs, 1072 verified endpoints.
# Add to your Claude Code skills
git clone https://github.com/Socialpranker/deepdiveLast scanned: 9/2/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-09-02T08:25:47.055Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}deepdive is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by Socialpranker. Deepdive skill for Claude Code — 12-phase research pipeline: plan-review gate, parallel sub-agent search, claims-ledger triangulation with dissent protection, relevance × authority evidence filter, multi-angle red team, four-layer citation verification. 105 blocks, 29 channels, 460+ stat sources, 47 APIs, 1072 verified endpoints. It has 393 GitHub stars.
Yes. deepdive passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/Socialpranker/deepdive" and add it to your Claude Code skills directory (see the Installation section above). deepdive ships a SKILL.md manifest, so compatible agents can discover and load it automatically.
deepdive is primarily written in Python. It is open-source under Socialpranker on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh deepdive against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Многошаговое исследование под вопрос или решение. Источник = файл, отчёт как Q&A, тезисы атомарны и пере-используемы.
Прямая просьба о ресёрче; сравнение N институций/продуктов/методологий/рынков; материал под стратегию, доклад, статью; проверка гипотезы внешними данными; «как устроен X», «карта области Y».
НЕ применять: быстрая фактоверка → отвечай напрямую · N конкурентов по фиксированной матрице → competitive-teardown · Anthropic SDK / Claude API → claude-api · брейншторм без данных → brainstorming/grill-me · ответ уже в проекте → сначала grep.
| Режим | Источников | Суб-агентов | Когда |
|---|---|---|---|
| shallow | 5–7 | 0 | первичная навигация, тема знакома, low-stakes |
| medium | 12–18 | 2–3 | нетривиальная тема, среднее решение |
| deep | 25–35+ | 4–5 | high-stakes решение, стратегия |
Объяви режим в начале с обоснованием. После Genre+Plan объяви model routing одной строкой (фазы по моделям + estimated cost + как перебить: «всё на opus» / «cheap mode»). Детали — model_routing.md.
До reframing (опционально — нет файлов, иди дальше): определи целевую папку → есть — перечисли содержимое, похожий slug ⇒ спроси «это update?» → прочитай CLAUDE.md/CLAUDE.local.md и memory/MEMORY.md, учти в reframing. Цель: не дублировать сделанное.
Плюс кросс-прогонная вики — python scripts/wiki_query.py --topic "<вопрос>": прошлые утверждения, уже оценённые источники (credibility не пересчитывать), открытые противоречия между прогонами. Непогашенное противоречие идёт в plan.md исследовательским вопросом. См. wiki.md.
Куда сохранять (не хардкодь): (1) research-папка из CLAUDE.md или существующая research/ · 06_Деск-ресёрч/ · docs/research/ · notes/research/; (2) иначе по типу проекта — манифест (pyproject.toml/package.json/Cargo.toml/go.mod) → research/, только документы → 06_Деск-ресёрч/; (3) не git-репо или пусто → ~/deep-research/<slug>/. Путь покажи ОДИН раз, дальше пиши молча.
Slug: латиница, цифры, дефисы («Postgres logical replication vs CDC» → postgres-replication-vs-cdc). Неочевиден — покажи в начале фазы 2.
Детали фаз — workflow.md, модель на фазу — model_routing.md. Здесь — что фаза обязана оставить после себя.
opus/high] — переписать вопрос; Decision Spec (решение глагол+объект+срок / потребитель→его следующий шаг / ≥1 if-then вилка «покажет X → делаю A»; ни одной вилки ⇒ честный даунгрейд в shallow); 2–4 опровергаемые гипотезы; medium/deep — персоны охвата (STORM) и router по типу вопроса. См. question_reframing.md.sonnet/medium] — жанр (qa/explainer/decision/landscape/validation/custom) + набор блоков, подтвердить одной строкой. См. genres.md, blocks/INDEX.md.opus/medium] — plan.md по шаблону §0–16 из workflow.md: HEADER → SCOPE → STRUCTURE → EXECUTION → TRACKING. Несущее: acceptance criteria, гипотезы, risk register, subtopic↔blocks mapping (least-to-most для многошаговых вопросов), sourcing strategy §12, opposition queries, stop-criteria.
3.5. Capability Discovery [sonnet/low] (deep — обязательна) — audit env vars, подтемы → доступные API, fallback на awesome-lists. См. capability_discovery.md.
3.7. Plan-review gate [sonnet/low] (shallow — skip) — единственная human-in-the-loop точка ПЕРЕД дорогой Фазой 4: показать сжатый план (вопрос, решение, жанр, гипотезы, каналы, стоп-критерий, routing). deep — ЖДАТЬ явного «Ок»; medium — soft. Плюс скаут-пасс (deep — рекомендуется): 3–4 Explore на haiku ищут непокрытые подвопросы, а не источники; выход — правки plan.md, ноль записей в sources/. См. plan_gate.md.sonnet/medium; sub-agents: haiku web/api, sonnet academic/long-source] — (4.0) Source Dispatch по матрице → plan.md §12; количественный подвопрос ⇒ primary-канал registry/API. (4.1) medium/deep — general-purpose суб-агенты параллельно, каждому свой диапазон id (s01-s09, s10-s19…) и своя ось поиска, не только подтема; shallow — главный поток. (4.2) Fetch, дедуп с замером overlap_rate. (4.3) Агент сам пишет sources/NN_slug.md, в главный поток — только index-строки. После раунда 1 — snowball. Loop: goal-check → bounded deviation → circuit breaker (2 раунда без нового ⇒ стоп, остаток в Open Questions). Окно раунда пересобирается, не накапливается (medium/deep): state.md ПЕРЕЗАПИСЫВАЕТСЯ перед каждым раундом (## Known статусами со ссылками · ## Gaps · ## Next, ≤6 КБ), планирование по нему, а не по транскрипту. См. source_dispatch.md, subagents_v2.md.haiku/low] — claims.csv (схема колонок — source_scoring.md). triangulated ⟺ ≥3 источника И ≥2 типа И ≥2 корня (root:) И ≥2 пути (discovery_path:); иначе single-type/single-root/single-path, потолок medium. Без primary — потолок medium; caveat (vendor/self-reported/disputed:sNN) — потолок medium, disputed без арбитра → low. Защита меньшинства: непогашенный dissent от Primary/credibility ≥ 4 ⇒ contested независимо от большинства, обе позиции в отчёт. Gap-волна на не-triangulated, max 2 круга, иначе data-insufficient. См. source_scoring.md.
5.5. Evidence-фильтр: relevance × authority [sonnet/low] (medium/deep — обязательно) — фильтр на ВХОДЕ синтеза. Relevance: пара (claim, source) → Correct/Ambiguous/Incorrect по дословным цитатам → relevant-only цитаты в evidence/CN.md; claim без relevant-источника → data-insufficient или до-поиск. Authority (несущие пары: claim в memo/F1/F9, ИЛИ с числом, ИЛИ источник единственный корень, ИЛИ caveat ≠ -): «вправе ли ЭТОТ источник утверждать ЭТО» → qualified/unqualified-for-this-claim/unknown → .verify/authority.json. unknown — карантин: не единственная опора, не high. См. evidence_filter.md.
5.7. Сверка с вики [sonnet/low] (medium/deep — обязательно) — claims.csv прогона против кросс-прогонной вики, до синтеза: расхождение с прошлым ресёрчем обязано попасть в отчёт, заметить его в Фазе 7 значит заметить поздно. wiki_pair.py build собирает пары и сам закрывает всё, что не требует суждения; остаток идёт в .verify/wiki_pairs.json → вердикт из четырёх (same-claim-agree/same-claim-conflict/different-claim/unknown) → wiki_pair.py record. unknown — карантин, не конфликт. Подтверждённый конфликт = обе позиции в отчёт, потолок medium, без арбитра contested. Срез по MAX_ADJUDICATED называть вслух. Пороги прескрина и правила record — wiki.md.opus/high для deep, sonnet/high для medium] — outline.md (section | block | claims из plan.md §8/§11 и фактического claims.csv) → собрать <date>_<genre>.md секция за секцией по outline, под каждую только её claim_id и её evidence/CN.md, не весь пул → числа в numbers.csv (verbatim/derived/share; у derived — formula+inputs) → финал «it depends» запрещён (рекомендация однозначная или условная по вилкам) → claim ledger → враждебные роли параллельно как general-purpose: R1 Skeptic, R2 Contrarian, R3 Gap-hunter, R4 Исполнитель, R5 Адвокат меньшинства → триаж severity → ОДИН раунд ремедиации HIGH → memo.md (рекомендация, вилки, 3 числа с [sNN]+as_of, риск, next actions, строка Урезано: — сработавший circuit breaker или даунгрейд вслух; иначе Урезано: —) → финал. Finder ≠ fixer. Гейт: shallow=R1 инлайн, medium=R1+R2+R4 (+R5 при dissent), deep=все пять. См. adversarial_pass.md, synthesis_outline.md, source_scoring.md (numbers.csv).
6.5. Verify [haiku/low] (medium/deep — обязательно) — четыре оси, вердикты в .verify/<ось>.json: liveness (check_citations.py); faithfulness (entailment claim⊨цитата по парам из evidence/CN.md → SUPPORTED/PARTIAL/UNSUPPORTED); qualifier preservation (F1/memo.md/Z12 против строк claims.csv → PRESERVED/BROADENED/SCOPE-DROPPED/UNTRACEABLE); construct provenance (именованные фреймворки/«законы»/термины против evidence/+sources/ → sourced/author-construct/unsourced; unsourced в memo.md/F1/F9 блокирует finish). Чинится отчёт, не ledger. Header F10 несёт все четыре оси плюс строку независимости источников; без него отчёт не «готов». См. runtime_verification.md.
6.9. Экспорт отчёта [haiku/low] (medium/deep) — uv run scripts/build_report.py <run>: HTML + PDF + DOCX из одного источника. Фигуры — только из numbers.csv, ```mermaid из E13/M9 — в «Схема N» через mmdc, [sNN] резолвятся в приложение из sources.csv. Битая ссылка или потерянная сноска роняют сборку. memo.md остаётся отдельной страницей. См. references/report_export.md.sonnet/medium] (medium/deep) — entities/numbers/hypotheses/topic-markers из отчёта в refresh_targets.md: точка входа для будущих update. Блок Z11 в blocks/close.md. Затем — четыре шага сбора наблюдений роя (collect_observations.py → update_priors.py → promote_candidates.py --track → изредка --write): порядок и флаги — references/swarm_postprocess.md, читать перед первым вызовом.opus/high, главный поток] (всегда, в shallow — 1 вилка) — отчёт не обсуждается, а исполняется: показать memo.md и провести пользователя по вилкам по одной. Исходы: принято (решение + next action + дата) / blocked (после 1 целевой gap-волны) / deferred. Артефакт application.md (любой status) + строка в ~/.claude/research/applications_ledger.csv. См. decision_walkthrough.md.Лимита на WebSearch/WebFetch нет.
Стоп когда: все гипотезы подтверждены/опровергнуты ≥3 разнотипными источниками либо помечены «данных мало» · прошёл и разобран ≥1 целевой поиск оппозиции («X criticism / counter-evidence / problems with X») · покрыты 4+ типа источников · последние 3–5 источников не дают нового.
Не стоп когда: источники противоречат (копай за причиной) · все одного типа · есть сильный контр-аргумент без разбора · оппозицию не искали.
Тупик: третий подряд поиск даёт источники total < 8 ⇒ стоп, в Open Questions «литература слабая», предложи интервью/эксперимент.
<slug>/
├── plan.md # Фаза 3 (+ changelog §16, notes §15)
├── state.md # Фаза 4 — окно раунда, ПЕРЕЗАПИСЫВАЕТСЯ каждый раунд (medium/deep)
├── sources.csv # индекс источников с оценками
├── claims.csv # Фаза 5 — claim-ledger
├── numbers.csv # Фаза 6 — реестр чисел отчёта (medium/deep)
├── outline.md # Фаза 6 — карта section → block → claim_id (medium/deep)
├── figures.csv # Фаза 6.9 — фигуры, только по num_id из numbers.csv (medium/deep)
├── sources/NN_slug.md # один файл = один источник (метаданные + цитаты)
├── evidence/CN.md # Фаза 5.5 — relevant-only цитаты под claim (medium/deep)
├── findings/FN_*.md # атомарные тезисы (опц., для крупных)
├── refresh_targets.md # Фаза 7 (medium/deep)
├── memo.md # Фаза 6 — decision-меморандум (всегда)
├── application.md # Фаза 8 — вердикт по вилкам + status (всегда)
├── .verify/ # I/O-контракт: один producer, много consumers
│ ├── authority.json # 5.5 — qualified/unqualified/unknown
│ ├── wiki_pairs.json # 5.7 — пары на адъюдикацию (medium/deep)
│ ├── wiki_ingest.json # finish-up — квитанция записи в вики (всегда)
│ └── citations|faithfulness|qualifiers|constructs.json # 6.5, по оси на файл
├── diffs/<date>_delta.md# дельты режима update
├── <YYYY-MM-DD>_<genre>.{html,pdf,docx} # Фаза 6.9 — собранный документ
└── <YYYY-MM-DD>_<genre>.md # финал: qa|explainer|decision|landscape|validation|custom
Кросс-прогонный слой лежит ВНЕ прогона — ~/.claude/research/wiki/, один на все ресёрчи. См. wiki.md.
Отдельный _changelog.md не создаётся — он в plan.md §16. Шаблоны: sources/NN.md, claims.csv — source_scoring.md; отчёт — genres.md + blocks/; findings/FN.md — Z6 в blocks/close.md.
python scripts/build_sources_csv.py --research-dir <root>/<slug> (единый источник колонок) · python eval/check_citations.py --research-dir <root>/<slug> --json --out <root>/<slug>/.verify/citations (без --out файл уйдёт в eval/output/ и gate его не найдёт).
0.2. Компиляция в вики — python scripts/wiki_ingest.py --research-dir <root>/<slug>, детерминированно и на любой глубине. Пишет квитанцию .verify/wiki_ingest.json, без неё phase-gate красный. Изредка python scripts/wiki_lint.py.
0.5. Числа — два прохода, --research-dir <root>/<slug> --strict: check_number_provenance.py (число без производителя; одно значение при разных корнях = ложная независимость) · check_number_arithmetic.py (пересчёт derived, доли к 100, производное число в memo без строки в numbers.csv).python scripts/validate_phases.py --research-dir <root>/<slug> --strict. Красный ⇒ фаза пропущена ⇒ вернись, доделай, перезапусти: не показывать путь, не писать резюме, не рапортовать «готово».memo.md (вход потребителя), затем отчёт.application.md.memory/ — предложи 1–3 кандидата (тезис + confidence + источники; авторитетный источник как [reference]).anthropic-skills:humanizer-ru — прогони им финальный отчёт (опционально).discover existing и reframing · Plan-review gate в medium/deep (для deep гейт без ожидания ответа = не гейт) · Фазы 5.5 и 5.7 и multi-angle red team в medium/deep · gap-волну · Фазу 8 «потому что и так ясно» — и не отвечать на вилки ЗА пользователя.as_of и корнями. Не трактовать unknown как «сойдёт» — карантин. Не молчать про срез по MAX_ADJUDICATED: непроверенные пары ≠ отсутствие противоречий. Не чинить противоречие выбором «более свежего».wiki_ingest по глубине.root: пустым и не копировать discovery_path: между источниками — это 3-е и 4-е условия триангуляции.overlap_rate в plan.md §15: совпадение это замер конформизма, а не подтверждение.state.md — он перезаписывается.triangulated строке с непогашенным dissent от Primary/credibility ≥ 4 (это contested) и не гасить dissent понижением credibility несогласного.unknown в authority как «сойдёт» — карантин. Не давать confidence выше medium без primary-источника. Не строить выводы на источниках с total < 8 и не оставлять утверждений без ссылки на sources/NN.md.origin_kind: unknown / chain_len ≥ 2 / нет data_as_of ⇒ не в memo.md/TL;DR/F9 и не high.numbers.csv с formula+inputs — оси 6.5 арифметику не проверяют.triangulated/contested claim вне outline.md.[sNN] или пометки «наша рамка».stat_sources//api_sources/.sources/NN.md (архив), не фильтровать по total вместо релевантности фрагмента к claim..verify/*.json и не пересчитывать в rubric/F10; пары брать из evidence/, не пересканировать sources/. Чинить отчёт, а не ledger.general-purpose с явным диапазоном номеров, не Explore (read-only, только разведка). Не запускать суб-агентов последовательно — только параллельно в одном сообщении.sources/ в один файл, не выводить результат только в чат.bash/curl. Единственный санкционированный fallback — scripts/fetch_source.py (Фаза 4.2): он читает robots.txt, санитайзит страницу от prompt injection и проставляет fetch_tier. Вывод ручного curl в sources/ не кладётся.fetch_source.py за средство против paywall и анти-бот-защиты: auth-wall и antibot для него — терминальный вердикт. Дальше — fallback-протокол channels.md или endpoint из api_sources/.numbers.csv, и не подбирать палитру фигур на глаз — она валидируется скриптом.memo.md в большой документ. Не считать DOCX форматом для чтения — читают PDF и HTML.update <slug> / «обнови ресёрч X» — дельта, не replay. Pre-flight: plan.md, refresh_targets.md (нет — сгенерируй по Z11), последний отчёт. Четыре категории дельты с date-фильтром от last_research_date: new entrants · entity diff · numbers refresh · adversarial trigger. Verified-no-change — тоже результат. Выход: diffs/<date>_delta.md; новый отчёт — только если дельта существенна (решает пользователь), старый получает status: superseded by …. Adversarial trigger HIGH ⇒ повторить только Фазу 6 на opus. Протокол — refresh_protocol.md.
Update тоже идёт в вики: wiki_ingest.py после дельты (квитанция обязательна) и wiki_pair.py build — расхождение НЕ по свежести и есть настоящая находка update'а.
Прогрессивная подгрузка: файл читается когда дошёл до фазы, не превентивно.
Базовые (читает любой medium/deep прогон): workflow.md (детали 13 фаз) · question_reframing.md (Фаза 1 + clarification-триаж) · plan_gate.md (Фаза 3.7 + скаут) · genres.md (6 жанров) · blocks/INDEX.md (106 блоков) · channels.md (29 каналов, query patterns, paywall fallbacks) · source_dispatch.md (обязательно перед launch суб-агентов) · model_routing.md · wiki.md (кросс-прогонный слой: чтение в discover existing, Фаза 5.7, запись в finish-up).
Условные — грузить, когда прогон дошёл до условия, а не заранее: capability_discovery.md и awesome_lists_registry.md — Фаза 3.5 (обязательна только на deep) · stat_sources/INDEX.md (33 категории) и api_sources/INDEX.md (47+ endpoints) — Фаза 4, когда подвопрос количественный или Source Dispatch ведёт в registry/API · refresh_protocol.md — только режим update.
По фазам: subagents_v2.md (4) · fetch_source.py + channels.md §«Bot-block ≠ paywall» (4.2, только когда WebFetch ответил «unable to fetch from …») · source_scoring.md (шкалы, provenance, claims-ledger, dissent, numbers.csv — 5–6) · evidence_filter.md (5.5) · synthesis_outline.md (6) · adversarial_pass.md (6) · runtime_verification.md (6.5) · report_export.md (6.9) · swarm_postprocess.md (после 7) · decision_walkthrough.md (8).
Блоки (по выбранному жанру): frame.md F1-F10 · explain.md E1-E14 · compare.md C1-C13 · map.md M1-M12 · validate.md V1-V10 · analyze.md A1-A13 · close.md Z1-Z12 · people.md P1-P7 · numbers.md N1-N8 · context.md X1-X7.
Stat/API источники (Фаза 4, точечно): stat_sources/core/*.md (14 cross-industry) · stat_sources/industries/*.md (19 отраслевых) · api_sources/. Читай INDEX, потом нужную категорию. Auth через env vars, ключи скилл не хранит; приоритет — free no-key API.
Stop ad-hoc Googling. Start documented investigation.
Docs · Install · How it works · Contribute
You: investigate the trade-offs between Postgres logical replication and CDC tooling
Claude: ✓ Reframed your question (3 hypotheses) + decision spec: what you'll do with the answer
✓ Picked genre: decision (comparison + validation)
✓ Wrote plan.md (17 sections)
✓ Checked your env: 4 APIs available, 2 fallback to HTML
✓ Launched 4 sub-agents across 12 channels
✓ Saved 23 sources to sources/ with quotes, provenance-checked
✓ Ran adversarial pass (3 counter-arguments + an "execute this" role)
✓ Report ready: research/postgres-replication-vs-cdc/2026-05-21_decision.md
✓ Walked through your decision forks — you picked CDC tooling, logged to application.md
A Claude Code skill that turns "research this topic" into a 13-phase pipeline with hypothesis testing, parallel sub-agent search, source triangulation, and adversarial review.
The output is a folder you can return to in a month. Every claim traces to a specific source file. The plan documents why you made every choice. No re-research needed.
New here? Start with the Quickstart — install → invoke → first result in ~5 min.
One-shot prompt → wall of text
Sources lost in chat history
No way to detect bias
No reuse next time
Generic Google results
Sources include... (vague)
17-section plan.md documents every choice
Each source = file with verbatim quotes
Mandatory adversarial pass + opposition queries
Atomic theses in findings/FN.md reusable
Every claim → [s12] link → specific quote
git clone https://github.com/Socialpranker/deepdive.git \
~/.claude/skills/deepdive
That's it. Now type any of these in a Claude Code session:
# Clone
git clone https://github.com/Socialpranker/deepdive.git
cd deepdive
# Package as .skill bundle
zip -r ../deepdive.skill . -x ".*" -x "*.zip"
# Upload via Claude.app → Settings → Skills → Add Skill
The 13-phase methodology is portable. Load SKILL.md + references/*.md into the LLM's context manually. Skip the sub-agent parts and use separate chat sessions per subtopic.
The skill runs 13 phases in order:
| Phase | Name | What happens |
|---|
| 1 | Reframing | opus / high | | 2 | Genre & block selection | sonnet / medium | | 3 | Plan | opus / medium | | 3.5 | Capability Discovery | sonnet / low | | 3.7 | Plan-review gate | sonnet / low | | 4 | Search | sonnet / medium | | 5 | Claims-ledger + triangulation | haiku / low | | 5.5 | Evidence filter | sonnet / low | | 5.7 | Wiki reconcile | sonnet / low | | 6 | Synthesis + multi-angle red team | opus / high | | 6.5 | Verify | haiku / low | | 7 | Refresh targets | sonnet / medium | | 8 | Decision walkthrough | opus / high |
Each phase runs on a model matched to its task — Opus where reasoning multiplies (1/3/6), Haiku for the parallel fan-out (4). The skill announces the routing and an estimated cost up front, once.
Every phase is transparent: you see what's happening, you confirm key decisions, and you get a folder you can return to. Before any search fires, the plan-review gate (3.7) shows you the reframing, hypotheses, genre, and channels and lets you approve or edit them — strictness scales with mode (deep waits for an explicit go-ahead, medium is a soft check, shallow skips it). Editing the plan before execution is the single highest-leverage step in the whole pipeline — Gemini Deep Research calls plan review its "biggest lever over output quality," and a wrong plan executed perfectly still produces a wrong report.
Reframing (1) doesn't just restate the question — a router classifies its profile (factual / multi-step / relational / comparative / landscape) and that classification picks the decomposition method: factual questions get flat independent subquestions, multi-step ones ("X given Y") get least-to-most leveling, comparative ones get a shared axis matrix with mandatory opposition queries per candidate. Picking the wrong decomposition for a question's shape is a silent failure mode — the router makes the choice explicit instead of defaulting to "flat parallel" for everything.
Phase 4 (Search) isn't a single pass — it's a bounded loop with three cheap safeguards so it doesn't quietly waste budget or silently give up:
met / partial / unmet with a one-line reason. This is what the expensive Opus evaluation reads instead of re-deriving the gap from scratch, and it's what targets the next round's dispatch.L1 → L2 instead of dispatched flat in parallel: L1 rounds run first, concrete facts they surface get carried forward, and L2 queries are launched already sharpened by that context. Independent subquestions still run flat.Scoring (5) doesn't stop at the usual Credibility/Recency/Bias rating — it also flags input-level skepticism: a source that measures its own product, self-reports a benchmark, or is directly disputed by another collected source gets a strict caveat: marker (vendor / self-reported / disputed:sNN) before the claim reaches claims.csv, not after synthesis has already built on it. A claim whose key number carries that marker is capped at confidence: medium (or low for an unresolved dispute) — the same rule shape as primary-first sourcing. Vendor benchmarks are the numbers that most often get quietly repeated as fact; catching them on the way in, not in the red team pass at the end, is the point.
For medium/deep depth, the pipeline runs two more machine-checked passes most one-shot research skips entirely:
SUPPORTED / PARTIAL / UNSUPPORTED verdicts to .verify/faithfulness.json. Citation fabrication is common enough industry-wide — the Tow Center found a >60% error rate in AI-generated citations — that checking for it, not just for dead links, is a real differentiator.None of this is enforced by discipline alone: scripts/validate_phases.py reads a finished run's mode: and checks that every phase mandatory for that mode actually left its file artifact (plan.md, claims.csv, evidence/, .verify/*.json, the dated report, ...). A skipped phase fails the check instead of silently passing — the model can't just claim "done." As of finish-up, this check is a blocker, not a suggestion: the skill won't report a research as done on a red gate, symmetrically to how a report isn't "done" without its verification header. sources.csv itself is now built the same deterministic way — scripts/build_sources_csv.py generates it from sources/NN.md frontmatter (with a --check mode for CI) instead of being assembled by hand each run.
A well-cited report that changes nothing is still a failure — the skill's answer to that is a decision spine running through the whole pipeline, not a bolt-on question at the end: