by joeseesun
让任何 Agent 调用 Codex 内置生图(MCP + CLI + Skill),内置小红书、视频封面、Mondo 海报技巧 · Codex image generation for any agent
# Add to your Claude Code skills
git clone https://github.com/joeseesun/qiaomu-codex-imagegenGuides for using ai agents skills like qiaomu-codex-imagegen.
See how qiaomu-codex-imagegen compares with popular alternatives.
qiaomu-codex-imagegen is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by joeseesun. 让任何 Agent 调用 Codex 内置生图(MCP + CLI + Skill),内置小红书、视频封面、Mondo 海报技巧 · Codex image generation for any agent. It has 54 GitHub stars.
qiaomu-codex-imagegen's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/joeseesun/qiaomu-codex-imagegen" and add it to your Claude Code skills directory (see the Installation section above). qiaomu-codex-imagegen ships a SKILL.md manifest, so compatible agents can discover and load it automatically.
qiaomu-codex-imagegen is primarily written in JavaScript. It is open-source under joeseesun on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh qiaomu-codex-imagegen against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
Codex ships an image generator that other agents cannot call. This package relays to it (codex app-server) and adds the craft that makes the pictures usable: a library of 24 templates (48 presets) distilled from 689 reference prompts, Mondo poster techniques, scenario presets, and a rule set for facts, text and identity.
suggest_directions, compose_prompt, generate_image, search_prompts, get_prompt, build_prompt, list_catalog; in Claude Code mcp__qiaomu-codex-imagegen__*): use them.node scripts/cli.mjs suggest|compose|generate|search|prompt|list (see --help).Extract: topic, deliverable (what it is for), ratio, text mode, facts that must be exact, references and the role of each (identity person, product real packaging, style, layout), must-keep, must-avoid.
none (no text), exact_short (only the short text you list), typeset_later (text-free base with room for the title).exact_short and give a complete short copy set: headline plus copy as lines separated by |. Take the lines from the user's facts. If they gave none, ask, or write clearly fictional sample lines and say so; never invent real prices, dates of real events, names, statistics.none fits portraits and photos (T09–T11), product macros that carry no information, and base images the user will typeset. Keep every string short and read the result: models misspell, especially long Chinese.Call suggest_directions with the topic, deliverable and any headline/text mode. It returns four or more directions with different visual mechanisms (typography-led, graphic structure, photographic, product macro, material/space, plus a Mondo artist option for poster-like work). Never mix two mechanisms into one image.
Present them to the user briefly, one numbered item each:
mechanism and the topic, not a copy of the lock text),references[].image thumbnail per direction when available, so your sentence matches the real look.Then ask which to make: one, several, "all", or "你定". Offer to exclude these and propose more (exclude).
Skip the question when the user already fixed the style ("用 Mondo 风格", "T03", "就这个风格") or said to go ahead ("直接出", "你定"). For "你定", pick the two best-fitting directions from different families and say which and why.
For each chosen template direction call compose_prompt with concrete variables:
missing_required non-empty → ask for exactly that. weak_variables → they still hold generic defaults, fill them first.structural_override → rewrite the prompt by hand: delete the base sentences the override replaces, keep no conflicting structure, then send the result as raw_prompt with the 避免:… line.copy as |-separated lines; when strings include dates or numbers, list them all so the whitelist matches exactly (or pass a full text_rule).Mondo directions skip this step: generate_image { prompt, style: <artist>, preset: <video-cover|poster|…> }.
Full rules and the canonical agent instruction: references/design-system/agent-guide.md.
Default one image per chosen direction, up to four, each its own generate_image call (they can run in parallel). Use count only for variants of the same direction. Say that it takes 30–180 s and spends Codex quota. out_dir goes inside the user's project when they name one; reference_images take absolute paths (identity photo, product shot, style reference, or the image to edit).
Series: lock template, type hierarchy, colour roles, texture placement and cross-over action; change one or two of scene, narrative focus, hue, module span. Pass the first image as reference_images for the rest.
Read each saved image and test it against acceptance_checks (the tool prints them): first-glance focus, relations that must hold, materials only where intended, exact text and identity, nothing invented. State pass or the specific failed relation.
With the local corpus pack installed, compare against the references: put the result next to one or two references[].image of the same template (suggest_directions lists them). If yours is plainly less designed (bare subject, no typographic hierarchy, no secondary layer), that is a failure too: complete the copy set or the missing layer and regenerate once.
On failure change only the sentence that carries that relation and regenerate once; do not stack "more premium" adjectives. Transparent regions in a result are flattened onto white automatically (transparent_background: true keeps them). Deliver absolute paths, the direction used, and the sidecar .json (exact prompt) next to each image. Only say "generated" for images you actually produced and looked at.
count above 3 only if asked.search_prompts / get_prompt read the local pack of 689 reference prompts (installed on the user's machine only). Use them to study how a mechanism is described; never send a corpus prompt unchanged as the user's prompt.prompt + preset + style (66 styles, 20 Mondo artists) when the user wants a quick image without the direction step.Reply with: absolute path(s), pixel size, direction (template/preset or Mondo artist) per image, assumptions made, acceptance result, and where the prompt sidecar is. On failure give the tool's error and the one fix to try (Codex login, shorten the prompt, check the reference path).
templates.json, routing.json; filled examplesreferences/styles.json (66 styles), mondo-artists.json (20), presets.json (13 scenarios)Copyright (c) 向阳乔木. X https://x.com/vista8, GitHub https://github.com/joeseesun/. Template library distilled from public-area samples of https://vip.xiaoxiaodong.ai/open-source; Mondo material from qiaomu-mondo-poster-design (MIT).
让任何 Agent 都能调用 Codex 内置的生图能力:每次生图先给出至少四种机制不同的风格方向,你选定后再出图。 内置 24 类场景模板(48 个预设)、20 位 Mondo 海报设计师风格,以及一套防止编造事实、乱写文字的规则。
Codex 的生图工具只在 Codex 里能用。这个项目把它接出来:一个 MCP 服务、一个命令行、一个 Agent Skill,共用同一套核心。Claude Code、Codex、Cursor,或者任何能执行命令的 Agent,都能说一句"做一张小红书配图"就拿到图片文件。
Use Codex's built-in image generation from any agent through MCP, a CLI or a skill. Every request first gets at least four divergent style directions (24 templates, 48 presets, 20 Mondo artists); you pick, it composes the prompt, generates, and checks the result.
对 Agent 说「给这期视频做一张封面」,它不会直接画,而是按下面的顺序走:
suggest_directions 给出 至少四个机制各异的方向,来自不同家族(文字主导、平面结构、摄影人物、产品近摄、材料空间,再加一个 Mondo 设计师方向),每个方向带风格锁、需要你补的信息和参考案例缩略图。compose_prompt 把变量填成完整提示词,同时给出避免项、验收条件、假设和缺失信息。带「结构替换」的预设会提示 Agent 改写,不会把冲突的句子硬拼在一起。你已经指定了风格(「用 Mondo 风格」「T03」)或说「直接出」时,会跳过选择这一步。
| 你说 | 它做 |
|---|---|
| 给这期视频做一张 16:9 视频封面 | 选 video-cover:缩略图逻辑、标题留白,用你指定的风格生成,返回文件路径和像素尺寸 |
| 做一张小红书配图,主题清晨读书 | 选 xiaohongshu:3:4 竖版,视觉重心偏上,顶部留标题位 |
| 用 Mondo 风格做《沙丘》海报,三个设计师各来一张 | 三个设计师风格并行生成,各自保存 |
| 把这张图的背景换成米色纸,主体不变 | 参考图改图(图生图) |
| 公众号头图、X 封面、朋友圈海报、书籍封面、专辑封面 | 各有预设比例和安全区 |
| 给我几个风格 / 帮我选个风格 | 四个以上机制不同的方向,附参考缩略图,选定后出图 |
| 找一下以前类似的案例 | 在本地 689 条参考提示词里检索,给出缩略图与原文 |
下面所有图片都是用这个工具调用 Codex 生成的,没有后期修图;活动、品牌、人物和文案全部是虚构的示例。每张图的提示词和参数在 docs/samples/prompts.json。
主题都是「咖啡店阅读月」。suggest_directions 给出的方向机制完全不同,所以出来的不是同一张图换个滤镜:
| 方向 | 机制 |
|---|---|
| 1 · T01 巨字与微小叙事 | 大字是场景的墙,水边的小读者提供尺度 |
| 2 · T02 物象破框 | 一枝咖啡树穿出细框,越界只发生一次 |
| 3 · T05 纸雕与织物地貌 | 书页的层叠被读成山河,微小读者坐在边缘 |
| 4 · T16 巨物微缩剧场 | 一杯拿铁放大成可进入的阅读小镇 |
| 5 · Mondo · Olly Moss | 两色丝网印,杯子的负空间里藏着一本书 |
| 类别与预设 | 样例 | 机制与文字 |
|---|---|---|
| T01 巨字与微小叙事预设 T01-1 旧刊慢场 | 大字是场景的墙,微小行动提供尺度;静止水平面被一条斜线打破。文字:exact_short:标题 + 副标题 + 信息行 | |
| T02 物象破框预设 T02-1 清透巨叶 | 框线建立秩序,实体穿过框线制造一次明确的越界。文字:exact_short | |
| T03 中央光隙预设 T03-1 清润春生 | 两侧巨大色域夹出中央通道,通道尽头的微小焦点表现希望。文字:exact_short | |
| T04 东方水墨编辑预设 T04-1 清透墨枝 | 水墨是空间与版式的一部分,宋体大字、透明色域和留白互相穿插。文字:exact_short | |
| T05 纸雕与织物地貌预设 T05-1 纸雕田垄 | 主题被转译成有材料厚度的地貌,微小物象为抽象层叠提供故事。文字:exact_short | |
| T06 民艺撞色招贴预设 T06-1 粗印民艺 | 两枚朴拙图符以冷暖强色对话,手写线把图像区和跳跃资讯区缝合。文字:exact_short:标题 + 活动名 + 时间地点 + 亮点 | |
| T07 童画与硬排版预设 T07-1 蜡笔展览 | 粗重现代字与松软儿童笔触互相挤压,细框提供第三种秩序。文字:exact_short | |
| T08 摄影与花境拼贴预设 T08-1 柔彩花境 | 刊头、摄影焦点、手绘前景、底栏构成连续层次,照片与插画必须相互遮挡。文字:exact_short | |
| T09 克制棚拍肖像预设 T09-1 柔灰专注 | 一主一辅的柔光塑造骨相,动作支撑与皮肤细节决定可信度。文字:none:纯肖像无字 | |
| T10 环境自然肖像预设 T10-1 暮阳天台 | 人物真实地处在环境中,动作先成立,光与景深再分离主体。文字:none:纯肖像无字 | |
| T11 婚礼与仪式肖像预设 T11-1 晴天珍珠 | 身份稳定是底座;薄纱、妆发和单一光线建立仪式感。文字:none:纯肖像无字 | |
| T12 食品触感近摄预设 T12-2 暖纸酥香 | 放大可食用的断面与真实触感,信息退到留白而不覆盖食物。文字:exact_short | |
| T13 饮品微距风味预设 T13-2 气泡切面 | 液体或果肉切面变成放大景观,气泡与水光体现风味而非替代产品事实。文字:exact_short | |
| T14 植物手绘产品广告预设 T14-1 清透水彩 | 真实产品与平面手绘并置,水彩路径环抱并贴附载体而非均匀铺花。文字:exact_short | |
| T15 科技轨道产品主视觉预设 T15-1 清透未来 | 产品是稳定中心,环形波纹和局部光线共同指向它。文字:exact_short | |
| T16 巨物微缩剧场预设 T16-1 食品小镇 | 巨物真正承担可进入的空间功能,微型人物的动作必须回应它。文字:exact_short | |
| T17 建筑制图编辑预设 T17-1 圆规素纸 | 真实结构件与有依据的几何线共享轴线,文字保持疏远而精密。文字:exact_short:竖排标题 | |
| T18 旅行与酒店框景预设 T18-1 温润拱窗 | 框景把观看者带入第二空间,说明围绕入口组织。文字:exact_short | |
| T19 文博材质巨像预设 T19-1 粗陶暗腔 | 材料孔隙与巨大的暗腔制造尺度,洁净空场和疏远文字维持静穆。文字:exact_short | |
| T20 会议与人物信息系统预设 T20-1 清爽斜带 | 重复单元共享节奏,一组定向斜切将照片与信息连接。文字:exact_short:四位虚构嘉宾的姓名与头衔 | |
| T21 九宫格日常手账预设 T21-1 绒线注释 | 统一矩阵中保留不同镜头密度,手绘标记必须回应格内内容。文字:exact_short | |
| T22 模块化演示视觉预设 T22-1 淡块作品集 | 标题与数据先建立信息层级,图形按页的叙事功能进入共同网格。文字:exact_short:标题 + 四项目录 | |
| T23 科普与商品信息卡预设 T23-1 友好研究卡 | 每组图形只解释一个命题,图与文的换位节拍代替装饰密度。文字:exact_short:标题 + 三条步骤 | |
| T24 日签与编辑纪念预设 T24-1 城市晨光 | 日期是锚点,实体照片是核心,一道问候跨过照片和纸面。文字:exact_short:标题 + 一句寄语 |
| 视频封面 16:9 | 竖屏封面 9:16 |
|---|---|
--preset video-cover --style saul-bass,标题压在左侧留白 |
--preset video-vertical --style kilian-eng,标题在上部安全区 |
| 公众号头图 2.35:1 | 小红书 3:4 |
|---|---|
--preset wechat-cover,标题在右侧天空 |
T12 食品近摄 + xiaohongshu 尺寸 |
| 原图 | 改后 |
|---|---|
| 一支铅笔 | --ref 原图,提示词:保持同一支铅笔,背景换成暖米色纸张,加柔和阴影 |
transparent_background: true 可保留)。前置条件:
node --version)codex --version 能运行)npx skills add joeseesun/qiaomu-codex-imagegen
Skill 会让 Agent 知道怎么选场景、风格、怎么写描述,并通过下面的 MCP 或自带 CLI 出图。
Claude Code:
claude mcp add --scope user qiaomu-codex-imagegen -- node /绝对路径/qiaomu-codex-imagegen/scripts/mcp-server.mjs
其他 MCP 客户端(Cursor、Codex 等),在配置里加:
{
"mcpServers": {
"qiaomu-codex-imagegen": {
"command": "node",
"args": ["/绝对路径/qiaomu-codex-imagegen/scripts/mcp-server.mjs"]
}
}
}
提供七个工具:
| 工具 | 作用 |
|---|---|
suggest_directions |
第一步:给出 ≥4 个机制各异的风格方向(免费、即时) |
compose_prompt |
把选定的模板和预设展开成完整提示词,附避免项、验收条件、缺失信息(免费) |
generate_image |
生成或改图;可用模板、raw_prompt(成品提示词)或简易的「描述 + 场景 + 风格」;返回路径、像素尺寸,并在图旁写 JSON |
search_prompts / get_prompt |
检索与读取本地参考提示词语料(需要本地数据包) |
build_prompt |
简易路线的提示词预览,不生成 |
list_catalog |
场景预设、24 个模板、风格键、语料状态 |
git clone https://github.com/joeseesun/qiaomu-codex-imagegen && cd qiaomu-codex-imagegen
node scripts/cli.mjs "窗边翻开的书,书页上升起咖啡的热气" --preset xiaohongshu
装好 Skill 和 MCP 后,对 Agent 说:
# 先要方向(免费),再选
node scripts/cli.mjs suggest "一个月的咖啡店阅读活动" --for 视频封面
# 展开模板,看完整提示词和验收条件(免费)
node scripts/cli.mjs compose --template T03-1 --var topic=谷雨 --var subject=嫩芽 --var headline=谷雨
# 带模板生成
node scripts/cli.mjs generate --template T03-1 --var topic=谷雨 --var subject=嫩芽 --var headline=谷雨
# 搜本地案例
node scripts/cli.mjs search "节气 茶" --limit 5
简易路线(不走方向推荐):
# 视频封面,三个方案先看
node scripts/cli.mjs "一个巨大的红色播放键像日出一样升起" --preset video-cover --style saul-bass --count 3
# Mondo 风格电影海报,只出图不放字
node scripts/cli.mjs "沙漠中巨大的沙虫轮廓,渺小的人影" --preset movie-poster --style olly-moss
# 必须带标题时,逐字给出
node scripts/cli.mjs "一座灯塔" --preset wechat-cover --text "夜航船"
# 图生图
node scripts/cli.mjs "保持同一支铅笔,背景换成暖米色纸张" --ref /abs/pencil.png --ratio 1:1
# 先看提示词,不生成
node scripts/cli.mjs "咖啡和书" --preset xiaohongshu --show-prompt
参数:--preset 场景,--style 风格(键或自由文字),--ratio 覆盖比例,--text 图内文字(可重复),--ref 参考图(可重复,绝对路径),--count 并行变体 1–4,--out 输出目录,--name 文件名,--model、--timeout,--json 机器可读输出。默认保存到 `~/Pictures/qiaomu-codex-ima