by shinpr
MCP server for AI image generation and editing with automatic prompt optimization and quality presets. Supports Nano Banana (Gemini), OpenAI GPT Image, and BytePlus Seedream.
# Add to your Claude Code skills
git clone https://github.com/shinpr/mcp-imageLast scanned: 5/30/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-05-30T16:32:17.089Z",
"npmAuditRan": true,
"pipAuditRan": true
}mcp-image is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by shinpr. MCP server for AI image generation and editing with automatic prompt optimization and quality presets. Supports Nano Banana (Gemini), OpenAI GPT Image, and BytePlus Seedream. It has 157 GitHub stars.
Yes. mcp-image passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/shinpr/mcp-image" and add it to your Claude Code skills directory (see the Installation section above).
mcp-image is primarily written in TypeScript. It is open-source under shinpr on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh mcp-image against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Generate and edit images from Codex, Cursor, Claude Code, or any MCP client. mcp-image adds visual direction to your request before sending it to Gemini, OpenAI, or BytePlus Seedream.
Tell it what image to create or what to change in an existing image, and what it is for. The result is saved to disk and returned to your assistant.
Before generating an image, mcp-image rewrites short requests into more specific prompts. It keeps what you asked for and fills in details such as composition, lighting, and camera angle. The more detail you provide, the less it changes.
You ask:
"A photo of a roast chicken dinner for a recipe site. It should look like it was actually cooked, and it should be partway through being carved so you can tell how juicy it is."
mcp-image sends to the image model:
"... a beautifully roasted whole chicken, golden-brown and glistening, resting on a rustic wooden cutting board. One leg is partially carved, revealing tender, succulent white meat and rich, glistening juices pooling around the carving knife ... shallow depth of field focused on the carved chicken."

Generated with Gemini using the default fast quality preset.
What carried through:
for a recipe site: one clear subject, with everything else kept subordinateactually cooked: uneven browning and juices across the boardpartway through being carved: the cut face and slices beside ithow juicy it is: close framing and shallow depth of field around the cutBaseline from the same request, with prompt enhancement disabled.
Set SKIP_PROMPT_ENHANCEMENT=true to send the original prompt to the image model unchanged.
You need Node.js 22 or later, an MCP-compatible client, and an API key for one image provider.
All three providers generate and edit images. Gemini is the default and requires the least configuration.
| Provider | Image size | Output format | Setup |
|---|---|---|---|
| Gemini (default) | 1K, 2K, 4K | Automatic | Get a key, then set GEMINI_API_KEY |
| OpenAI | 1K, 2K, 4K | PNG or JPEG | Get a key, then set IMAGE_PROVIDER=openai and OPENAI_API_KEY |
| BytePlus Seedream | 1K, 2K | PNG or JPEG | Get an AP region key, then set IMAGE_PROVIDER=seedream and ARK_API_KEY |
Google Search grounding is available with Gemini only. OpenAI may require organization verification before it can generate images.
The examples below use Gemini. Replace the provider settings if you prefer OpenAI or Seedream.
Add this to ~/.codex/config.toml:
[mcp_servers.mcp-image]
command = "npx"
args = ["-y", "mcp-image"]
[mcp_servers.mcp-image.env]
GEMINI_API_KEY = "your_gemini_api_key_here"
IMAGE_OUTPUT_DIR = "/absolute/path/to/images"
Add this to ~/.cursor/mcp.json for all projects, or .cursor/mcp.json in a project:
{
"mcpServers": {
"mcp-image": {
"command": "npx",
"args": ["-y", "mcp-image"],
"env": {
"GEMINI_API_KEY": "your_gemini_api_key_here",
"IMAGE_OUTPUT_DIR": "/absolute/path/to/images"
}
}
}
}
Run this in your project directory:
claude mcp add mcp-image --env GEMINI_API_KEY=your-api-key --env IMAGE_OUTPUT_DIR=/absolute/path/to/images -- npx -y mcp-image
Add --scope user after mcp-image to make it available in every project.
Never commit API keys to version control. Use an absolute IMAGE_OUTPUT_DIR in MCP configuration because the server's working directory depends on the client. If omitted, images are written to ./output relative to that working directory.
Restart your MCP client after changing its configuration, then ask your AI assistant:
Generate a product photo of a ceramic coffee mug on a wooden desk.
The generated file is saved in the configured output directory and returned to the assistant as an MCP resource.
pnpm install
pnpm run build
Configure the MCP client to run the local build instead of npx -y mcp-image:
node /absolute/path/to/mcp-image/dist/index.js
Give the assistant an absolute path to the source image:
Edit /path/to/image.jpg so the person is facing right.
Generate a high-quality product photo of a smartphone with clear text on the screen.Generate a cinematic desert landscape in a 21:9 aspect ratio.Keep the knight's appearance consistent with the previous image.See the tool reference for the options your assistant can pass explicitly.
Changing the provider changes both prompt enhancement and image generation. The way you ask for an image stays the same.
IMAGE_QUALITY accepts fast (default), balanced, or quality. Set it in the MCP server environment:
IMAGE_QUALITY=balanced
A request-level quality option takes precedence. Each provider maps the three values to its own image settings.
| Variable | Default | Description |
|---|---|---|
IMAGE_PROVIDER |
gemini |
Default provider: gemini, openai, or seedream |
GEMINI_API_KEY |
- | API key for Gemini |
OPENAI_API_KEY |
- | API key for OpenAI |
ARK_API_KEY |
- | ModelArk AP API key for Seedream |
IMAGE_OUTPUT_DIR |
./output |
Directory where generated images are saved; use an absolute path in MCP configuration |
IMAGE_QUALITY |
fast |
Default quality preset: fast, balanced, or quality |
SKIP_PROMPT_ENHANCEMENT |
false |
Set to true to send prompts through unchanged |
You can configure keys for more than one provider and switch per request. A request-level provider option takes precedence over IMAGE_PROVIDER.
Your MCP client calls this tool for you. Open the reference when you need to check an option or provider limitation.
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt |
string | Yes | Image description or editing instruction |
quality |
string | No | fast, balanced, or quality; overrides IMAGE_QUALITY |
provider |
string | No | gemini, openai, or seedream; overrides IMAGE_PROVIDER |
inputImagePath |
string | No | Absolute path to an input image for editing |
fileName |
string | No | Output filename; .png, .jpg, or .jpeg selects the format for OpenAI and Seedream |
aspectRatio |
string | No | 1:1 (default), 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 1:4, 1:8, 4:1, or 8:1 |
imageSize |
string | No | 1K, 2K, or 4K; availability depends on the provider |
blendImages |
boolean | No | Add blending guidance when combining visual elements |
maintainCharacterConsistency |
boolean | No | Keep a character's appearance consistent across images |
useWorldKnowledge |
boolean | No | Add context for historical figures, landmarks, and factual scenes |
useGoogleSearch |
boolean | No | Gemini only. Use Google Search grounding for current information |
purpose |
string | No | Intended use, such as cookbook cover or social media post |
Check that the key for the selected provider is present in the MCP server's environment:
GEMINI_API_KEYOPENAI_API_KEYARK_API_KEYRestart the MCP client after changing its configuration.
Use an absolute path and make sure the MCP server can read the file. Input images can be PNG, JPEG, or WebP and must be no larger than 10 MB. Seedream editing accepts PNG and JPEG only.
Check the requested size in the provider table. useGoogleSearch works with Gemini only, and Seedream does not support 4K. For OpenAI permission errors, check your organization settings. For quota or rate-limit errors, check the selected provider account.
This repository also includes an Agent Skill for assistants that already have access to an image generation tool. It teaches the prompt-writing approach used by mcp-image and works independently of this server.
Install it with:
npx mcp-image skills install --path <skills-directory>
For example, use ~/.codex/skills, ~/.cursor/skills, or ~/.claude/skills as the destination.
MIT License. See LICENSE for details.
Need help? Open an issue or check Troubleshooting.