by waybarrios
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
# Add to your Claude Code skills
git clone https://github.com/waybarrios/vllm-mlxGuides for using mcp servers skills like vllm-mlx.
Last scanned: 5/1/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-05-01T06:39:51.820Z",
"semgrepRan": false,
"npmAuditRan": true,
"pipAuditRan": true
}See how vllm-mlx compares with popular alternatives.
vllm-mlx is an open-source mcp servers skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by waybarrios. High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support. It has 1,617 GitHub stars.
Yes. vllm-mlx passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/waybarrios/vllm-mlx" and add it to your Claude Code skills directory (see the Installation section above).
vllm-mlx is primarily written in Python. It is open-source under waybarrios on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other MCP Servers skills you can browse and compare side by side. Open the MCP Servers category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh vllm-mlx against similar tools.
No comments yet. Be the first to share your thoughts!
Top skills in this category by stars
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference.
Read this in other languages: English · Español · Français · 中文
A vLLM-style inference server for Apple Silicon Macs. Unlike Ollama or mlx-lm used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI /v1/* and Anthropic /v1/messages from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step.
New users: follow the isolated release install and first-response check. The example below assumes that environment is activated. Initial model download/loading time depends on the machine and connection.
The version below is this walkthrough's pinned reference release, not an automatically updated latest version.
python -m pip install 'vllm-mlx==0.4.1'
vllm-mlx serve mlx-community/Llama-3.2-3B-Instruct-4bit --host 127.0.0.1 --port 8000
After the first response, restart with --continuous-batching to try
continuous batching and its cache options.
The minimal command above uses the default engine without that flag.
OpenAI SDK:
Install openai in your Python client environment first; it is not installed
by the server package. See the client setup.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
r = client.chat.completions.create(model="mlx-community/Llama-3.2-3B-Instruct-4bit", messages=[{"role": "user", "content": "Hi!"}])
print(r.choices[0].message.content)
Anthropic SDK / Claude Code:
export ANTHROPIC_BASE_URL=http://localhost:8000
export ANTHROPIC_API_KEY=not-needed
claude
Validated with OpenCode, pi, Codex, Claude Code, GitHub Copilot CLI, Cline CLI, and OpenClaw's embedded agent. All seven completed a streamed tool interaction and an exact file edit in one local run with Qwen3.8-27B-4bit on September 19, 2026. Results apply to the tested client versions and settings.
See the validated CLI matrix and setup guide for versions, API transports, reproduction commands, and coverage limits.
/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/rerank, /v1/responses/v1/messages (streaming, tool use, system prompts)response_format (lm-format-enforcer)--ssd-cache-dir)--warm-prompts) for 1.3-2.25x TTFTaudio_url content blocks)--reasoning-parser)--moe-top-k for +7-16% on Qwen3-30B-A3B--enable-mtp for supported models--specprefill for supported model/draft combinations/metrics endpoint with --enable-metricsvllm-mlx bench-serve for prompt sweeps with CSV/JSON outputLLM decode (M4 Max, 128 GB, greedy, single stream):
| Model | Tok/s | Memory |
|---|---|---|
| Qwen3-0.6B-8bit | 417.9 | 0.7 GB |
| Llama-3.2-3B-Instruct-4bit | 205.6 | 1.8 GB |
| Qwen3-30B-A3B-4bit | 127.7 | ~18 GB |
Audio speech-to-text (M4 Max, RTF = real-time factor):
| Model | RTF | Use case |
|---|---|---|
| whisper-tiny | 197x | Real-time / low latency |
| whisper-large-v3-turbo | 55x | Quality + speed |
| whisper-large-v3 | 24x | Highest accuracy |
See docs/benchmarks/ for continuous-batching results, KV-cache quantization (4-bit / 8-bit / fp16), and MoE top-k sweeps.
vllm-mlx serve mlx-community/Qwen3-8B-4bit --port 8000
export ANTHROPIC_BASE_URL=http://localhost:8000
export ANTHROPIC_API_KEY=not-needed
claude
vllm-mlx serve mlx-community/Qwen3-8B-4bit --reasoning-parser qwen3
r = client.chat.completions.create(
model="mlx-community/Qwen3-8B-4bit",
messages=[{"role": "user", "content": "What is 17 * 23?"}],
)
print("Thinking:", r.choices[0].message.reasoning)
print("Answer:", r.choices[0].message.content)
vllm-mlx serve mlx-community/Qwen3-VL-4B-Instruct-3bit --port 8000
r = client.chat.completions.create(
model="mlx-community/Qwen3-VL-4B-Instruct-3bit",
messages=[{"role": "user", "content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}},
]}],
)
With the text server from the quick start:
r = client.chat.completions.create(
model="mlx-community/Llama-3.2-3B-Instruct-4bit",
messages=[{"role": "user", "content": "List 3 colors."}],
response_format={
"type": "json_schema",
"json_schema": {
"schema": {"type": "object", "properties": {"colors": {"type": "array", "items": {"type": "string"}}}}
},
},
)
/v1/rerank)curl http://localhost:8000/v1/rerank -H 'Content-Type: application/json' -d '{
"model": "default",
"query": "apple silicon inference",
"documents": ["MLX is Apples framework", "Metal kernels on M-series", "CUDA on NVIDIA"]
}'
The built-in MLX reranker forward path supports standard BERT/XLM-RoBERTa
sequence-classification weights with gelu, gelu_new/gelu_fast, relu, or
silu/swish hidden_act values. Other activations fail explicitly so custom
reranker architectures can add a dedicated adapter instead of silently using the
wrong activation.
vllm-mlx serve <llm-model> --embedding-model mlx-community/all-MiniLM-L6-v2-4bit
emb = client.embeddings.create(model="mlx-community/all-MiniLM-L6-v2-4bit", input=["Hello", "World"])
pip install vllm-mlx[audio]
brew install espeak-ng # macOS, needed for non-English TTS
python examples/tts_example.py "Hello, how are you?" --play
python examples/tts_multilingual.py "Hola mundo" --lang es --play
vllm-mlx bench-serve --url http://localhost:8000 --concurrency 5 --prompts prompts.txt --output results.csv
# Product-style workload with quality checks and metrics deltas
vllm-mlx bench-serve --url http://localhost:8000 --workload workload.json --repetitions 5 --output results.json
# Append workload rows into SQLite for longitudinal comparisons
vllm-mlx bench-serve --url http://localhost:8000 --workload workload.json --repetitions 5 --format sqlite --output bench.db
# Inspect repo metadata, file sizes, config, and rough fit before downloading weights
vllm-mlx model inspect mlx-community/Llama-3.2-3B-Instruct-4bit
# Acquire with resumable Hugging Face transfer and write a local artifact manifest
vllm-mlx model acquire mlx-community/Llama-3.2-3B-Instruct-4bit --target-dir ./models/llama-3b-4bit
# Wrap mlx-lm conversion and record the exact recipe in the converted artifact
vllm-mlx model convert meta-llama/Llama-3.2-3B-Instruct --output ./models/llama-3b-mlx-q4 --quantize --q-bits 4 --q-group-size 64 --q-mode affine
vllm-mlx serve <model> --enable-metrics
curl http://localhost:8000/metrics
Use the release walkthrough for an isolated environment and a pinned server version. If you already manage CLI tools with uv, its isolated equivalent is:
uv tool install 'vllm-mlx==0.4.1'
Keep development checkouts separate. See the Installation Guide for optional extras and Audio Guide for audio setup.
Browse the