by Jakedismo
100% Rust implementation of code graphRAG with blazing fast AST+FastML parsing, surrealDB backend and advanced agentic code analysis tools through MCP for efficient code agent context management
# Add to your Claude Code skills
git clone https://github.com/Jakedismo/codegraph-rustGuides for using ai agents skills like codegraph-rust.
Last scanned: 5/14/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-05-14T06:46:22.883Z",
"semgrepRan": false,
"npmAuditRan": true,
"pipAuditRan": true
}See how codegraph-rust compares with popular alternatives.
codegraph-rust is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by Jakedismo. 100% Rust implementation of code graphRAG with blazing fast AST+FastML parsing, surrealDB backend and advanced agentic code analysis tools through MCP for efficient code agent context management. It has 894 GitHub stars.
Yes. codegraph-rust passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/Jakedismo/codegraph-rust" and add it to your Claude Code skills directory (see the Installation section above).
codegraph-rust is primarily written in Rust. It is open-source under Jakedismo on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh codegraph-rust against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.

Your codebase, understood.
[!TIP] Give your coding agents lasting project memory.
Meet CodeGraph Agent Memory, our sibling project that builds on CodeGraph with persistent memory for architectural decisions, findings, failed approaches, and reusable procedures. Enable memory retrieval to bring that knowledge into agent workflows alongside current code evidence across sessions.
CodeGraph transforms your entire codebase into a semantically searchable knowledge graph that AI agents can actually reason about—not just grep through.
Ready to get started? Jump to the Installation Guide for step-by-step setup instructions.
Already set up? See the Usage Guide for tips on getting the most out of CodeGraph with your AI assistant.
Prefer shell commands? Run
codegraph agent context "your question". The same four agentic tools are available through the CLI, with project-local Claude Code/Codex hooks.codegraph initoffers hook setup and adds agent instructions before indexing.
AI coding assistants are powerful, but they're flying blind. They see files one at a time, grep for patterns, and burn tokens trying to understand your architecture. Every conversation starts from zero.
What if your AI assistant already knew your codebase?
Most semantic search tools create embeddings and call it a day. CodeGraph builds a real knowledge graph:
Your Code → Build Context → AST + FastML → LSP Resolution → Enrichment → Graph + Embeddings
↓ ↓ ↓ ↓ ↓ ↓
Packages Nodes/edges Type-aware API surface Graph Semantic
Features Fast patterns linking Module graph traversal search
Targets Spans Definitions Dataflow/Docs (hybrid)
When you search, you don't just get "similar code"—you get code with its relationships intact. The function that matches your query, plus what calls it, what it depends on, and where it fits in the architecture.
Indexing enrichment adds:
defines, uses, flows_to, returns, mutates) for impact analysisREADME.md, docs/**/*.md, and schema/**/*.surqlIndexing is tiered so you can choose between speed/storage and graph richness. The default is fast.
| Tier | What it enables | Typical use |
|---|---|---|
fast |
AST nodes + core edges only (no LSP or enrichment) | Quick indexing, low storage |
balanced |
LSP symbols + docs/enrichment + module linking | Richer navigation and documentation with moderate analyzer cost |
full |
All analyzers + LSP definitions + dataflow + architecture | Maximum graph richness and analyzer coverage |
Agent answer accuracy is evaluated separately; see the CLI accuracy results.
Tier behavior details:
fast: disables build context, LSP, enrichment, module linking, dataflow, docs/contracts, and architecture; filters out Uses/References edges.balanced: enables build context, LSP symbols, enrichment, module linking, and docs/contracts; filters out References edges.full: enables all analyzers and LSP definitions; no edge filtering.Configure the tier:
codegraph index /path/to/project --index-tier balancedCODEGRAPH_INDEX_TIER=balanced[indexing] tier = "balanced"Directory indexing scans subdirectories by default, so
codegraph index --languages Rust --index-tier balanced . finds Rust sources in workspace crates. Use
--no-recursive for an intentional root-only scan; -r/--recursive remain accepted.
codegraph estimate uses the same traversal defaults.
When the tier enables LSP (balanced/full), indexing fails fast if required external tools are missing.
Rust indexing also runs rust-analyzer --version from the target project before parsing:
an existing rustup shim does not guarantee that its active toolchain has the component.
With rustup, run rustup component add rust-analyzer from that project directory,
then verify rust-analyzer --version. Language-server failures retain the final 4 KiB
of stderr so startup and runtime errors include the server's diagnostic.
Required tools by language:
rust-analyzernode and typescript-language-servernode and pyright-langservergoplsjdtlsclangdWarm language-server sessions retain versioned documents and deduplicate/pipeline
definition requests. Symbol and definition requests retry transient ContentModified
(-32801) responses up to five times with backoff within one 30-second deadline.
Changed document versions, exhausted retries and other errors still fail indexing,
with the request method and file URI in the diagnostic. CODEGRAPH_LSP_REQUESTS bounds
outstanding requests per server (default 32). CODEGRAPH_ANALYZERS=0 disables analyzers
independently of tier. CODEGRAPH_SCIP_INDEX=/path/index.scip can substitute a compiler
index for LSP; source validation and sidecar requirements are described in the
indexing implementation guide.
--batch-size sets the maximum number of embedding texts per batch for both local
and cloud providers. An explicit value wins over CODEGRAPH_EMBEDDINGS_BATCH_SIZE,
its legacy alias CODEGRAPH_EMBEDDING_BATCH_SIZE, and [embedding] batch_size in
TOML, in that order; the default is 64. Memory-based tuning does not change explicit
values, including --batch-size 100.
codegraph index --languages Rust --index-tier balanced --batch-size 512 .
Token/byte budgets, cache hits and provider API limits can produce smaller actual
requests. For larger requests, adjust CODEGRAPH_EMBEDDING_BATCH_TOKENS (default at
least the resolved input context, with local/remote floors of 8192/32768) and
CODEGRAPH_EMBEDDING_BATCH_BYTES (default 1 MiB) as needed.
Logs show these inference limits separately from database write batches;
CODEGRAPH_CHUNK_DB_BATCH_SIZE defaults to at most 32 rows and is capped at 512.
Ollama and LM Studio have no extra fixed 256-text cap.
Full, single-file and watch indexing share complete-project reconciliation. Unchanged
sources reuse cached AST/analyzer artifacts while the full catalog retains callers
across edits, renames and deletions. Only changed graph records and file metadata are
written. Source snapshots, parsing, inference and writer queues have independent
resource bounds; completion follows durable acknowledgements and final input checks.
--force prepares and reconciles again without trusting the previous catalog.
Embedding and semantic-resolution work is independent of extraction tier. Each accepts
sync (default when its feature is compiled), deferred or off:
CODEGRAPH_EMBEDDING_POLICY=deferred CODEGRAPH_SEMANTIC_RESOLUTION=off \
codegraph index /path/to/project --index-tier fast --stats-json indexing.json
codegraph index /path/to/project --complete-deferred --stats-json completed.json
Deferred runs persist a resumable job and distinguish graph readiness from pending
inference. Prepared-text embedding caches include model/task/tokenizer/runtime identity;
chunking preserves Unicode and enforces provider token budgets. Mutable model aliases
expire; CODEGRAPH_MODEL_REVISION declares an immutable revision.
Chunk planning starts from AST node source spans, such as functions and classes. A unit that fits the complete input budget stays intact. Oversized units split at Tree-sitter statement/block boundaries, then merge adjacent pieces while they fit. Oversized leaves and unsupported syntax use UTF-8-safe line/token splitting. Unicode and structural whitespace are preserved; overlap uses token counts.
For Ollama, the budget follows model metadata and serving context. Qwen3 embedding
models support 32K inputs;
nomic-embed-text-v2-moe
supports 512 tokens. Counting includes document/query prefixes and special tokens.
Nomic text models receive search_document: / search_query: ; Qwen3 queries
receive a retrieval instruction. Every Ollama request uses truncate=false:
context mismatch fails visibly instead of silently dropping source.
Recognized models automatically load their matching publisher tokenizer, caching
tokenizer.json and downloading it on first use if absent. Model weights are not
downloaded. Custom/offline models can supply CODEGRAPH_TOKENIZER_PATH; an unknown
model requires that path or CODEGRAPH_TOKENIZER_REPO.
| Control | Behavior |
|---|---|
CODEGRAPH_CHUNK_MAX_TOKENS |
Lower the complete-input target; capped by serving context. Legacy CODEGRAPH_MAX_CHUNK_TOKENS has lower precedence. |
CODEGRAPH_CHUNK_SMART_SPLIT=0 |
Use token splitting instead of AST boundaries for oversized units. Default: AST splitting. |
CODEGRAPH_CHUNK_OVERLAP_TOKENS |
Maximum suffix overlap in tokens; default 64, zero disables. Shrunk to fit t |