GitNexus
The context engine for Enterprise Codebases
Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow — then exposes it through smart MCP tools so AI agents never miss code.
Gitnexus_CLI.1.mp4
Like DeepWiki, but deeper. DeepWiki helps you understand code. GitNexus lets you analyze it — a knowledge graph tracks every relationship, not just descriptions.
TL;DR: The CLI + MCP makes your AI agent reliable — it gives Cursor, Claude Code, Antigravity, Codex, and friends a deep architectural view of your codebase so they stop missing dependencies, breaking call chains, and shipping blind edits. Even smaller models get full architectural clarity. The Web UI is a quick way to chat with any repo in the browser.
Quick Start
# 1. Index your repo (run from repo root)
npx gitnexus analyze
# 2. Connect your editors (one-time, auto-detects Claude Code, Cursor, Codex, …)
npx gitnexus setupThat's it. analyze indexes the codebase, installs agent skills, registers Claude Code hooks, and creates AGENTS.md / CLAUDE.md context files — all in one command. setup writes the MCP config so your AI agent can use the graph.
Install problems? npm 11 crash · slow cold install · no C++ toolchain
On npm 11.x?
npxcan crash during install withCannot destructure property 'package' of 'node.target'(an npm/arborist bug, before GitNexus runs). Use pnpm instead — it builds the native deps explicitly:pnpm --allow-build=@ladybugdb/core --allow-build=gitnexus --allow-build=tree-sitter dlx gitnexus@latest analyzeOr install globally (
npm install -g gitnexus@latest) and rungitnexus analyze. See #1939.
Fastest MCP startup: install globally (
npm i -g gitnexus) before runninggitnexus setup— this writes an absolute-path MCP config that bypassesnpxentirely. On a cold cache, annpx-based MCP install can exceed Claude Code'sMCP_TIMEOUTdefault (~30s).
No C++ toolchain? Set
GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1beforenpm install -g gitnexusto skip the vendored grammar materialize/build fortree-sitter-dart,tree-sitter-proto,tree-sitter-swift, andtree-sitter-kotlin— those four languages won't be parsed, but install completes in seconds withoutpython3/make/g++. Strict=1only — any other value falls through to the rebuild.
Local embeddings are opt-in. Default
npm installdoes not fetch@huggingface/transformersoronnxruntime-node. Rungitnexus embeddings install(orgitnexus analyze --embeddings, which auto-heals) to fetch the stack through your npm registry config into~/.gitnexus/embedding-runtime. CUDA GPU binaries still use NuGet via--cuda(#2370). The prefix needs Node withmodule.registerHooks(≥ 22.15 on 22.x, ≥ 23.5 on 23.x). A leftover 1.6.12 package-first tree is residual until a clean reinstall;--forceonly refreshes prefix overrides.
About
tree-sitter-kotlin: like Dart/Proto/Swift, Kotlin is a vendored grammar (undergitnexus/vendor/tree-sitter-kotlin). Upstream ships source only (no prebuilt binaries), so GitNexus cross-builds the platform prebuilds itself (via thebuild-tree-sitter-prebuildsGitHub Actions workflow) and vendors them — the same uniform pipeline used for Dart, Proto, and Swift.node-gyp-buildselects the right.nodeat require time, so no C/C++ toolchain is needed. If no prebuild matches your platform-arch, only Kotlin (.kt/.kts) parsing is unavailable; the rest ofgitnexusis unaffected.
Deploy to Render
Deploy GitNexus in one click:
The Blueprint creates two services. gitnexus-server runs gitnexus serve as a private service: no public URL, reachable only over Render's private network, with a persistent disk for indexes and cloned repos. gitnexus-web is the public one. It serves the UI and reverse-proxies /api/* to the server, so the browser talks to a single origin.
At the Blueprint's defaults this runs about $35/month: $25 for the server's standard instance, $7 for the web service's starter instance, and $2.50 for the 10 GB disk. See Render's pricing for other plans.
The deploy generates an access token, and the UI asks for it on first use:
- Open the
gitnexus-webservice in your Render dashboard. - Copy
GITNEXUS_SERVE_AUTH_TOKENfrom its Environment tab. - Load the site and paste the token into the prompt (or the settings panel).
Every /api/* request carries that token as a header, and the proxy answers 401 without it. The browser keeps it in sessionStorage, so a new tab asks again. To rotate it, edit the environment variable and redeploy.
The proxy strips Origin before forwarding, so the server's CSRF guard does nothing for proxied traffic; it passes Origin-less requests through by design. The token is the only control on this deploy, not a second layer behind the guard. Anyone holding it can read every indexed repo. See SECURITY.md.
Indexing is memory-bound. If gitnexus-server runs out of memory on a large repo, raise its plan, which sets available RAM: standard is 2 GB, pro is 4 GB. Raise sizeGB only if the disk fills with clones and indexes.
Deploy to RepoCloud
Two Ways to Use GitNexus
| CLI + MCP (recommended) | Web UI | |
|---|---|---|
| What | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
| For | Daily development with Cursor, Claude Code, Antigravity, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
| Scale | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
| Install | npm install -g gitnexus |
No install — gitnexus.vercel.app |
| Storage | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
| Parsing | Tree-sitter native bindings | Tree-sitter WASM |
| Privacy | Everything local, no network | Everything in-browser, no server |
Bridge mode:
gitnexus serveconnects the two — the web UI auto-detects the local server and can browse all your CLI-indexed repos without re-uploading or re-indexing.
Why a Knowledge Graph?
Tools like Cursor, Claude Code, Codex, Cline, Roo Code, and Windsurf are powerful — but they don't truly know your codebase structure. So this happens:
- AI edits
UserService.validate() - Doesn't know 47 functions depend on its return type
- Breaking changes ship
Traditional Graph RAG gives the LLM raw graph edges and hopes it explores enough. GitNexus precomputes structure at index time — clustering, tracing, scoring — so tools return complete context in one call:
flowchart TB
subgraph Traditional["Traditional Graph RAG"]
direction TB
U1["User: What depends on UserService?"]
U1 --> LLM1["LLM receives raw graph"]
LLM1 --> Q1["Query 1: Find callers"]
Q1 --> Q2["Query 2: What files?"]
Q2 --> Q3["Query 3: Filter tests?"]
Q3 --> Q4["Query 4: High-risk?"]
Q4 --> OUT1["Answer after 4+ queries"]
end
subgraph GN["GitNexus Smart Tools"]
direction TB
U2["User: What depends on UserService?"]
U2 --> TOOL["impact UserService upstream"]
TOOL --> PRECOMP["Pre-structured response:
8 callers, 3 clusters, all 90%+ confidence"]
PRECOMP --> OUT2["Complete answer, 1 query"]
end
Core innovation: Precomputed Relational Intelligence
- Reliability — the LLM can't miss context; it's already in the tool response
- Token efficiency — no 10-query chains to understand one function
- Model democratization — smaller LLMs work because the tools do the heavy lifting
What Your AI Agent Gets
19 MCP tools (17 per-repo + 2 group)
| Tool | What It Does |
|---|---|
list_repos |
Discover all indexed repositories (paginated — limit/offset) |
query |
Process-grouped hybrid search (BM25 + semantic + RRF) |
context |
360-degree symbol view — categorized refs, process participation |
impact |
Blast radius analysis with depth grouping and confidence |
trace |
Shortest directed path between two symbols (call + class-member edges) |
detect_changes |
Git-diff impact — maps changed lines to affected processes |
check |
Read-only structural checks against the indexed graph |
rename |
Multi-file coordinated rename with graph + text search |
cypher |
Raw Cypher graph queries |
route_map |
API route map — which components fetch which endpoints, and handlers |
tool_map |
MCP/RPC tool definitions — where they're defined and handled |
shape_check |
Validate API response shapes against consumers' property accesses |
api_impact |
Pre-change impact report for an API route handler |
explain |
Explain persisted taint findings (source→sink flows, --pdg indexes) |
pdg_query |
Query control/data dependence at statement level (--pdg indexes) |
read_file |
Read a checkout file (optional 0-indexed slice; maxLines cap) |
grep |
Regex search of the working tree for indexed files (1-based hits) |
group_list |
List configured repository groups |
group_sync |
Rebuild a group's Contract Registry and cross-repo links |
Per-repo read-only tools take an optional
repoparameter. Omit it when only one repo is indexed, an MCP default is configured, or the GitNexus process cwd is inside a registered path without crossing into an unindexed nested Git checkout; otherwise pass it explicitly. Mutating tools requirerepowhen multiple repos are indexed and no MCP default exists. Per-repo tools also take an optionalbranchfor indexes pinned withgitnexus analyze --branch, exceptread_fileandgrep, which read the checkout and do not acceptbranch. Omittingbranchqueries the workspace index, which follows your checked-out working tree — switching branches and re-runninggitnexus analyzeupdates it incrementally.explainandpdg_queryneed an index built withgitnexus analyze --pdg.
Resources for instant context
| Resource | Purpose |
|---|---|
gitnexus://repos |
List all indexed repositories (read this first) |
gitnexus://setup |
Setup and usage guidance for agents |
gitnexus://repo/{name}/context |
Codebase stats, staleness check, and available tools |
gitnexus://repo/{name}/clusters |
All functional clusters with cohesion scores |
gitnexus://repo/{name}/cluster/{name} |
Cluster members and details |
gitnexus://repo/{name}/processes |
All execution flows |
gitnexus://repo/{name}/process/{name} |
Full process trace with steps |
gitnexus://repo/{name}/schema |
Graph schema for Cypher queries |
gitnexus://group/{name}/contracts |
A group's extracted contracts and cross-links |
gitnexus://group/{name}/status |
Staleness of repos in a group |
2 MCP prompts for guided workflows
| Prompt | What It Does |
|---|---|
detect_impact |
Pre-commit change analysis — scope, affected processes, risk level |
generate_map |
Architecture documentation from the knowledge graph with mermaid diagrams |
Agent skills installed to .claude/skills/ and .agents/skills/ (if .agents/ exists) automatically
- Exploring — navigate unfamiliar code using the knowledge graph
- Debugging — trace bugs through call chains
- Impact Analysis — analyze blast radius before changes
- Refactoring — plan safe refactors using dependency mapping
- Guide — GitNexus tool/resource/schema reference for the agent
- CLI — run analyze/status/clean/wiki commands on request
- PDG Query — statement-level control/data dependence queries (
--pdgindex) - Taint Analysis — source→sink data-flow findings (
--pdgindex) - Plan (
/gitnexus-plan) — implementation-ready engineering plans backed by the graph and PDG slices - Work (
/gitnexus-work) — executes a plan as impact-checked,detect_changes-gated atomic commits - Review (
/gitnexus-review) — graph-backed review of a PR, branch, range, or local diff, with taint pass and per-domain expert lenses - LFG (
/gitnexus-lfg) — the full pipeline: plan → user gate → work → review
Repo-specific skills — run gitnexus analyze --skills and GitNexus detects the functional areas of your codebase (via Leiden community detection) and generates each one as a direct project skill under .claude/skills/gitnexus-area-<name>/. Each skill describes a module's key files, entry points, execution flows, and cross-area connections, and is regenerated on each --skills run to stay current.
When a repo contains an .agents/ directory, the standard and generated skills are also mirrored to .agents/skills/ (e.g. .agents/skills/gitnexus-cli/, .agents/skills/gitnexus-area-<name>/) so agents that read repo-local .agents/skills/ (like Codex) stay in sync.
Editor Setup
gitnexus setup auto-detects your editors and writes the correct global MCP config. Run it once. To configure only selected integrations, pass --coding-agent/-c with a comma-separated list, e.g. gitnexus setup -c cursor,codex.
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|---|---|---|---|---|
| Claude Code | Yes | Yes | Yes (PreToolUse + PostToolUse) | Full |
| Cursor | Yes | Yes | Yes (postToolUse, manual install) | Full |
| Antigravity (Google) | Yes | Yes | Yes (AfterTool, Gemini CLI hooks schema)¹ | Full |
| Codex | Yes | Yes | Yes (PreToolUse + PostToolUse, Codex hooks) | Full |
| Factory (Droid) | Yes | Yes | Yes (PostToolUse, plugin) | Full |
| OpenCode | Yes | Yes | — | MCP + Skills |
| CodeBuddy (Tencent) | Yes | Yes | — | MCP + Skills |
| Qoder (Alibaba) | Yes | Yes | — | MCP + Skills |
| Windsurf | Yes | — | — | MCP |
Full means MCP tools + agent skills + hooks that enrich searches with graph context. Claude Code and Codex go deepest: their PreToolUse hooks enrich the search before it runs, and their PostToolUse hooks also detect a stale index after commits and prompt the agent to reindex. Cursor, Antigravity, and Factory augment from a post-tool hook only, so they enrich the result rather than the query and do not carry the stale-index hint.
¹ Antigravity hooks follow the Gemini CLI hooks reference (Antigravity 2.0 is the documented successor to Gemini CLI). Augmentation runs in
AfterToolbecauseBeforeToolhas no context-injection channel in the Gemini contract — the agent sees graph context appended to the tool result viahookSpecificOutput.additionalContext. Stale-index hints land in the same channel after a successfulgit commit/merge/rebase/cherry-pick/pull. The schema may evolve if Antigravity-specific hook docs diverge from Gemini CLI's; the implementation will track those changes.
Manual MCP configuration (if you prefer not to run gitnexus setup)
Claude Code (full support — MCP + skills + hooks):
# macOS / Linux
claude mcp add gitnexus -- npx -y gitnexus@latest mcp
# Windows
claude mcp add gitnexus -- cmd /c npx -y gitnexus@latest mcpCodex (full support — MCP + skills + hooks):
codex mcp add gitnexus -- npx -y gitnexus@latest mcpOr via ~/.codex/config.toml (system scope) / .codex/config.toml (project scope):
[mcp_servers.gitnexus]
command = "npx"
args = ["-y", "gitnexus@latest", "mcp"]Codex hooks (PreToolUse graph enrichment + PostToolUse stale-index detection in ~/.codex/hooks.json, same schema as Claude Code) need the bundled adapter script, so they are installed by gitnexus setup -c codex rather than manually.
Alternatively, install everything as a Codex plugin (MCP + skills + hooks in one step):
codex plugin marketplace add abhigyanpatwari/GitNexus
# then inside Codex: /plugins → install "GitNexus"Codex notes: SessionStart is intentionally not registered — Codex reads AGENTS.md natively, which already carries the GitNexus context block. Newly installed hooks need a one-time approval in Codex via
/hooksbefore they run. Pick one install route (gitnexus setup -c codexor the plugin): plugin hooks load alongside~/.codex/hooks.json, so installing both can fire duplicate hooks per tool call.
Factory (Droid) — MCP + skills via gitnexus setup -c droid, or add the server manually to ~/.factory/mcp.json (user scope, applies to all projects):
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}gitnexus setup -c droid also installs skills to ~/.factory/skills/. For the PostToolUse search-augment hook, install the bundled gitnexus-factory-plugin/ — from a marketplace that includes this repo, run droid plugin install gitnexus@<marketplace>, or point Droid at it via extraKnownMarketplaces in .factory/settings.json. Factory reads AGENTS.md natively, which already carries the GitNexus context block.
Cursor (~/.cursor/mcp.json — global, works for all projects):
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}Antigravity (Google) — ~/.gemini/antigravity/mcp_config.json:
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}
gitnexus setupalso merges anAfterToolentry into~/.gemini/settings.json(under the canonical Gemini CLI hooks schema) and installs skills to~/.gemini/antigravity/skills/. Existing user hooks are preserved. The hook adapter's path is rewritten at install time, so rungitnexus setuprather than hand-editing.
OpenCode (~/.config/opencode/config.json):
{
"mcp": {
"gitnexus": {
"type": "local",
"command": ["gitnexus", "mcp"]
}
}
}CodeBuddy (Tencent) — priority chain, edit the first non-empty file that exists: ~/.codebuddy/.mcp.json (recommended) → ~/.codebuddy/mcp.json (deprecated) → ~/.codebuddy.json (legacy). CodeBuddy reads only the first existing file, so adding servers to a higher-priority file than the one currently in use would hide the servers below it. Create ~/.codebuddy/.mcp.json only if none exist:
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}Qoder (Alibaba) — ~/.qoder.json:
{
"mcpServers": {
"gitnexus": {
"command": "npx",
"args": ["-y", "gitnexus@latest", "mcp"]
}
}
}MCP read-only mode
Set GITNEXUS_MCP_READ_ONLY=1 before starting the MCP server to expose only the proven single-repository read surface. Raw cypher, rename and group tools, group routing, and group resources are omitted from discovery and rejected before backend dispatch. Tool descriptions and generated setup/context resources are scrubbed so they do not recommend unavailable routes.
The default is unchanged when the variable is unset or 0. Any other value fails server startup rather than silently weakening the policy.
MCP repository policy
Set GITNEXUS_MCP_ALLOWED_REPOS to a comma-separated list of canonical registry names or absolute indexed paths. Entries are trimmed, resolved against the registry, and deduplicated at startup. When exactly one repository is allowed it becomes the implicit default; when several are allowed, callers must select one unless GITNEXUS_MCP_DEFAULT_REPO is also set.
The default repository must resolve to an allowed repository. Invalid, ambiguous, blank, or mismatched configuration fails startup before stdio or HTTP begins serving. The allowlist applies to tools, aliases, discovery, resources, templates, implicit resolution, and embedded HTTP; hidden repository details are not included in selection errors. Setting only GITNEXUS_MCP_DEFAULT_REPO chooses a default without restricting explicit repository selections. An allowed repository whose name is duplicated in the registry must be configured by path, and its context resource is only served for the unique name form.
MCP response budgets
The query, context, and impact tools accept an optional positive-integer maxTokens argument. It bounds the complete formatted MCP response, including hints and error text, using a deterministic four-UTF-8-bytes-per-token estimate. When truncation is required, the response ends with … and remains valid UTF-8.
Set GITNEXUS_MCP_DEFAULT_MAX_TOKENS to apply the same guardrail when callers do not send maxTokens. An explicit tool argument takes precedence. Leaving both unset preserves the existing response byte-for-byte; this is a transport guardrail, not semantic pagination or an exact model-specific tokenizer limit.
CLI Reference
Everyday commands:
gitnexus setup # Configure MCP for detected editors (one-time; -c to select)
gitnexus analyze [path] # Index a repository (or update a stale index)
gitnexus analyze [path] --watch # Watch local files and serialize incremental refreshes
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
gitnexus eval-server # Start lightweight evaluation HTTP tools (loopback by default)
gitnexus list # List all indexed repositories
gitnexus status # Show index status for current repo
gitnexus clean # Delete index for current repo
gitnexus wiki [path] # Generate repository wiki from knowledge graph
gitnexus uninstall # Preview removal of GitNexus MCP/skills/hooks (--force to apply)You can also query the graph directly from the terminal — gitnexus query, context, impact, trace, cypher, detect-changes, and check mirror the MCP tools of the same names, and gitnexus doctor prints runtime platform capabilities.
gitnexus analyze --watch requires a Git repository. It runs one initial
analysis, then debounces scanner-admitted working-tree changes for 300 ms by
default and applies serialized incremental refreshes. Events arriving during a
refresh remain queued, and retryable failures retain the same batch with bounded
backoff. Invalid .gitnexusrc or ignore-file reloads pause ordinary refreshes
until the control file is fixed. Stop the watcher with Ctrl+C.
Watch mode accepts --debounce, --workers, --worker-timeout,
--max-file-size, --max-processes, --max-process-branching,
--max-process-trace-depth, --max-entry-point-candidates, --branch, --pdg, --skip-fts, --name, --allow-duplicate-name, and
--verbose. Explicit one-shot options such as --force, --repair-fts,
embedding flags, --skills, --self-commit, --index-only, and --skip-git
are rejected. Unsupported defaults from .gitnexusrc are ignored with a
warning rather than making an otherwise valid repository unwatchable.
POSIX requests clone-first copy-and-swap publication when the live index has no
orphan sidecars. Windows and sidecar fallback runs update in place: failures
known to occur before writes are retried, while a failure that may have mutated
the live index stops the watcher. Watch mode does not pull remotes. Running MCP
and serve processes reopen a newly published index automatically; MCP observes
the replacement on its next tool call, typically within five seconds, so no
restart is required.
Authenticated eval-server binding
gitnexus eval-server binds to 127.0.0.1 by default. Loopback bindings do not require authentication. Any non-loopback bind, including 0.0.0.0, a LAN address, or a hostname that resolves to a LAN IPv4 address, requires GITNEXUS_AUTH_TOKEN. Every endpoint then requires an exact Authorization: Bearer <token> header.
GITNEXUS_AUTH_TOKEN='replace-me' gitnexus eval-server --host 0.0.0.0The token may be set in the shell, .env.local, or .env in the working directory. Precedence is shell > .env.local > .env. Only GITNEXUS_AUTH_TOKEN is read from those files; their other values are not added to the process environment. Keep token files uncommitted.
All analyze flags
gitnexus analyze --force # Full graph + FTS rebuild (reuses unchanged parser output)
gitnexus analyze --no-parse-cache # Full rebuild that re-parses every source file
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
gitnexus analyze --skip-fts # Index graph/embeddings without loading FTS or building keyword indexes
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
gitnexus analyze --embeddings [limit] # Enable embedding generation (slower, better search)
gitnexus analyze --skip-agents-md # Preserve custom AGENTS.md/CLAUDE.md gitnexus section edits
gitnexus analyze --skip-skills # Skip installing standard skill files under .claude/skills/ and .agents/skills/
gitnexus analyze --skip-git # Index folders that are not Git repositories
gitnexus analyze --default-branch develop # Branch used in the generated regression-compare example (base_ref)
gitnexus analyze --verbose # Log skipped files when parsers are unavailable
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
gitnexus analyze --workers <n> # Parse worker pool size (>=1; default: cores-1, capped at 16,
# auto-sized to the repo). 0 is rejected — there is no sequential mode.
gitnexus analyze --max-processes <n> # Process-detection process cap (replaces dynamic max(20, round(symbols/10)))
gitnexus analyze --max-entry-point-candidates <n> # Ranked entry-point pool (default 200; raise when the warning names it)
gitnexus analyze --spring-actuator ./actuator # Enrich with local Spring Boot Actuator JSON snapshots
gitnexus analyze --asyncapi-spec ./docs/asyncapi # Resolve broker addresses from AsyncAPI 3.x documents
gitnexus analyze --memory-budget 3000 # Main-thread V8 heap in MB (>= 200); overrides the auto-sizer and --max-old-space-size
gitnexus analyze --wal-checkpoint-threshold 67108864 # LadybugDB WAL auto-checkpoint threshold in bytes
# (default 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB)--skip-fts (or GITNEXUS_SKIP_FTS=1) disables FTS extension loading and keyword-index construction for this analysis. Graph queries, communities, processes, and existing embeddings remain available. Status and search report "FTS disabled for this index". Remove both the flag and environment setting and run analyze again to restore keyword search, even at the same commit. Only the exact environment value 1 enables the opt-out; the flag takes precedence. It cannot be combined with --repair-fts. Disabling an existing FTS index may require one graph-store rebuild to avoid unsafe writes through native indexes.
--spring-actuator is explicitly opt-in and accepts either a JSON bundle keyed by mappings, beans, conditions, configprops, and/or env, or a directory containing endpoint-named JSON files. It confirms matching static nodes and adds conservative runtime-only routes, beans, and property keys. The configured input is excluded from source scanning; only normalized repository-relative exclusions are retained for future scans, never absolute paths. Env/configprops values, origins, condition messages, and source names are never persisted or printed. Because snapshots are external runtime state, an enabled run always rebuilds; the first later run without the option rebuilds once to remove runtime evidence. The same path can be set as springActuator in .gitnexusrc.
--asyncapi-spec is explicitly opt-in and accepts a directory of AsyncAPI documents or a single document; the path is resolved against the repository root, so a committed docs/asyncapi and an absolute cache written by something else both work. Each operations[] entry of an AsyncAPI 3.x document can contribute a Destination node keyed by broker and address, with action: send emitting PUBLISHES_TO and action: receive emitting CONSUMES_FROM, so a document and source code that name one address on one broker land on the same node. Edges start at the document, not at a callable — a document states that the service talks to an address, not which method does — and no address a document names is ever attached to an unresolved source site.
An operation must name a protocol, either through its own bindings or through the servers[].protocol of the servers its channel resolves to (a channel that lists no servers resolves to all of them); operations that name none are refused, as are operations whose two readings name different brokers, and channels that inherit a multi-protocol server set without choosing. HTTP and WebSocket documents are refused for destination minting: there the host rather than the address names the place, and an HTTP endpoint is already modelled as a Route. A parameterized address — a channel declaring parameters, or an address containing { — is refused rather than keyed: two services publishing {env}.orders share a pattern, not a queue. AsyncAPI 2.x is refused under its own counted reason and never mapped, because its publish/subscribe are inverted relative to 3.x send/receive and a naive mapping would reverse the async graph while leaving it connected. Every refusal is counted, and a configured path that yields nothing is reported rather than passed over in silence.
Like Actuator snapshots, documents are external to git freshness — replacing one moves no commit and dirties no file — so an enabled run always rebuilds, and the first later run without the option rebuilds once to remove document-derived evidence. There is no glob-based auto-discovery, and the option is unsupported with --watch.
If analyze reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use --worker-timeout 60 or set GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000. For very large files, GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES controls the worker job byte budget.
Embeddings node limit — gitnexus analyze --embeddings generates semantic search vectors with a default 50,000-node safety cap to protect memory on large repositories:
gitnexus analyze --embeddings # default 50,000 node safety cap
gitnexus analyze --embeddings 0 # disable the cap entirely
gitnexus analyze --embeddings 100000 # custom capIf embeddings are skipped on a large repository, the indexed graph likely exceeds the default cap — re-run with --embeddings 0 or a higher limit.
Keep remote repositories indexed with gitnexus auto-sync
gitnexus auto-sync clones or pulls configured repositories, analyzes new commits, and optionally syncs their group. It runs once immediately, then repeats on the configured interval. It runs in the foreground; use your process manager if it must survive a shell session. gitnexus watch is reserved and prints this split; it does not start auto-sync or local file watching.
# 1. Create the config once. It never overwrites an existing file.
gitnexus auto-sync init
# 2. Edit $GITNEXUS_HOME/watch_config.yml, then start it.
gitnexus auto-sync start # `gitnexus auto-sync` is equivalent
gitnexus auto-sync status
gitnexus auto-sync restart # Required after config changes
gitnexus auto-sync stop
gitnexus auto-sync reset # Clear failure state; leaves clones and indexes intactGITNEXUS_HOME defaults to ~/.gitnexus. A minimal configuration:
sync_interval_minutes: 10
analyze_timeout: 5m
# Extra hosts beyond github.com, gitlab.com, and gitee.com. Exact names only.
# allowed_hosts: [gitlab.mycompany.com]
projects:
- local_path: /absolute/path/to/clones
branches: [main, master]
# pdg: omit = preserve live index mode; true = keep PDG current;
# false = init default (warns, then strips PDG on the next successful rebuild).
# Do not paste pdg: false onto an existing watch file unless you intend to drop PDG.
pdg: false
overwrite_local_changes: false
remote_urls:
- git@github.com:owner/repo.gitsync_interval_minutesmust be at least5;local_pathmust be an absolute path. Clones are stored below it ashost/namespace/repo.- Remote URLs may use SSH SCP or HTTPS. Hosts are github.com, gitlab.com, and gitee.com unless listed in top-level
allowed_hosts(exact DNS names, no wildcards). The CLI image includes OpenSSH; mount keys yourself. Invalidwatch_config.ymlskips auto-sync immediately. Auto-sync honors.gitnexusrcembeddings (HTTP embeddings env still required in the image). branchesare tried in order. The legacybranchfield is supported, but do not set both.- Set per-project
pdg: trueto keep the full control-flow, control/data-dependence, and taint layers current. Untouched configs that omitpdgpreserve an existing index's mode and cannot silently strip PDG data. Do not pastepdg: falsefrom this example onto an existing watch file unless you intend to drop PDG; an explicitfalseopt-out logs a warning before removing existing PDG data. Auto-sync requests atomic incremental publication where supported, so readers keep using the previous graph until a successful update is ready and a failed staged analysis leaves it intact; unsupported paths retain the analyzer's existing in-place behavior. - Analysis runs in an isolated worker;
analyze_timeoutdefaults to half ofsync_interval_minutes, but may be longer (for example, a30manalysis timeout with5minute polling) up to Node's timer limit. If a polling tick arrives while analysis is active, it is coalesced into one immediate follow-up run using the newest commit. If the parent times out and leaves that worker running, the follow-up is deferred to the next interval so a leftover lock holder is not counted as a hard analyze failure. Timeout andauto-sync stoprequest safe cancellation; a worker in native work exits after reaching a JS-visible safe point. Until then, auto-sync reportscancellingorstoppingand retains ownership so another auto-sync cannot take over, for up to 5 seconds — after that the parent stops waiting and leaves the worker to exit on its own rather than killing it mid-write. This behavior is the same on macOS and Windows.overwrite_local_changesdefaults tofalse, so a dirty local clone is skipped rather than overwritten; setting it totruealso deletes untracked files in the clone, while keeping ignored paths. - Add
group_nameonly after creating that group withgitnexus group create <name>. Partial clone output is isolated and removed after 14 days.
See the full auto-sync configuration and runtime reference for concurrency, timeouts, failure thresholds, and runtime files.
Repository groups (multi-repo / monorepo service tracking)
gitnexus group create <name> # Create a repository group
gitnexus group add <group> <groupPath> <registryName> # Add a repo. <groupPath> is a hierarchy path
# (e.g. hr/hiring/backend); <registryName> is the
# repo's name from the registry (see `gitnexus list`)
gitnexus group remove <group> <groupPath> # Remove a repo by its hierarchy path
gitnexus group list [name] # List groups, or show one group's config
gitnexus group sync <name> # Extract contracts and match across repos/services
gitnexus group contracts <name> # Inspect extracted contracts and cross-links
gitnexus group query <name> <q> # Search execution flows across all repos in a group
gitnexus group status <name> # Check staleness of repos in a group
gitnexus group impact <name> --target <symbol> --repo <groupPath> # Cross-repo blast radiusProject config (.gitnexusrc)
Commit a .gitnexusrc JSON file at the repo root to preconfigure recurring analyze options per project, instead of re-passing the same flags every run. It is read from the resolved repo root (not .gitnexus/, which is gitignored index storage). CLI flags always override .gitnexusrc.
A nested analyze block is also accepted (and overrides flat keys for the same option):
{ "analyze": { "defaultBranch": "develop", "skipSkills": true } }Notes:
- The default branch is resolved as:
--default-branch>.gitnexusrcdefaultBranch/branch> auto-detectedorigin/HEAD>main. skipContextFiles/skipAiContextare aliases forskipAgentsMd— they skip theAGENTS.md/CLAUDE.mdblock only. They do not implyskipSkills.indexOnlyis the stronger option that skips all file injection.- Supported keys:
defaultBranch(branch),skipAgentsMd(skipContextFiles,skipAiContext),skipSkills,indexOnly,stats/noStats,embeddings,dropEmbeddings,name,allowDuplicateName,maxFileSize,workerTimeout,walCheckpointThreshold,workers,maxProcesses,maxProcessBranching,maxProcessTraceDepth,maxEntryPointCandidates,springActuator,embeddingThreads,embeddingBatchSize,embeddingSubBatchSize,embeddingDevice. - The file is JSON only. Unknown keys and wrong JSON types fail fast with an actionable error before analysis starts. Process-detection knobs (
maxProcesses,maxProcessBranching,maxProcessTraceDepth,maxEntryPointCandidates) that are not a positive integer warn and fall through to env, then the built-in default.
Environment variables
Most analyze knobs are also CLI flags (--workers, --worker-timeout, --max-file-size, --verbose). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over .gitnexusrc, which takes precedence over env vars, which take precedence over built-in defaults.
| Variable | Default | Effect | Tune when… | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
GITNEXUS_WORKER_POOL_SIZE |
cores - 1, capped at 16 |
Parse worker pool size (must be ≥ 1). Equivalent to --workers <n>. The worker pool is the sole parse path — there is no sequential parser, so 0 is rejected with an actionable error (the pool self-heals via quarantine + respawn). |
Constrained containers (cgroup CPU limits) or CI runners with explicit quotas. To narrow down a worker crash set 1 for a single-worker pool — not 0. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_PARSE_CHUNK_CONCURRENCY |
2 |
Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_VERBOSE |
unset | When 1, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to --verbose. |
Debugging an analyze that "completed" but seems to have missed files; tuning --workers / chunk concurrency against observable throughput. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_EMBEDDING_RETRY_TIMEOUTS |
unset | When truthy (1/true/yes), per-attempt HTTP embedding timeouts (TimeoutError on fetch or body read) go through the bounded GITNEXUS_EMBEDDING_MAX_ATTEMPTS retry loop instead of failing the job. Any other value leaves it off, so cloud/default timeouts remain terminal. |
Local accelerators that drop a device lock when the client disconnects and succeed on the next request (observed with FastFlowLM on Ryzen AI). | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_EMBEDDING_SIDECAR_TIMEOUT_MS |
180000 (3 minutes) |
Per-request IPC timeout for local embedding sidecar embed batches. On overrun the parent SIGKILLs the sidecar child and rejects the batch. Init still uses the HF download budget (HF_DOWNLOAD_TIMEOUT_MS × attempts), not this knob. |
Large embed batches or slow local ONNX inference cause sidecar request timeouts during analyze --embeddings, embeddings sync, serve, or MCP. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_ANALYZER_IDENTITY_IN_PROCESS_GUARDS |
unset | When truthy (1/true/yes), forces in-process cache-guard validation once a batch has ≥128 requests. In-process mode also auto-selects when packageRoot/buildRoot fail W_OK with EACCES/EROFS. Otherwise those large batches use a Node subprocess probe. Batches under 128 always stay in-process. |
Trusted or read-only installs where two identity subprocess spawns per analyze dominate wall time; leave unset to keep the default isolation path on writable trees. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_RESOLVE_DEF_GRAPH_ID_MEMO |
on (unset) | Memoizes resolveDefGraphId per nodeLookup instance (WeakMap). Enabled by default. Set to 0/false/off/no to disable and recompute on every call (debug / bisect memo bugs). |
Suspecting stale graph-id resolution after a lookup rebuild, or comparing memo vs uncached cost on a large index. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_AUTH_TOKEN |
unset | Bearer token required when eval-server binds beyond loopback. May also be read from .env.local or .env; shell values take precedence. |
Exposing the evaluation HTTP tools to a container, VM, or LAN. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_MCP_AUTH_TOKEN |
unset | Bearer token for the dedicated gitnexus mcp --http server, for a directly reachable gitnexus serve /api/mcp route, and for the docker-server / web proxy in front of one. A non-loopback dedicated MCP bind requires it; serve enables protocol-layer MCP auth when it is set. Behind a proxy, set the same value on both services: the proxy spends the edge GITNEXUS_SERVE_AUTH_TOKEN, then replaces Authorization with this token on /api/mcp only. |
Dedicated MCP, a serve the client can reach directly, or a proxied deploy (Render Blueprint) where the backend runs protocol-layer MCP auth — configure it on the proxy too. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_PROFILE_DEFERRED |
unset | When 1, emits [deferred-profile] timing/progress logs for the post-chunk deferred resolution band (imports → heritage → buildHeritageMap → legacy call resolution). Implied by GITNEXUS_VERBOSE. |
Diagnosing analyze stalls in "Resolving calls (all chunks)" on large Java/Kotlin repos (issue #1741) without the full verbose ingestion noise. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_PROFILE_DEFERRED_SLOW_MS |
3000 (verbose) / 5000 |
Per-file threshold in ms above which processCallsFromExtracted emits a slow file … log line. Parsed via Number(): accepts integers (5000), scientific notation (2.5e3), decimals (.5), and hex (0x10). Non-finite or non-positive values fall back to the default. |
Hunting a few outlier files dominating the deferred call-resolution stage; lower to surface more, raise to focus only on the worst. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
PROF_LBUG_LOAD |
unset | When 1, emits one [lbug-load prof] summary line per loadGraphToLbug call breaking the graph-DB persistence wall into stages (csv-emit / copy-nodes / copy-rels / fallback / total) plus node & edge counts. Zero-cost when unset. |
Attributing large-repo analyze wall time across CSV generation vs. LadybugDB COPY (issue #2203) — the analyze "emit" timing is the scope-resolution bucket, not this DB-write path. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_MAX_FILE_SIZE |
512 (KB) |
Walker skip threshold in KB. Hard cap is 32768 (tree-sitter buffer ceiling). Equivalent to --max-file-size <kb>. |
Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_MAX_PROCESSES |
dynamic (max(20, round(symbols/10))) |
Analyze-time process-detection process cap. Equivalent to --max-processes <n> / .gitnexusrc maxProcesses. Explicit values replace the dynamic formula (not a multiplier). 0 is invalid, not unlimited. Changing this re-detects flows on the next analyze without --force. Distinct from query-time IMPACT_MAX_CHUNKS. |
[processes] … whole flows are MISSING names --max-processes after entry points were never traced or flows were dropped. Tracing does not start the next entry once collected traces already reach maxProcesses * 2; a started entry can still emit every trace that entry produces. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_MAX_PROCESS_BRANCHING |
4 |
Analyze-time per-node branching cap during flow tracing. Equivalent to --max-process-branching <n>. Shape-only: raising it shortens fewer traces; it does not restore whole missing flows. |
A flow is present but calleesDropped is high at debug. |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_MAX_PROCESS_TRACE_DEPTH |
10 |
Analyze-time DFS depth cap during flow tracing. Equivalent to --max-process-trace-depth <n>. Shape-only. |
A reported flow is shorter than the code path (tracesDepthCapped at debug). |
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GITNEXUS_MAX_ENTRY_POINT_CANDIDATES |
200 |
Ranked entry-point candidate pool. Equ
(README truncated) Recent activitycommits and pull requests
Recent open issuesview all
Discussionsall 33
Releases and announcements
Code frequencyadditions and deletionsCommits per weeklast 52 weeksWhen work happensweekday and hourWho is committinglast 52 weeksMaintainer commits270 (13%) Community commits1,857 (87%) 2,127 commits in total over the last year. Trending appearances
Related repositories
|
{ // Default branch used in the generated regression-compare example (base_ref). // Use this so a project on `develop`/`master` doesn't get "main" rewritten // over its fix on every analyze. (Alias: "branch".) "defaultBranch": "develop", "skipContextFiles": true, // alias of skipAgentsMd: keep your own AGENTS.md/CLAUDE.md "skipSkills": true, // don't install standard skill files under .claude/skills/ and .agents/skills/ "embeddings": true, // generate embeddings by default "springActuator": "./actuator", // optional local runtime snapshot directory or bundle "workerTimeout": 60, }