Graphmemdocs GitHub ↗

MCP stdio

Run gmem mcp as an MCP server over newline-delimited JSON-RPC. The database is the same SQLite database used by the CLI and is selected with GRAPHMEM_HOME.

The server exposes nine memory tools:

  • remember: stores content, with optional memory_type, importance, and scopes, verified entities, and verified directed relations. Entity and relation endpoints are linked to the new memory atomically. With embeddings enabled, the vectors of the memory and of any new entity or relation are stored in the same transaction; if the model cannot load or embed, the call fails and nothing is stored. Defaults are observation, 0.0, and the server repository scope (or global outside a repository). Only the current repository and global may be written; a foreign scope rejects the whole call before embedding or storage.
  • recall: searches query with optional scopes, limit (default 10), use_embeddings (default true), and memory_type. Set use_embeddings = false to force SQLite FTS5 lexical ranking for an apples-to-apples comparison without loading the embedding model. using local semantic seeds and Personalized PageRank over linked memories, entities, and relations. Omitted scopes or [] search the server repository plus global, with repository matches first. An explicit list searches those scopes plus global, without automatically adding the current repository. Reading foreign scopes does not change the current scope. The scope filter is applied before ranking, so results never include unselected repositories. memory_type narrows the candidates the same way and ignores case, so a filtered memory neither seeds nor propagates rank; it applies to both the semantic and the lexical path. Each returned memory increments access_count and updates last_accessed_at; candidates outside the response are not counted.
  • stats: returns the counts of memories, scopes, entities, and edges.
  • list_scopes: read-only discovery of all stored scopes plus the current scope and global, including when empty. Returns current_scope and a sorted, deduplicated scopes list; each entry has name, is_current, and writable. Only the current scope and global are writable. Listing does not select scopes, change the current scope, create scope rows, or record memory accesses.
  • relate: stores a directed relationship between two entities and embeds it. On an embedding error the call fails; the entities and edge may already exist without vectors, and a retry reuses them.
  • graph: inspects incoming, outgoing, or both relationship paths for an entity.
  • inspect: returns a memory and its scopes by numeric id, with optional scopes. A successful inspection increments access_count and updates last_accessed_at.
  • update: revises a memory by numeric id, with optional scopes. content, memory_type, and importance replace the stored values; omitted fields are kept. Scopes, entities, and relations are not changed. With embeddings enabled, changed content is re-embedded in the same transaction, so a model that cannot load or embed leaves the memory untouched; with embeddings disabled the stale vector is dropped and recall embeds the new content on its next semantic pass. A change that leaves the content alone never loads the model and keeps the stored vector, because memory_type and importance are not part of the embedded document.
  • forget: deletes a memory by numeric id, with optional scopes.

Code navigation tools

Binaries built with --features code (release binaries are) add four tools, listed unless [code] enabled = false or GRAPHMEM_CODE=off; the initialize instructions then add a short note on when to use them. They use a separate, rebuildable index ($GRAPHMEM_HOME/code-v{N}.sqlite, where N is the schema/extractor version) keyed by Git checkout root, so each linked worktree has its own. Different index versions use separate files; older caches, including legacy code.sqlite, are left untouched and can be deleted when no older binary uses them. They never read or write memories, and recall ranking is unaffected. A corrupt code index is recreated; if the index still cannot be opened or [code] is invalid, the server logs a warning on stderr and starts with the memory tools only.

There is no MCP index tool: the tools refresh the index on demand, re-reading only files whose size or modification time changed. To build the index of a large checkout ahead of time, or to see which files could not be indexed, run gmem code index in it.

  • code_outline: lists the definitions in path (relative to the checkout root, or absolute inside it), with optional offset, limit (1 to 500, default 200), depth (0 for top-level symbols only, 1 adds their direct children; omitted for every level), and root (an absolute directory inside another checkout; defaults to the server’s startup directory). The file is re-indexed first when it changed, so the result always matches it. Like the other code tools, the result is plain text with no structured content:

    lib/billing.ex	elixir	complete
    1-3	module Billing
    2-2	  function charge/1	def charge(amount), do: amount
    next_offset 200 of 214
    

    The first line is the path, language, and coverage: complete, or partial: <reason> when syntax errors or template limits may hide symbols. Each symbol line is its 1-based line range, a tab, two spaces per nesting level, its kind and name, and a tab and its first source line (the signature) when that shows parameters, types, or values the name does not (never for imports, whose name is the declaration). A next_offset <n> of <total> line follows when more symbols remain; with depth, both count only the symbols within that depth. Paths containing .., symlinks, and paths that resolve outside the checkout are refused.

  • code_imports: lists the imports declared in path, with optional offset and limit (1 to 500, default 200), and root using the same path rules as code_outline. The text starts with the same path, language, and coverage line, then one start-end<TAB>name line per import, and the same next_offset line when more remain. The file is refreshed before returning results. coverage is partial only when syntax errors or template limits may hide imports; complete does not mean every way of loading code is recognised. Listed forms include Rust use and extern crate; JavaScript/TypeScript import, export ... from, require(), import(), and import x = require(); Ruby require, require_relative, and load; PHP use, require, and include (with _once); Bash source and .. Loading through other APIs is not listed, so use text search when absence matters. Names are as written (quotes are kept where the language has them) and cut at 160 characters; use the line range to read a long one. Imports nested in functions or modules are listed without their scope. Imports are not resolved to files.

  • find_symbol: finds definitions by query, with optional kind, limit (1 to 100, default 20), and root. It refreshes the checkout’s index first (only changed files are re-read), so a checkout that was never indexed is indexed on the first call and results always cover every source file. Exact names come first (an Elixir name also matches name/arity), then names ending in .query (ConsentLive finds MyAppWeb.ConsentLive, and users finds public.users), then prefix matches; matching ignores case. Imports, HEEx component/slot usages, and EEx expressions are returned only when kind asks for them. Every match is returned, one line each, and none is presented as the resolved target:

    lib/billing.ex:2-2	function charge/1	in Billing	def charge(amount), do: amount
    app/service.py:9-10	function load_config	def load_config(path):
    20 of 35 matches; raise limit for more
    

    A match line is the path and 1-based line range, a tab, and the kind and name, followed by tab-separated in <parent> when the symbol is nested, stale or missing when the file changed between the refresh and the read, and the signature under the same rule as code_outline. An <n> of <total> matches line follows when limit cut matches off, and a truncated: line when the checkout has more source files than [code] max_files, so a definition may be missing. No match returns empty text.

  • code_diff: lists the symbols added, removed, or modified between two Git revisions, with base, optional head (default HEAD), limit (1 to 1000 changed files, default 200), and root. When more files changed, the first limit are reported and a final <n> of <total> changed files line says how many were left out. Like a pull request, the merge base of base and head is compared with head. Both versions of each changed file are outlined in memory; the index is neither read nor written. The result is plain text:

    merge base a9c4b016dd43 head a51919589c13
    minibot/llm/tools/bash.py	modified
      ~ 67-148	method _handle	async def _handle(self, payload: dict[str, Any], _: ToolContext) -> dict[str, Any]:
      - 215-232	method _truncate_output
    minibot/shared/subprocess_utils.py	modified
      + 11-32	function truncate_subprocess_output	def truncate_subprocess_output(
    skipped (unsupported file type): todos/ROADMAP.md
    

    Each file line is the path, a tab, and its status: added, deleted, modified, or renamed from <old path>. A third column is added when coverage is partial: <reason>. Each symbol line is +, -, or ~, its line range, a tab, its kind and name, and, for added and modified symbols, a tab and its signature when that shows parameters, types, or values the name does not. Lines refer to head, or to the merge base for removed symbols. HEEx component and slot usages inside a definition (such as a ~H sigil) are left out, so a template edit shows on the enclosing function; top-level components of .heex files are kept. The diff is structural, not semantic: a symbol is identified by its kind, its name and its ancestors' names, and its occurrence among symbols sharing those, and it is modified when its signature or its own text changed. Its own text includes the comment and attribute lines directly above it and excludes nested symbols, so an edited method is reported without its impl or module. Blank lines, repeated spaces inside a line, and the , and ; that separate nested symbols are ignored; indentation is compared, so re-indenting Python counts as a change. A renamed symbol appears as removed plus added. Changed files that are not supported, minified, binary, not UTF-8 (content or path), larger than 1 MiB, or whose type changed (for example into a symlink) are listed on skipped (<reason>): … lines. Paths with control characters are printed quoted and escaped. Revisions starting with - are refused. gmem code diff <BASE> [HEAD] prints the same text.

Supported files: Rust, Go, Zig, C, C++ (.h headers are parsed as C++), Python, JavaScript/JSX, TypeScript/TSX, Elixir (including ~H sigils), HEEx, EEx (directives only), Ruby, PHP, Racket, SQL, Bash, CSS, SCSS, and the <script> and <style> elements of HTML and HEEx. CSS rule sets are named by their selector list. Symbols are syntactic definitions and imports; nested Elixir modules are named in full (Parent.Child), and a JavaScript/TypeScript export default { ... } object is listed with its methods. Call sites and cross-file references are not indexed. Changed files are parsed in parallel by [code] index_threads threads ("auto", the default, uses the CPUs available to the process).

Memory IDs are small positive integers generated by SQLite and returned as JSON numbers. inspect uses the same selection as recall: the requested scopes plus global and legacy unscoped memories; omitted scopes or [] use the current repository. Reading an unselected repository by id reports memory not found. To read another project’s memory, select its scope explicitly.

remember, update, and forget allow writes only to the current repository and global. Explicit foreign scopes are rejected, even for a known id. Every scope attached to an existing memory must be writable: a memory shared between global or the current repository and another project remains read-only here. Legacy unscoped memories remain readable and writable like global memories. Validation happens before embedding or mutation, so a rejected batch changes nothing. Errors identify the rejected scope and the allowed current/global scopes. The CLI follows the same policy.

For example, a successful remember response contains "id": 1.

Use global for reusable knowledge and repo:/absolute/path/to/repository for project-specific knowledge. If scopes are omitted, the server resolves its startup working directory to Git’s canonical common directory (git rev-parse --git-common-dir); linked worktrees share that scope, separate clones do not. The default scope appears in the MCP initialize instructions. A bare repository uses its own Git directory as the scope; the server does not modify Git or turn a checkout into a bare repository. If Git is unavailable or the startup directory is outside a repository, the default is global.

Explicit repo: paths pointing to a checkout root or its common Git directory resolve to the same canonical scope. Other absolute paths remain literal scopes. The default is fixed at server startup: a client switching to another repository without restarting the MCP server may read that repository by selecting it explicitly, but must restart in that repository to write its memories. Explicit scopes must be global or an absolute repo: path.

Discover another project’s scope with list_scopes, then pass its name to recall or inspect, for example:

{"query":"authentication decisions","scopes":["repo:/path/to/other-project/.git"]}

That recall includes the selected project and global, not the current project.

Embeddings default to sentence-transformers/msmarco-MiniLM-L6-cos-v5 through Candle. BERT, DistilBERT, Qwen3, and ModernBERT checkpoints with CLS pooling are supported; for multilingual memories (for example Spanish queries over English memories) use ibm-granite/granite-embedding-97m-multilingual-r2. The model is downloaded and loaded on first use (the first remember, relate, or recall) and cached under models/ in the Graphmem data directory. Content is truncated to the model’s max_position_embeddings (512 tokens for the default model), capped at 2048 tokens, before embedding, so only the first tokens up to that limit of a very long memory, entity, or edge document affect its embedding. This cap bounds the quadratic attention memory used by long-context models. The initialize response’s instructions name the active model (or report embeddings disabled), and the gmem://embedding resource returns the active model, revision, and the exact max_tokens limit read from its config.json (only that file is fetched, from cache once present; max_tokens is null when embeddings are disabled). Truncation is logged as a warning and reported back in the warnings field of the remember, recall, and relate responses (an empty array when nothing was truncated). To disable embeddings and retain SQLite FTS5 lexical ranking, use either:

# ~/.graphmem/config.toml
[embedding]
enabled = false
backend = "auto"

backend accepts auto, cpu, or cuda; auto selects CUDA when the binary was built with Candle’s CUDA feature and a device is available, then falls back to CPU. Use cpu to force CPU or cuda to fail instead of falling back. Build CUDA support with cargo build --features cuda when the CUDA toolkit is installed. GRAPHMEM_EMBEDDINGS=off disables embeddings entirely. GRAPHMEM_EMBEDDING_MODEL, GRAPHMEM_EMBEDDING_REVISION, GRAPHMEM_EMBEDDING_CACHE_DIR, and GRAPHMEM_EMBEDDING_BATCH_SIZE override the remaining model settings. batch_size is how many texts go through the model per call. It defaults to 1 on CPU, where padded batches are slower, and 16 on CUDA; gmem --embedding-batch-size <n> mcp overrides it for one server, but GRAPHMEM_EMBEDDING_BATCH_SIZE still wins over the flag. The [retrieval] keys accept the same --retrieval-* flags and GRAPHMEM_RETRIEVAL_* environment variables; see cli.md for the exact precedence. If loading or inference fails, recall reports the failure on stderr and falls back to lexical ranking.

Switching model (or GRAPHMEM_EMBEDDING_MODEL) does not convert existing embeddings — different models produce incompatible vectors, so a switch means recomputing them. This happens lazily: remember, relate, and recall only re-embed what they actually touch, so memories in a repo: scope nobody queries can stay embedded under the old model indefinitely. Run gmem reembed once after changing the model to eagerly recompute every memory, entity, and edge across every scope in the store. It fails loudly (instead of silently falling back to lexical ranking) if the new model can’t load, and is safe to re-run — already-current rows are skipped. It processes every record regardless of individual failures; it prints each failure with the record’s id and exits non-zero if any occurred, but a single failing record never blocks the rest of the store from migrating.

Tokio uses four workers by default. Override it in the same file when needed:

[runtime]
worker_threads = 4

GRAPHMEM_TOKIO_WORKER_THREADS overrides that value for one process.

All process logs append to logs/graphmem.log below the active Graphmem data directory: ~/.graphmem/logs/graphmem.log by default, or $GRAPHMEM_HOME/logs/graphmem.log when that override is set. MCP protocol messages remain on stdout; the log records model-cache/download, model-ready, fallback, and command-failure events. For a live view, run:

tail -f ~/.graphmem/logs/graphmem.log

Example configuration:

{
  "mcpServers": {
    "graphmem": {
      "command": "gmem",
      "args": ["mcp"]
    }
  }
}