MCP stdio
Run gmem mcp as an MCP server over newline-delimited JSON-RPC. The
database is the same SQLite database used by the CLI and is selected with
GRAPHMEM_HOME.
The server exposes nine memory tools:
remember: storescontent, with optionalmemory_type,importance, andscopes, verifiedentities, and verified directedrelations. Entity and relation endpoints are linked to the new memory atomically. With embeddings enabled, the vectors of the memory and of any new entity or relation are stored in the same transaction; if the model cannot load or embed, the call fails and nothing is stored. Defaults areobservation,0.0, and the server repository scope (orglobaloutside a repository). Only the current repository andglobalmay be written; a foreign scope rejects the whole call before embedding or storage.recall: searchesquerywith optionalscopes,limit(default 10),use_embeddings(defaulttrue), andmemory_type. Setuse_embeddings = falseto force SQLite FTS5 lexical ranking for an apples-to-apples comparison without loading the embedding model. using local semantic seeds and Personalized PageRank over linked memories, entities, and relations. Omitted scopes or[]search the server repository plusglobal, with repository matches first. An explicit list searches those scopes plusglobal, without automatically adding the current repository. Reading foreign scopes does not change the current scope. The scope filter is applied before ranking, so results never include unselected repositories.memory_typenarrows the candidates the same way and ignores case, so a filtered memory neither seeds nor propagates rank; it applies to both the semantic and the lexical path. Each returned memory incrementsaccess_countand updateslast_accessed_at; candidates outside the response are not counted.stats: returns the counts of memories, scopes, entities, and edges.list_scopes: read-only discovery of all stored scopes plus the current scope andglobal, including when empty. Returnscurrent_scopeand a sorted, deduplicatedscopeslist; each entry hasname,is_current, andwritable. Only the current scope andglobalare writable. Listing does not select scopes, change the current scope, create scope rows, or record memory accesses.relate: stores a directed relationship between two entities and embeds it. On an embedding error the call fails; the entities and edge may already exist without vectors, and a retry reuses them.graph: inspects incoming, outgoing, or both relationship paths for an entity.inspect: returns a memory and its scopes by numericid, with optionalscopes. A successful inspection incrementsaccess_countand updateslast_accessed_at.update: revises a memory by numericid, with optionalscopes.content,memory_type, andimportancereplace the stored values; omitted fields are kept. Scopes, entities, and relations are not changed. With embeddings enabled, changed content is re-embedded in the same transaction, so a model that cannot load or embed leaves the memory untouched; with embeddings disabled the stale vector is dropped and recall embeds the new content on its next semantic pass. A change that leaves the content alone never loads the model and keeps the stored vector, becausememory_typeandimportanceare not part of the embedded document.forget: deletes a memory by numericid, with optionalscopes.
Code navigation tools
Binaries built with --features code (release binaries are) add four tools,
listed unless [code] enabled = false or GRAPHMEM_CODE=off; the
initialize instructions then add a short note on when to use them. They use
a separate, rebuildable index ($GRAPHMEM_HOME/code-v{N}.sqlite, where N
is the schema/extractor version) keyed by Git checkout root, so each linked
worktree has its own. Different index versions use separate files; older
caches, including legacy code.sqlite, are left untouched and can be deleted
when no older binary uses them. They never read or write memories, and
recall ranking is unaffected. A corrupt code index is recreated; if the
index still cannot be opened or [code] is invalid, the server logs a warning
on stderr and starts with the memory tools only.
There is no MCP index tool: the tools refresh the index on demand, re-reading
only files whose size or modification time changed. To build the index of a
large checkout ahead of time, or to see which files could not be indexed, run
gmem code index in it.
code_outline: lists the definitions inpath(relative to the checkout root, or absolute inside it), with optionaloffset,limit(1 to 500, default 200),depth(0 for top-level symbols only, 1 adds their direct children; omitted for every level), androot(an absolute directory inside another checkout; defaults to the server’s startup directory). The file is re-indexed first when it changed, so the result always matches it. Like the other code tools, the result is plain text with no structured content:lib/billing.ex elixir complete 1-3 module Billing 2-2 function charge/1 def charge(amount), do: amount next_offset 200 of 214The first line is the path, language, and
coverage:complete, orpartial: <reason>when syntax errors or template limits may hide symbols. Each symbol line is its 1-based line range, a tab, two spaces per nesting level, its kind and name, and a tab and its first source line (the signature) when that shows parameters, types, or values the name does not (never for imports, whose name is the declaration). Anext_offset <n> of <total>line follows when more symbols remain; withdepth, both count only the symbols within that depth. Paths containing.., symlinks, and paths that resolve outside the checkout are refused.code_imports: lists the imports declared inpath, with optionaloffsetandlimit(1 to 500, default 200), androotusing the same path rules ascode_outline. The text starts with the same path, language, and coverage line, then onestart-end<TAB>nameline per import, and the samenext_offsetline when more remain. The file is refreshed before returning results.coverageispartialonly when syntax errors or template limits may hide imports;completedoes not mean every way of loading code is recognised. Listed forms include Rustuseandextern crate; JavaScript/TypeScriptimport,export ... from,require(),import(), andimport x = require(); Rubyrequire,require_relative, andload; PHPuse,require, andinclude(with_once); Bashsourceand.. Loading through other APIs is not listed, so use text search when absence matters. Names are as written (quotes are kept where the language has them) and cut at 160 characters; use the line range to read a long one. Imports nested in functions or modules are listed without their scope. Imports are not resolved to files.find_symbol: finds definitions byquery, with optionalkind,limit(1 to 100, default 20), androot. It refreshes the checkout’s index first (only changed files are re-read), so a checkout that was never indexed is indexed on the first call and results always cover every source file. Exact names come first (an Elixirnamealso matchesname/arity), then names ending in.query(ConsentLivefindsMyAppWeb.ConsentLive, andusersfindspublic.users), then prefix matches; matching ignores case. Imports, HEExcomponent/slotusages, and EExexpressions are returned only whenkindasks for them. Every match is returned, one line each, and none is presented as the resolved target:lib/billing.ex:2-2 function charge/1 in Billing def charge(amount), do: amount app/service.py:9-10 function load_config def load_config(path): 20 of 35 matches; raise limit for moreA match line is the path and 1-based line range, a tab, and the kind and name, followed by tab-separated
in <parent>when the symbol is nested,staleormissingwhen the file changed between the refresh and the read, and the signature under the same rule ascode_outline. An<n> of <total> matchesline follows whenlimitcut matches off, and atruncated:line when the checkout has more source files than[code] max_files, so a definition may be missing. No match returns empty text.code_diff: lists the symbolsadded,removed, ormodifiedbetween two Git revisions, withbase, optionalhead(defaultHEAD),limit(1 to 1000 changed files, default 200), androot. When more files changed, the firstlimitare reported and a final<n> of <total> changed filesline says how many were left out. Like a pull request, the merge base ofbaseandheadis compared withhead. Both versions of each changed file are outlined in memory; the index is neither read nor written. The result is plain text:merge base a9c4b016dd43 head a51919589c13 minibot/llm/tools/bash.py modified ~ 67-148 method _handle async def _handle(self, payload: dict[str, Any], _: ToolContext) -> dict[str, Any]: - 215-232 method _truncate_output minibot/shared/subprocess_utils.py modified + 11-32 function truncate_subprocess_output def truncate_subprocess_output( skipped (unsupported file type): todos/ROADMAP.mdEach file line is the path, a tab, and its status:
added,deleted,modified, orrenamed from <old path>. A third column is added when coverage ispartial: <reason>. Each symbol line is+,-, or~, its line range, a tab, its kind and name, and, for added and modified symbols, a tab and its signature when that shows parameters, types, or values the name does not. Lines refer tohead, or to the merge base for removed symbols. HEEx component and slot usages inside a definition (such as a~Hsigil) are left out, so a template edit shows on the enclosing function; top-level components of.heexfiles are kept. The diff is structural, not semantic: a symbol is identified by its kind, its name and its ancestors' names, and its occurrence among symbols sharing those, and it is modified when its signature or its own text changed. Its own text includes the comment and attribute lines directly above it and excludes nested symbols, so an edited method is reported without itsimplor module. Blank lines, repeated spaces inside a line, and the,and;that separate nested symbols are ignored; indentation is compared, so re-indenting Python counts as a change. A renamed symbol appears as removed plus added. Changed files that are not supported, minified, binary, not UTF-8 (content or path), larger than 1 MiB, or whose type changed (for example into a symlink) are listed onskipped (<reason>): …lines. Paths with control characters are printed quoted and escaped. Revisions starting with-are refused.gmem code diff <BASE> [HEAD]prints the same text.
Supported files: Rust, Go, Zig, C, C++ (.h headers are parsed as C++),
Python, JavaScript/JSX, TypeScript/TSX, Elixir (including ~H sigils), HEEx,
EEx (directives only), Ruby, PHP, Racket, SQL, Bash, CSS, SCSS, and the <script> and
<style> elements of HTML and HEEx. CSS rule sets are named by their selector
list. Symbols are syntactic definitions and imports; nested Elixir modules
are named in full (Parent.Child), and a JavaScript/TypeScript
export default { ... } object is listed with its methods. Call sites and
cross-file references are not indexed. Changed files are parsed in parallel
by [code] index_threads threads ("auto", the default, uses the CPUs
available to the process).
Memory IDs are small positive integers generated by SQLite and returned as JSON
numbers. inspect uses the same selection as recall: the requested scopes plus
global and legacy unscoped memories; omitted scopes or [] use the current
repository. Reading an unselected repository by id reports memory not found.
To read another project’s memory, select its scope explicitly.
remember, update, and forget allow writes only to the current repository
and global. Explicit foreign scopes are rejected, even for a known id. Every
scope attached to an existing memory must be writable: a memory shared between
global or the current repository and another project remains read-only here.
Legacy unscoped memories remain readable and writable like global memories.
Validation happens before embedding or mutation, so a rejected batch changes
nothing. Errors identify the rejected scope and the allowed current/global
scopes. The CLI follows the same policy.
For example, a successful remember response contains "id": 1.
Use global for reusable knowledge and repo:/absolute/path/to/repository
for project-specific knowledge. If scopes are omitted, the server resolves its
startup working directory to Git’s canonical common directory (git rev-parse --git-common-dir); linked worktrees share that scope, separate clones do not.
The default scope appears in the MCP initialize instructions. A bare repository
uses its own Git directory as the scope; the server does not modify Git or turn
a checkout into a bare repository. If Git is unavailable or the startup directory
is outside a repository, the default is global.
Explicit repo: paths pointing to a checkout root or its common Git directory
resolve to the same canonical scope. Other absolute paths remain literal scopes.
The default is fixed at server startup: a client switching to another repository
without restarting the MCP server may read that repository by selecting it
explicitly, but must restart in that repository to write its memories.
Explicit scopes must be global or an absolute repo: path.
Discover another project’s scope with list_scopes, then pass its name to
recall or inspect, for example:
{"query":"authentication decisions","scopes":["repo:/path/to/other-project/.git"]}
That recall includes the selected project and global, not the current project.
Embeddings default to sentence-transformers/msmarco-MiniLM-L6-cos-v5 through Candle. BERT,
DistilBERT, Qwen3, and ModernBERT checkpoints with CLS pooling are supported;
for multilingual memories (for example Spanish queries over English memories)
use ibm-granite/granite-embedding-97m-multilingual-r2. The model is
downloaded and loaded on first use (the first remember, relate, or
recall) and cached under models/ in the Graphmem data directory. Content is truncated to the model’s
max_position_embeddings (512 tokens for the default model), capped at 2048
tokens, before embedding, so only the first tokens up to that limit of a very
long memory, entity, or edge document affect its embedding. This cap bounds
the quadratic attention memory used by long-context models. The initialize response’s instructions
name the active model (or report embeddings disabled), and the
gmem://embedding resource returns the active model, revision, and the exact
max_tokens limit read from its config.json (only that file is fetched, from
cache once present; max_tokens is null when embeddings are disabled).
Truncation is logged as
a warning and reported back in the warnings field of the remember,
recall, and relate responses (an empty array when nothing was truncated). To
disable embeddings and retain SQLite
FTS5 lexical ranking, use either:
# ~/.graphmem/config.toml
[embedding]
enabled = false
backend = "auto"
backend accepts auto, cpu, or cuda; auto selects CUDA when the
binary was built with Candle’s CUDA feature and a device is available, then
falls back to CPU. Use cpu to force CPU or cuda to fail instead of falling
back. Build CUDA support with cargo build --features cuda when the CUDA
toolkit is installed. GRAPHMEM_EMBEDDINGS=off disables embeddings entirely.
GRAPHMEM_EMBEDDING_MODEL,
GRAPHMEM_EMBEDDING_REVISION, GRAPHMEM_EMBEDDING_CACHE_DIR, and
GRAPHMEM_EMBEDDING_BATCH_SIZE override the remaining model settings.
batch_size is how many texts go through the model per call. It defaults to
1 on CPU, where padded batches are slower, and 16 on CUDA;
gmem --embedding-batch-size <n> mcp overrides it for one server, but
GRAPHMEM_EMBEDDING_BATCH_SIZE still wins over the flag. The [retrieval]
keys accept the same --retrieval-* flags and GRAPHMEM_RETRIEVAL_*
environment variables; see cli.md
for the exact precedence. If
loading or inference fails, recall reports the failure on stderr and falls
back to lexical ranking.
Switching model (or GRAPHMEM_EMBEDDING_MODEL) does not convert existing
embeddings — different models produce incompatible vectors, so a switch means
recomputing them. This happens lazily: remember, relate, and recall
only re-embed what they actually touch, so memories in a repo: scope
nobody queries can stay embedded under the old model indefinitely. Run
gmem reembed once after changing the model to eagerly recompute every
memory, entity, and edge across every scope in the store. It fails loudly
(instead of silently falling back to lexical ranking) if the new model can’t
load, and is safe to re-run — already-current rows are skipped. It processes
every record regardless of individual failures; it prints each failure with
the record’s id and exits non-zero if any occurred, but a single failing
record never blocks the rest of the store from migrating.
Tokio uses four workers by default. Override it in the same file when needed:
[runtime]
worker_threads = 4
GRAPHMEM_TOKIO_WORKER_THREADS overrides that value for one process.
All process logs append to logs/graphmem.log below the active Graphmem data
directory: ~/.graphmem/logs/graphmem.log by default, or
$GRAPHMEM_HOME/logs/graphmem.log when that override is set. MCP protocol
messages remain on stdout; the log records model-cache/download, model-ready,
fallback, and command-failure events. For a live view, run:
tail -f ~/.graphmem/logs/graphmem.log
Example configuration:
{
"mcpServers": {
"graphmem": {
"command": "gmem",
"args": ["mcp"]
}
}
}