Graphmemdocs GitHub ↗

CLI reference

gmem is the human-facing CLI over the same SQLite store the MCP server (gmem mcp, see docs/mcp.md ) uses. Every command opens the database at ~/.graphmem by default; set GRAPHMEM_HOME to point at a different data directory (for example a separate dev store), and put config.toml there to configure embeddings ([embedding], see docs/mcp.md ) and retrieval tuning ([retrieval], see the README ). Run gmem version to print the installed binary version.

Commands that access the store open their own connection. gmem tui keeps its connection open until you quit; press r to reload changes made by another process.

Global options

FlagOverridesNotes
--embedding-batch-size <n>GRAPHMEM_EMBEDDING_BATCH_SIZE, then [embedding] batch_sizeTexts per embedding model call. Defaults to 1 on CPU, 16 on CUDA. Must be at least 1.
--retrieval-seed-top-k <n>GRAPHMEM_RETRIEVAL_SEED_TOP_K, then [retrieval] seed_top_kMemories kept as PageRank seeds. Must be at least 1.
--retrieval-seed-temperature <f>GRAPHMEM_RETRIEVAL_SEED_TEMPERATURE, then [retrieval] seed_temperatureSoftmax temperature over the top-k seeds; lower sharpens. Must be in (0, 1].
--retrieval-memory-seed-weight <f>GRAPHMEM_RETRIEVAL_MEMORY_SEED_WEIGHT, then [retrieval] memory_seed_weightShare of seed mass kept on memories vs. the graph. Must be in [0, 1].
--retrieval-entity-anchor-weight <f>GRAPHMEM_RETRIEVAL_ENTITY_ANCHOR_WEIGHT, then [retrieval] entity_anchor_weightBoost for entities named in the query. Must be in [0, 1].
--retrieval-damping <f>GRAPHMEM_RETRIEVAL_DAMPING, then [retrieval] dampingPersonalized PageRank damping. Must be in (0, 1).

All six are accepted before or after the subcommand, and apply to gmem mcp too. Resolution is per field, highest priority first:

environment variable  >  command-line flag  >  config.toml  >  built-in default

So --retrieval-damping 0.7 wins over [retrieval] damping in config.toml, but GRAPHMEM_RETRIEVAL_DAMPING=0.9 in the environment wins over both. --help shows the same overrides column. Values from every source are validated the same way, so an out-of-range flag is rejected before recall runs.

gmem remember <content>

Store a new narrative memory.

FlagDefaultNotes
--type <memory_type>observationFree-text category, e.g. decision, convention, constraint.
--importance <0.0-1.0>0.0How costly it would be for a future task to miss this.
--scope <scope> (repeatable)the current repo scope, or global outside a repoOnly the current repository and global may be attached. Foreign scopes reject the entire call.

The CLI cannot attach entities or relations to a memory, and there is no CLI equivalent of the relate tool — both are MCP-only (the remember tool’s entities/relations fields, and the relate tool).

When embeddings are enabled, the memory’s vector is computed and stored in the same transaction as the memory, loading the model on first use. If the model cannot load or embed, the command exits non-zero and nothing is stored.

Prints remembered: <id>.

gmem remember "Use cargo nextest for integration tests" \
  --type convention --importance 0.7 --scope repo:/absolute/path/to/project

gmem list

List memories, newest first, no ranking.

FlagDefault
--scope <scope> (repeatable)current repository + global
--limit <n>50

Explicit scopes select those projects plus global, without adding the current repository. Checkout roots and common Git directories resolve to the same scope.

Output: one line per memory, tab-separated id, memory_type, content.

gmem show <id>

Print one memory’s full detail (id, type, importance, timestamps, scopes, content). A successful show updates last_accessed_at and increments access_count; list does not. Exits non-zero with memory not found if the id doesn’t exist or belongs to an unselected project. Repeat --scope <scope> to read another project’s memory; global memories are always included. Without that flag, only the current repository plus global is selected.

gmem search <query>

Semantic recall (embeddings + graph context via Personalized PageRank). If embeddings are disabled, or the model fails to load, it reports the failure on stderr and falls back to SQLite FTS5 lexical ranking — the same fallback the MCP recall tool uses. Unlike recall, the CLI has no flag to force lexical search. Each returned result updates last_accessed_at and increments access_count.

FlagDefaultNotes
--scope <scope> (repeatable)current repository + globalExplicit scopes replace the project selection, but still include global. Reads do not change the destination of later writes.
--limit <n>10

Output: one line per result, tab-separated score, id, memory_type, content.

gmem graph <kind> <name>

Inspect the unscoped entity graph around one entity (does not search narrative memories).

FlagDefaultRange
--direction <incoming|outgoing|both>both
--max-depth <n>11-3
--limit <n>251-100

Output: first line is the entity (kind, name, canonical_name), then one line per hop per path (depth, direction, relation, entity kind, entity name), paths separated by blank lines.

gmem forget <id>

Permanently delete one memory. Prints forgot: <id>. Exits non-zero with memory not found if the id doesn’t exist. Only memories whose scopes are all current-repository or global may be deleted. A known foreign id, including a memory shared with global, produces a clear read-only-scope error without changes.

gmem tui

Browse memories and unscoped graph nodes in an interactive terminal. The memory list shows the current repository plus global, as gmem list does. The graph pane shows each node’s incoming and outgoing relations, up to 100 per node. Memory details render common Markdown formatting with terminal colors.

KeyAction
TabSwitch between memories and graph nodes
j / k, arrow keysMove through the list
PageUp / PageDownScroll the detail pane
/Filter memories by text or type as you type; Enter keeps the filter, Esc restores the previous filter, Ctrl-U clears it
eEdit the selected memory’s text with $VISUAL, then $EDITOR
dAsk to delete the selected memory or graph node; y confirms, n or Esc cancels
rReload both lists
qQuit

Deleting a graph node also removes its edges and links to memories, but keeps the memories. Editing preserves a memory’s type and importance. The editor command may include arguments and runs through a POSIX shell.

gmem flush

Delete everything — all memories, scopes, entities, and edges in the active data directory, provided every memory is writable from the current repository. If any memory belongs to a foreign scope (including shared memories), the entire operation is refused without deletion. Requires --yes; without it, exits non-zero with refusing to flush; rerun with --yes. There is no undo.

gmem scopes

List every scope in the store: one line per scope, tab-separated id, name.

gmem reembed

Eagerly recompute every memory, entity, and edge embedding under the currently configured model, across every scope in the store (not just the current repo + global). This is the migration command for switching embedding models.

Why it’s needed: embeddings otherwise only get (re)computed lazily, as a side effect of whatever remember/relate/recall happens to touch. recall/search re-embeds every entity and edge on every call (both are unscoped), but for memories it only visits the scopes it was asked about — so after changing model, a memory in a repo: scope nobody queries stays embedded under the old model indefinitely. Worse, search/recall treats a failed embed as soft failure and silently falls back to lexical ranking, so a broken model swap can go unnoticed. gmem reembed fails loudly instead (see below) and reaches the whole store in one explicit pass.

There is no vector conversion — different models produce incompatible vector spaces, so “migrate” always means recompute, never convert. Each embedding row is looked up and stored keyed by (model, revision) (memory_id/entity_id/edge_id is the primary key, model/revision are plain columns updated in place — see src/infrastructure/sqlite.rs), so switching models just overwrites the row for that id; nothing needs deleting first.

Batched: records are embedded --embedding-batch-size at a time. If a batch fails, its records are retried one by one so each failure is still reported individually. Progress goes to stderr: first a message before loading (or downloading) the model, then memories: 20/100 processed, followed by entities and edges. There is no percentage for model download; progress is reported at the start and end of each type, and otherwise at most once every two seconds. Counts include items already cached and items whose embedding failed; the final summary on stdout and any failures retain their existing format.

Idempotent and safe to re-run: every item is cache-checked against (model, revision) before being recomputed, so re-running gmem reembed after the store is already current costs one SELECT per row, not a recompute.

Fails loudly, on purpose — unlike search, it does not fall back to lexical ranking. If [embedding] enabled = false (or GRAPHMEM_EMBEDDINGS=off), it exits non-zero with embeddings are disabled; set [embedding] enabled = true (or unset GRAPHMEM_EMBEDDINGS) to reembed. If the configured model fails to download or load, it exits non-zero with that error instead of silently doing nothing.

One failing record never blocks the rest of the store: reembed still visits every memory, entity, and edge even if some individual record fails to embed. Each failure is printed with the record’s kind and id, and the command exits non-zero if any occurred, but every other record still gets migrated in the same pass — re-run gmem reembed afterward and only the still-failing records are retried (everything else is already current and skipped per the idempotency check above).

Typical migration, after selecting the persistent model in config.toml (the default is sentence-transformers/msmarco-MiniLM-L6-cos-v5):

gmem reembed
# reembedded 14 memories, 16 entities, 7 edges under the current embedding model

If some records fail:

gmem reembed
# reembedded 13 memories, 16 entities, 7 edges under the current embedding model
# 1 item(s) failed to reembed:
#   memory 42: model inference failed: ...

Note that revision in the cache key doesn’t include the compute backend (CPU/CUDA) — switching backend alone does not invalidate existing embeddings (correctly: the resulting vector is expected to be device-independent), so gmem reembed after only a backend change is a near-instant no-op, not a recompute.

Not exposed as an MCP tool — this is a one-shot operator/maintenance action taken after a config change, not something an MCP client needs mid-session.

gmem doctor

Print the resolved database path and a health status line. Currently a fixed status: healthy if the database opened at all — there is no deeper check yet. Store counts are available via the stats MCP tool; there is no CLI equivalent.

gmem code

Only in binaries built with --features code. Disabled by [code] enabled = false in config.toml or GRAPHMEM_CODE=off. It works on the Git checkout containing the current directory; each linked worktree is indexed separately into $GRAPHMEM_HOME/code-v{N}.sqlite (N is the schema/extractor version), which can be deleted and rebuilt at any time. Different index versions use separate files. Older caches, including legacy code.sqlite, are left untouched; delete them only when no older binary uses them.

  • gmem code index [PATH] indexes the tracked and untracked, non-ignored source files of the checkout containing PATH (default: the current directory). Unchanged files are skipped by size and modification time, and files that disappeared are removed. Changed files are parsed in parallel by [code] index_threads threads ("auto" by default: the CPUs available to the process; also GRAPHMEM_CODE_INDEX_THREADS). Symlinks, binary files, *.min.js, files over 1 MiB, and anything past the first [code] max_files source files (default 20,000; also GRAPHMEM_CODE_MAX_FILES) are not indexed. Progress goes to stderr.
  • gmem code outline <FILE> prints path, language, and coverage, then one line per symbol as start-end<TAB>kind name, indented by nesting, followed by a tab and the symbol’s first source line when that shows parameters, types, or values the name does not. The file is re-indexed first if it changed.
  • gmem code imports <FILE> prints path, language, and coverage, then one start-end<TAB>name line per declared import. The file is re-indexed first if it changed. Imports are reported as written (quotes included where the language has them), cut at 160 characters, and not resolved to files. Only import forms the parser recognises are listed, even when coverage is complete; see code_imports in mcp.md for the forms, and use text search for loading through other APIs.
  • gmem code find <QUERY> [--kind KIND] [--limit N] refreshes the index, then prints path:start-end<TAB>kind name for each match, followed by tab-separated in <parent> when it is nested, stale or missing when the file changed during the read, and the signature under the same rule as outline: exact names first, then names ending in .QUERY, then prefixes. Imports and template usages (import, component, slot, expression) appear only with --kind. When more matches exist than --limit, a last line says N of TOTAL matches; raise limit for more.
gmem code index
gmem code outline lib/demo_web/components/core_components.ex
# lib/demo_web/components/core_components.ex	elixir	complete
# 1-40	module DemoWeb.CoreComponents
# 8-15	  function flash/1	def flash(assigns) do
# 11-11	    component .icon	<.icon name="hero-x-mark" />
gmem code find flash
# lib/demo_web/components/core_components.ex:8-15	function flash/1	in DemoWeb.CoreComponents	def flash(assigns) do

gmem mcp

Run the MCP server over stdio (newline-delimited JSON-RPC 2.0). See docs/mcp.md for the tool contract, and the gmem-development skill’s scripts/mcp_smoke.py for a dependency-free way to smoke-test it by hand.