CLI reference
gmem is the human-facing CLI over the same SQLite store the MCP server
(gmem mcp, see docs/mcp.md
) uses. Every command opens the
database at ~/.graphmem by default; set GRAPHMEM_HOME to point at a
different data directory (for example a separate dev store), and put
config.toml there to configure embeddings ([embedding], see
docs/mcp.md
) and retrieval tuning ([retrieval], see the
README
). Run gmem version to print the installed
binary version.
Commands that access the store open their own connection. gmem tui keeps its connection open until
you quit; press r to reload changes made by another process.
Global options
| Flag | Overrides | Notes |
|---|---|---|
--embedding-batch-size <n> | GRAPHMEM_EMBEDDING_BATCH_SIZE, then [embedding] batch_size | Texts per embedding model call. Defaults to 1 on CPU, 16 on CUDA. Must be at least 1. |
--retrieval-seed-top-k <n> | GRAPHMEM_RETRIEVAL_SEED_TOP_K, then [retrieval] seed_top_k | Memories kept as PageRank seeds. Must be at least 1. |
--retrieval-seed-temperature <f> | GRAPHMEM_RETRIEVAL_SEED_TEMPERATURE, then [retrieval] seed_temperature | Softmax temperature over the top-k seeds; lower sharpens. Must be in (0, 1]. |
--retrieval-memory-seed-weight <f> | GRAPHMEM_RETRIEVAL_MEMORY_SEED_WEIGHT, then [retrieval] memory_seed_weight | Share of seed mass kept on memories vs. the graph. Must be in [0, 1]. |
--retrieval-entity-anchor-weight <f> | GRAPHMEM_RETRIEVAL_ENTITY_ANCHOR_WEIGHT, then [retrieval] entity_anchor_weight | Boost for entities named in the query. Must be in [0, 1]. |
--retrieval-damping <f> | GRAPHMEM_RETRIEVAL_DAMPING, then [retrieval] damping | Personalized PageRank damping. Must be in (0, 1). |
All six are accepted before or after the subcommand, and apply to gmem mcp
too. Resolution is per field, highest priority first:
environment variable > command-line flag > config.toml > built-in default
So --retrieval-damping 0.7 wins over [retrieval] damping in
config.toml, but GRAPHMEM_RETRIEVAL_DAMPING=0.9 in the environment wins
over both. --help shows the same overrides column. Values from every
source are validated the same way, so an out-of-range flag is rejected
before recall runs.
gmem remember <content>
Store a new narrative memory.
| Flag | Default | Notes |
|---|---|---|
--type <memory_type> | observation | Free-text category, e.g. decision, convention, constraint. |
--importance <0.0-1.0> | 0.0 | How costly it would be for a future task to miss this. |
--scope <scope> (repeatable) | the current repo scope, or global outside a repo | Only the current repository and global may be attached. Foreign scopes reject the entire call. |
The CLI cannot attach entities or relations to a memory, and there is no CLI
equivalent of the relate tool — both are MCP-only (the remember tool’s
entities/relations fields, and the relate tool).
When embeddings are enabled, the memory’s vector is computed and stored in the same transaction as the memory, loading the model on first use. If the model cannot load or embed, the command exits non-zero and nothing is stored.
Prints remembered: <id>.
gmem remember "Use cargo nextest for integration tests" \
--type convention --importance 0.7 --scope repo:/absolute/path/to/project
gmem list
List memories, newest first, no ranking.
| Flag | Default |
|---|---|
--scope <scope> (repeatable) | current repository + global |
--limit <n> | 50 |
Explicit scopes select those projects plus global, without adding the current
repository. Checkout roots and common Git directories resolve to the same scope.
Output: one line per memory, tab-separated id, memory_type, content.
gmem show <id>
Print one memory’s full detail (id, type, importance, timestamps, scopes,
content). A successful show updates last_accessed_at and increments
access_count; list does not. Exits non-zero with memory not found if the
id doesn’t exist or belongs to an unselected project. Repeat --scope <scope>
to read another project’s memory; global memories are always included. Without
that flag, only the current repository plus global is selected.
gmem search <query>
Semantic recall (embeddings + graph context via Personalized PageRank). If
embeddings are disabled, or the model fails to load, it reports the failure
on stderr and falls back to SQLite FTS5 lexical ranking — the same fallback
the MCP recall tool uses. Unlike recall, the CLI has no flag to force
lexical search.
Each returned result updates last_accessed_at and increments access_count.
| Flag | Default | Notes |
|---|---|---|
--scope <scope> (repeatable) | current repository + global | Explicit scopes replace the project selection, but still include global. Reads do not change the destination of later writes. |
--limit <n> | 10 |
Output: one line per result, tab-separated score, id, memory_type,
content.
gmem graph <kind> <name>
Inspect the unscoped entity graph around one entity (does not search narrative memories).
| Flag | Default | Range |
|---|---|---|
--direction <incoming|outgoing|both> | both | |
--max-depth <n> | 1 | 1-3 |
--limit <n> | 25 | 1-100 |
Output: first line is the entity (kind, name, canonical_name), then one
line per hop per path (depth, direction, relation, entity kind,
entity name), paths separated by blank lines.
gmem forget <id>
Permanently delete one memory. Prints forgot: <id>. Exits non-zero with
memory not found if the id doesn’t exist. Only memories whose scopes are all
current-repository or global may be deleted. A known foreign id, including a
memory shared with global, produces a clear read-only-scope error without changes.
gmem tui
Browse memories and unscoped graph nodes in an interactive terminal. The
memory list shows the current repository plus global, as gmem list does. The graph pane shows each
node’s incoming and outgoing relations, up to 100 per node. Memory details
render common Markdown formatting with terminal colors.
| Key | Action |
|---|---|
Tab | Switch between memories and graph nodes |
j / k, arrow keys | Move through the list |
PageUp / PageDown | Scroll the detail pane |
/ | Filter memories by text or type as you type; Enter keeps the filter, Esc restores the previous filter, Ctrl-U clears it |
e | Edit the selected memory’s text with $VISUAL, then $EDITOR |
d | Ask to delete the selected memory or graph node; y confirms, n or Esc cancels |
r | Reload both lists |
q | Quit |
Deleting a graph node also removes its edges and links to memories, but keeps the memories. Editing preserves a memory’s type and importance. The editor command may include arguments and runs through a POSIX shell.
gmem flush
Delete everything — all memories, scopes, entities, and edges in the
active data directory, provided every memory is writable from the current
repository. If any memory belongs to a foreign scope (including shared memories),
the entire operation is refused without deletion. Requires --yes; without it,
exits non-zero with refusing to flush; rerun with --yes. There is no undo.
gmem scopes
List every scope in the store: one line per scope, tab-separated id,
name.
gmem reembed
Eagerly recompute every memory, entity, and edge embedding under the currently configured model, across every scope in the store (not just the current repo + global). This is the migration command for switching embedding models.
Why it’s needed: embeddings otherwise only get (re)computed lazily, as a
side effect of whatever remember/relate/recall happens to touch.
recall/search re-embeds every entity and edge on every call (both are
unscoped), but for memories it only visits the scopes it was asked about —
so after changing model, a memory in a repo: scope nobody queries stays
embedded under the old model indefinitely. Worse, search/recall treats a
failed embed as soft failure and silently falls back to lexical ranking, so
a broken model swap can go unnoticed. gmem reembed fails loudly instead
(see below) and reaches the whole store in one explicit pass.
There is no vector conversion — different models produce incompatible
vector spaces, so “migrate” always means recompute, never convert. Each
embedding row is looked up and stored keyed by (model, revision)
(memory_id/entity_id/edge_id is the primary key, model/revision
are plain columns updated in place — see src/infrastructure/sqlite.rs), so
switching models just overwrites the row for that id; nothing needs
deleting first.
Batched: records are embedded --embedding-batch-size at a time. If a
batch fails, its records are retried one by one so each failure is still
reported individually. Progress goes to stderr: first a message before loading
(or downloading) the model, then memories: 20/100 processed, followed by
entities and edges. There is no percentage for model download; progress is
reported at the start and end of each type, and otherwise at most once every
two seconds. Counts include items already cached and items whose embedding
failed; the final summary on stdout and any failures retain their existing
format.
Idempotent and safe to re-run: every item is cache-checked against
(model, revision) before being recomputed, so re-running gmem reembed
after the store is already current costs one SELECT per row, not a
recompute.
Fails loudly, on purpose — unlike search, it does not fall back to
lexical ranking. If [embedding] enabled = false (or
GRAPHMEM_EMBEDDINGS=off), it exits non-zero with embeddings are disabled; set [embedding] enabled = true (or unset GRAPHMEM_EMBEDDINGS) to reembed. If the configured model fails to download or load, it exits
non-zero with that error instead of silently doing nothing.
One failing record never blocks the rest of the store: reembed still
visits every memory, entity, and edge even if some individual record fails
to embed. Each failure is printed with the record’s kind and id, and the
command exits non-zero if any occurred, but every other record still gets
migrated in the same pass — re-run gmem reembed afterward and only the
still-failing records are retried (everything else is already current and
skipped per the idempotency check above).
Typical migration, after selecting the persistent model in config.toml
(the default is sentence-transformers/msmarco-MiniLM-L6-cos-v5):
gmem reembed
# reembedded 14 memories, 16 entities, 7 edges under the current embedding model
If some records fail:
gmem reembed
# reembedded 13 memories, 16 entities, 7 edges under the current embedding model
# 1 item(s) failed to reembed:
# memory 42: model inference failed: ...
Note that revision in the cache key doesn’t include the compute backend
(CPU/CUDA) — switching backend alone does not invalidate existing
embeddings (correctly: the resulting vector is expected to be
device-independent), so gmem reembed after only a backend change is a
near-instant no-op, not a recompute.
Not exposed as an MCP tool — this is a one-shot operator/maintenance action taken after a config change, not something an MCP client needs mid-session.
gmem doctor
Print the resolved database path and a health status line. Currently a
fixed status: healthy if the database opened at all — there is no deeper
check yet. Store counts are available via the stats MCP tool; there is no
CLI equivalent.
gmem code
Only in binaries built with --features code. Disabled by [code] enabled = false in config.toml or GRAPHMEM_CODE=off. It works on the Git checkout
containing the current directory; each linked worktree is indexed separately
into $GRAPHMEM_HOME/code-v{N}.sqlite (N is the schema/extractor version),
which can be deleted and rebuilt at any time. Different index versions use
separate files. Older caches, including legacy code.sqlite, are left untouched;
delete them only when no older binary uses them.
gmem code index [PATH]indexes the tracked and untracked, non-ignored source files of the checkout containingPATH(default: the current directory). Unchanged files are skipped by size and modification time, and files that disappeared are removed. Changed files are parsed in parallel by[code] index_threadsthreads ("auto"by default: the CPUs available to the process; alsoGRAPHMEM_CODE_INDEX_THREADS). Symlinks, binary files,*.min.js, files over 1 MiB, and anything past the first[code] max_filessource files (default 20,000; alsoGRAPHMEM_CODE_MAX_FILES) are not indexed. Progress goes to stderr.gmem code outline <FILE>printspath, language, and coverage, then one line per symbol asstart-end<TAB>kind name, indented by nesting, followed by a tab and the symbol’s first source line when that shows parameters, types, or values the name does not. The file is re-indexed first if it changed.gmem code imports <FILE>printspath, language, and coverage, then onestart-end<TAB>nameline per declared import. The file is re-indexed first if it changed. Imports are reported as written (quotes included where the language has them), cut at 160 characters, and not resolved to files. Only import forms the parser recognises are listed, even when coverage iscomplete; seecode_importsin mcp.md for the forms, and use text search for loading through other APIs.gmem code find <QUERY> [--kind KIND] [--limit N]refreshes the index, then printspath:start-end<TAB>kind namefor each match, followed by tab-separatedin <parent>when it is nested,staleormissingwhen the file changed during the read, and the signature under the same rule asoutline: exact names first, then names ending in.QUERY, then prefixes. Imports and template usages (import,component,slot,expression) appear only with--kind. When more matches exist than--limit, a last line saysN of TOTAL matches; raise limit for more.
gmem code index
gmem code outline lib/demo_web/components/core_components.ex
# lib/demo_web/components/core_components.ex elixir complete
# 1-40 module DemoWeb.CoreComponents
# 8-15 function flash/1 def flash(assigns) do
# 11-11 component .icon <.icon name="hero-x-mark" />
gmem code find flash
# lib/demo_web/components/core_components.ex:8-15 function flash/1 in DemoWeb.CoreComponents def flash(assigns) do
gmem mcp
Run the MCP server over stdio (newline-delimited JSON-RPC 2.0). See
docs/mcp.md
for the tool contract, and the gmem-development
skill’s scripts/mcp_smoke.py for a dependency-free way to smoke-test it by
hand.