Graphmemdocs GitHub ↗

Code-tool agent benchmark: gmem vs a plain agent

Does a coding agent reach the same answer using less context when gmem’s code tools are available? The harness is in eval/ . Raw attempts and generated summaries are in eval/results/; the generated cross-model comparison reports median tokens, mean estimated cost, accuracy, and resource settings.

Current results

All three matrices contain 75 fresh attempts: five tasks × three variants × five runs. They use the corrected version-2 harness, temperature 0, eight jobs, and a 16,384-token generation ceiling.

modelcorrectestimated matrix costadditional resource caps
DeepSeek V4.1 Flash75/75$0.8823none; baseline predates caps
Nemotron Lightning 3.5 30B A3B53/75$0.4462none; baseline predates caps
GLM 5.3 Flash70/75$0.5124shared cost budget and per-task token caps

DeepSeek and Nemotron have no invocation errors. GLM has four deliberate per-run token-budget stops and one incorrect symbol answer; no other invocation errors. All retained attempts have known provider usage. Failed and stopped attempts remain in accuracy, token, and cost aggregates.

GLM’s preliminary provider check cost $0.00099149 and is not included in its 75-attempt matrix. It was deducted from the agreed $1.15 allowance, leaving $1.14900851 for the matrix. Combined estimated GLM spend was $0.51336; the cost ceiling did not bind.

Do not treat these as controlled model-quality rankings: GLM ran with additional token budgets while the two baselines did not. Compare variants within each matrix, and retain the budget-stop counts alongside accuracy. No failed samples were selectively replaced.

Conclusions

  1. Savings are task- and model-dependent. Neutral GLM reduces diff mean cost and tokens by about 55%, with 5/5 correct in both cells. Neutral DeepSeek instead uses 5% more mean diff tokens and costs 1% more.
  2. Cost is not total tokens. Neutral DeepSeek costs 48% less on symbol lookup and 20% less on outline, despite using 30% and 46% more mean tokens. Cache hits and model-specific cached-input rates matter.
  3. Less shell/file reading is not always less context. DeepSeek’s gmem variants read fewer shell/file lines on four of five tasks, but both read more on workflow. MCP text, tool definitions, and guidance add context.
  4. Guidance is not a universal savings switch. Guided GLM imports finishes only 2/5: three runs hit the token allowance. Guided Nemotron imports reaches 5/5 at 79% lower mean cost than its control.
  5. Cheaper incorrect answers are not a win. Nemotron’s diff task has only one fully correct answer across all 15 attempts, despite lower gmem costs.
  6. Protect the budget before requests, not after a whole agent finishes. Shared reservations include all eight workers and retries. Historical medians, rather than runaway maxima, define new per-task allowances.

Method

  • Corpus: openclaw/openclaw@8f5c33c3; diff base 4de57f22.
  • Control: shell and read_file, retaining rg/sed/git. It is not handicapped; gmem tools are additive.
  • Neutral gmem: control tools plus find_symbol, code_outline, code_imports, and code_diff; memory tools are filtered out. Same shared system prompt as control.
  • Guided gmem: neutral gmem plus shipped plugin guidance and an on-demand skill loader. Each attempt starts a fresh agent.
  • Scoring: deterministic against eval/gold.json, not an LLM judge. Only the terminal assistant response is scored. Unfinished tool turns and errored invocations cannot count correct. Repository-relative paths must match fully; basename-only, unrelated, absolute, and traversal paths fail. Import F1 normalizes matching surrounding quotes/backticks.
  • Usage: accumulated provider input/output/cache counts are retained on failure. Unknown complete usage is null/n/a; known_usage retains any measured lower bounds. Cost uses eval/pricing.json list rates, not invoices.
  • Tool delivery: all text, including MCP and skills, shares a 50,000-character per-result cap including the truncation marker. Character metrics are collected at delivery. Shell/file line counts exclude metadata and discarded lines, and do not include MCP source text.
  • Retries: explicit LiteLLM/native throttling backoff, six attempts with waits of 4/8/16/32/64 seconds. Retrying a model call retains the existing agent and paid history; authentication/configuration errors are not retried.

DeepSeek V4.1 Flash

Means per attempt; every cell is 5/5 correct. tool chars measures delivered text; src lines includes only shell/file output, not MCP.

taskvarianttokenscost usdtool charssrc linescycles
diff-001control36,719$0.006019,7574507.6
diff-001gmem38,680$0.006117,3272316.4
diff-001gmem-guided51,989$0.005923,3152056.6
imports-001control12,919$0.000722,5085002.8
imports-001gmem11,790$0.00082,749543.4
imports-001gmem-guided14,863$0.00189,8191543.0
outline-001control9,159$0.00207,5272054.4
outline-001gmem13,377$0.00166,050703.4
outline-001gmem-guided26,246$0.001612,332454.4
symbol-001control9,083$0.00188,4621594.2
symbol-001gmem11,802$0.00091,175204.0
symbol-001gmem-guided15,098$0.001210,455643.0
workflow-001control967,219$0.0477130,7182,38830.6
workflow-001gmem1,030,933$0.0481162,7083,22728.4
workflow-001gmem-guided1,003,091$0.0503156,9892,57627.0

Neutral gmem increases mean total tokens on four tasks and guided gmem on all five. Neutral is cheaper on two tasks; guided on three. Workflow costs rise about 1% and 6%, respectively. This matrix does not establish a general workflow or structural-navigation win.

Nemotron Lightning 3.5 30B A3B

Median tokens / mean USD / correct attempts. Do not confuse these token medians with the DeepSeek means above.

taskcontrolgmemgmem-guided
diff-001219,897 / $0.0193 / 0/5130,536 / $0.0050 / 1/5172,188 / $0.0043 / 0/5
imports-00147,149 / $0.0020 / 5/525,402 / $0.0030 / 4/59,142 / $0.0004 / 5/5
outline-00171,437 / $0.0019 / 4/599,003 / $0.0025 / 4/567,050 / $0.0029 / 5/5
symbol-00149,643 / $0.0011 / 3/56,202 / $0.0001 / 5/57,493 / $0.0004 / 5/5
workflow-001272,629 / $0.0058 / 5/5797,354 / $0.0327 / 4/5453,126 / $0.0077 / 3/5

Both gmem variants improve symbol lookup to 5/5 and reduce its cost. Guided imports is also substantially cheaper with unchanged accuracy. Neither variant improves workflow accuracy or cost. Long loops reached 132 cycles; median tokens limit their influence on the displayed central tendency, but mean costs retain their spend.

GLM 5.3 Flash

Median tokens / mean USD / correct attempts. Every variant shares the same per-task token allowance. Stopped attempts remain included.

taskcontrolgmemgmem-guided
diff-00128,768 / $0.0046 / 5/517,296 / $0.0020 / 5/534,884 / $0.0036 / 5/5
imports-0017,536 / $0.0007 / 5/58,469 / $0.0008 / 5/514,094 / $0.0015 / 2/5
outline-0018,696 / $0.0010 / 5/513,321 / $0.0014 / 5/521,441 / $0.0021 / 5/5
symbol-0017,842 / $0.0009 / 4/59,245 / $0.0006 / 5/513,462 / $0.0011 / 5/5
workflow-001215,887 / $0.0236 / 5/5427,251 / $0.0345 / 4/5348,693 / $0.0240 / 5/5

Neutral diff is 55% cheaper by mean cost, with 40% fewer median tokens; mean tokens fall 55%. Guided diff is 22% cheaper but has 21% more median tokens. Neutral workflow has one budget stop and costs 46% more on average. Three guided import attempts stop before another call would exhaust the 17,099-token allowance. The remaining non-correct attempt is a control symbol answer scoring 0.75. Budget stops do not prove those answers would remain incorrect with an unlimited allowance.

Resource controls and limitations

eval/limits.json fixes new defaults at $1.15 per invocation and the larger baseline model/task median +30% for run tokens. Before each request, a shared ledger reserves conservative uncached-input and maximum-output cost, including in-flight calls and retries. Unknown usage consumes its reservation rather than becoming free. Missing prices or input estimates block calls. Paid usage survives stops.

The per-run token guard uses projected input and clamps output to remaining allowance; inaccurate estimation can overshoot tokens on the last call. Unexpected cost-reservation overruns block further matrix calls. Controls use list rates and conservative local bounds, not an exact provider invoice or account-balance guarantee. Explicit overrides are documented in eval/README.md .

Other limitations:

  • Five samples per cell and a single corpus; no evidence of universal savings or reduced variance. Means and medians can suggest different relative changes.
  • GLM’s resource policy differs from the uncapped baselines. A future controlled cross-model study would need fresh, consistently bounded matrices.
  • Cache state and concurrent request order are uncontrolled.
  • Fewer shell/file lines do not imply less total delivered text or billed context.
  • LiteLLM warns that multi-turn reasoningContent is unsupported; its behavioral effect was not isolated.
  • Records capture model, generation/concurrency/resource settings and assistant turns, but not a full binary/index/prompt-hash manifest.

Future work: larger structural tasks, more corpora, repeated bounded matrices, and a complete reproducibility manifest. Local regression tests require no provider credentials, and no eval GitHub Actions job is added.