Measurement report · Amendo

Does an Obsidian vault actually save the agent tokens?

A measured comparison of text search versus link-graph traversal as the retrieval path for Claude Code working on the Amendo Symfony codebase. Every number below was taken from the live vault and the live CLI — none are estimates unless labelled as such.

Codebase
Symfony 3.4 / PHP 7.4
Scale
1,115 files · 265,263 lines
Vault
C:\Amendo\docs · 246 notes
CLI
Obsidian.com · cli=true
Date
2026-07-27

Verdict

Yes — but not for the reason most people assume

The vault does not make reading a document cheaper. A note costs the same number of tokens whether ripgrep printed it or the Obsidian CLI printed it.

What it makes dramatically cheaper is answering relationship questions — what touches this?, what breaks if I change it? — because a [[wikilink]] is a human judgement that text matching cannot reconstruct. On that class of question the measured advantage is roughly an order of magnitude.

Relationship query 4.7× Fewer tokens to answer "what relates to Product?" via backlinks than via grep.
Precision 22% → 100% Grep returned 125 notes; 28 were genuinely related. Backlinks returned exactly the 28.
Structural peek 48× Cost of outline versus reading the whole note to find out what is in it.

Baseline · how retrieval works with no vault

There is no index. There is only reading.

Claude Code has no embeddings, no vector store, and no pre-scan of the repository. Nothing about the codebase is in context until it is read. Three primitives do all the work: Glob for filename patterns, Grep for content, and Read for targeted file access.

A typical investigation — "the Vipps refund isn't writing a transaction row" — runs grep, infers structure from Symfony's conventions, reads two or three candidates, follows constructor injection to the next hop, and checks the service wiring. Four to eight tool calls, roughly 25,000–40,000 tokens of context for a medium task.

This loop is competent at finding files. It is structurally incapable of four things:

  • Vocabulary gaps. A Norwegian business term whose class is named something else. Grep returns nothing and the search stalls.
  • Cross-cutting flows. POS device → ws host → tenant database → back-office screen spans four bundles. Grep sees files, never a flow.
  • Live versus dead. Around twenty integrations exist in the tree. Nothing in the code says which three are in production.
  • Relationships. Which features break if an entity changes. This is the expensive one, and it is measured below.

Key constraint

There is exactly one input channel: text in the context window. Every fact the agent learns arrives as tokens it reads. No tool — CLI, MCP server, or REST API — can deliver knowledge through a side channel. This is why swapping the retrieval mechanism changes nothing, and only changing what gets returned changes the bill.

Finding · a common assumption, tested

“The CLI reads a JSON graph” — false, and it matters why

A reasonable hypothesis is that .obsidian/graph.json holds the link graph, and that reading it would be cheaper than searching. The file exists. Here it is in full, all 607 bytes of it:

{
  // C:\Amendo\docs\.obsidian\graph.json
  "showTags": true,          "showOrphans": true,
  "textFadeMultiplier": 0,   "nodeSizeMultiplier": 1,
  "centerStrength": 0.5,     "repelStrength": 10,
  "linkStrength": 1,         "linkDistance": 250,
  "scale": 0.18462326225858108,
  "colorGroups": [{ "query": "surface:SA", "color": { "rgb": 4886754 } }]
}

Those are graph-view rendering settings — force-simulation physics, zoom level, node colours. There is no node list, no edge list, no adjacency. It configures how the graph is drawn, not what the graph is.

But the underlying instinct was right: a structured index does exist. It simply lives in Obsidian's in-memory metadata cache, not on disk as JSON, and the CLI is the way to reach it. backlinks, links, outline, properties, and aliases all query that cache. That is where the real advantage turns out to be.

Measurement 1 · relationship query

“What relates to Product?”

The single most common question when planning a change. Run both ways against the live 246-note vault.

Tokens returned · lower is better · single measure, two methods
rg "Product" docs/ -l 680 tok · 125 notes
backlinks file=Product 145 tok · 28 notes

Bars scaled to the larger value. Measured 2026-07-27 on C:\Amendo\docs.

Measurement 1 — full data
Method Notes returned Output chars Tokens (approx.) Precision
Grep, lexical match on “Product” 125 of 246 2,709 ~680 ~22%
CLI, curated backlink graph 28 572 ~145 100%
Advantage 4.5× tighter — 4.7× cheaper ~21× signal/token

Why the gap is structural, not incidental

Grep matched half the vault. “Product” appears inside ProductUnit, ProductOrders, ProductTransfer, and in any note that merely uses the word — this is lexical coincidence. The 28 backlinks are notes where a human deliberately wrote [[Product]]. That is curated semantic structure, and no amount of text matching recovers it. The gap does not close with a better regex.

Measurement 2 · structural peek

Deciding whether a note is worth reading

Traversal surfaces candidates; it does not absorb them. The cost of triage is what makes or breaks the loop. outline returns a note's heading tree without its body.

Tokens to evaluate API/Product.md · 7,166 bytes, 60 lines
read file=Product
full note body
1,900 tok
outline file=Product
heading tree only
40 tok

Bars scaled to the larger value. 155 characters of output versus 7,166 bytes of file.

$ obsidian outline file=Product format=tree

└── ProductController (BO)
    ├── Responsibility
    ├── Routes / Actions
    ├── Key dependencies
    ├── Related features
    └── Notes / gotchas

Forty tokens establish that this note documents a controller and carries a Notes / gotchas section — enough to decide whether to spend the other 1,860. Grep has no equivalent: it can show matching lines, but it cannot show a document's shape. This is genuine progressive disclosure at the note level, and it is what makes traversing 28 backlinks affordable instead of ruinous.

Worked example · where the savings actually come from

“Add a discount field to the product import”

The same task costed three ways. Scenario B and C are deliberately paired to isolate one variable: whether the tool matters, or whether the documentation matters.

Estimated token cost by scenario
Step A · No docs B · Docs + Grep C · Docs + CLI
Locate the feature2,9008080
Read the feature note—2,0002,000
Read ProductImportService.php8,0001,3001,300
Read ProductImportType.php2,700——
Trace service wiring300——
Read Product entity5,4005,4005,400
Total19,3008,8008,800

The isolation result

A → B saves 10,500 tokens. That saving comes entirely from the note existing and naming ProductImportService.php:45, which converts an 8,000-token full read into a 1,300-token targeted one. B and C are identical. The CLI changed nothing here, because this task is a straight document retrieval. Write the docs and the win is already banked — the tooling is a separate question, answered by Measurements 1 and 2.

Protocol · the loop the measurements imply

“What breaks if I change the Product entity?”

This is the question grep cannot answer at any price — it would have to read notes to discover relationships that the links already encode. With traversal plus cheap triage, it costs about seven thousand tokens and returns full recall across forty related notes.

  1. Resolve the entry pointRead 00-Index.md, or search query=… limit=5. Traversal needs a starting node. ~500 tok
  2. Expand the neighbourhoodlinks file=Product + backlinks file=Product → 40 curated candidates. ~200 tok
  3. Triage structurallyoutline each candidate at ~40 tokens. Read shape, not substance. ~1,600 tok
  4. Read selectivelyOnly the two or three notes that actually matter, then verify against source. ~5,000 tok

Total ≈ 7,300 tokens, with recall across 40 notes. The critical property is step 3: without cheap triage, step 2's precision becomes a liability — 28 backlinks cost 145 tokens to list and roughly 53,000 to read. Precision surfaces more real work, and only outline makes that affordable.

Design rule · where each fact belongs

One question decides it

Does this fact change when the code changes? This is not a stylistic preference. A vault cannot be version-controlled against a branch — put a code fact in it and it desyncs the moment someone refactors, and the agent will act confidently on the stale version.

Placement by volatility
Durable — vault / second brain Code-coupled — repo docs/ & CLAUDE.md
Why the master/dynamic DB split exists, and what was rejectedWhich entity manager each bundle maps to
Which of the ~20 integrations clients actually pay forIntegration entry-point paths
Norwegian domain glossary (business term → class name)Entity and field mappings
“Tenant 47 has a hand-patched schema — do not run migrations”Migration command conventions
Debugging war stories, ticket context, client quirksBackOfficeBundle subdirectory map
Meeting decisions and architectural rationaleRoute tables, service wiring

Note the asymmetry. The right column is largely what grep would have found anyway in ~30,000 tokens. The left column can never be recovered from source at any token price. That is where the leverage is — and it is exactly what a vault is good at.

Precedence when sources disagree

Source code > repo docs > vault. Put this line in CLAUDE.md so it applies without being asked. Without an explicit precedence rule, two knowledge stores are strictly worse than one.

Risk · constraints before adopting this

Two things that will bite

The CLI requires Obsidian to be running

It is a shim that talks to the live application — which is why it has an “active file” concept and a vault= selector. Headless, CI, remote, or app-closed means it is unavailable. Grep never is. The protocol must therefore always define a grep fallback, never assume the CLI.

The same binary can destroy the vault

delete … permanent skips the trash entirely. Alongside it sit move, rename, property:set, plugin:install, plugins:restrict off and restart. One malformed argument mutates 246 real notes with no undo.

Required mitigation — allowlist, do not blocklist

Grant only read verbs. A blocklist fails open the moment Obsidian ships a new command; an allowlist fails closed. The full protocol is specified in Claude-Vault-CLI-Protocol.md.

readoutlinesearchbacklinks linkspropertiesproperty:readaliases filesfoldersorphansdeadends base:query
✕ delete✕ move✕ rename ✕ create✕ append✕ prepend ✕ property:set✕ property:remove ✕ plugin:*✕ plugins:restrict ✕ command✕ restart✕ history:restore

Current state · the Amendo vault today

The capability is unlocked by authoring, not by installing

The vault at C:\Amendo\docs is already well built — 246 notes, densely linked, organised by surface, with a written explanation of how the agent should use it. The measurements above are only possible because that link structure exists.

Three gaps are worth closing, in this order:

Observed gaps
Observation Measured Consequence
aliases total 0 Norwegian business terms cannot resolve to notes. This is precisely the vocabulary gap grep cannot bridge — the highest-value unclaimed win.
bases none No typed structured queries. base:query … format=paths is the one capability with no grep equivalent at all.
00-How-Claude-Uses-This-Vault.md outdated States that Claude “never launches Obsidian” and that wikilinks are “not a traversed graph.” True when written; cli=true is now set and the CLI is verified working.

Recommended sequence

1. Add aliases: frontmatter carrying the Norwegian terms — this alone converts a dead-end search into a direct hit. 2. Add status:, bundle: and covers: properties, then define one Base over them. 3. Merge the CLI protocol into the vault guide and correct the superseded claims. 4. For the v2 repo, write src/BackOfficeBundle/CLAUDE.md — 61% of the codebase, currently unmapped.

Skip step 1 and 2 and the CLI degrades to a slower grep that needs the app open. The conclusion of this report is not install the tool. It is write the notes so the tool has something to traverse — the graph is the asset, and the CLI is only how it is read.