MCP server · tree-sitter · SQLite

Cartograph

Turn any repository into a queryable code graph and serve it to coding agents over MCP — so an agent can ask “what breaks if I change this?” instead of grepping and hoping.

View on GitHub no embeddings no API keys $0 to run

Generated from a live index of this project · v0.4.0 · 2026-09-09

The problem

Give a coding agent a large unfamiliar repo and watch it work: grep, read a file, grep again. It burns context rebuilding structure a parser could have supplied in one call — and still misses the caller three modules away that its change just broke.

The reflex is RAG: embed the code, retrieve “similar” chunks. But “who calls this function?” is not a similarity question. It has an exact answer, and that answer lives in the call graph.

Live numbers from this repo

43
files indexed
8,775
lines parsed
594
symbols
1,731
graph edges
64%
internal resolution
3
languages
229
tests
0
API keys needed

Impact analysis, before the edit

The query that earns the whole project. One call names the dependent files and the tests to run.

$ cartograph blast src/cartograph/graph/store.py

Blast radius — file src/cartograph/graph/store.py

17 dependent file(s), 20 affected symbol(s), 7 test file(s).

Tests to run first

  • tests/test_cli.py
  • tests/test_docs.py
  • tests/test_incremental.py
  • tests/test_mcp.py
  • tests/test_resolver.py
  • tests/test_traversal.py
  • tests/test_views.py

Dependent files (by import distance)

  • tests/test_cli.py · d1 · test
  • tests/test_incremental.py · d1 · test
  • tests/test_resolver.py · d1 · test
  • src/cartograph/graph/resolver.py · d1
  • src/cartograph/indexer/pipeline.py · d1
  • src/cartograph/service.py · d1
  • tests/test_mcp.py · d2 · test
  • tests/test_traversal.py · d2 · test
  • tests/test_views.py · d2 · test
  • scripts/bench.py · d2
  • scripts/build_docs.py · d2
  • scripts/smoke_mcp.py · d2
  • src/cartograph/cli.py · d2
  • src/cartograph/mcp_server/server.py · d2
  • src/cartograph/views.py · d2
  • tests/conftest.py · d2
  • tests/test_docs.py · d3 · test

Affected symbols

  • cartograph.graph.resolver:resolve_imports · src/cartograph/graph/resolver.py:107 · 0.97 (inferred-type)
  • cartograph.graph.resolver:resolve_references · src/cartograph/graph/resolver.py:135 · 0.97 (inferred-type)
  • cartograph.service:Cartograph._one · src/cartograph/service.py:118 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.candidates · src/cartograph/service.py:124 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.close · src/cartograph/service.py:94 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.find_symbol · src/cartograph/service.py:105 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.related · src/cartograph/service.py:181 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.search · src/cartograph/service.py:115 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.stats · src/cartograph/service.py:293 · 0.97 (inferred-type)
  • cartograph.service:Cartograph.symbol_detail · src/cartograph/service.py:127 · 0.97 (inferred-type)
  • cartograph.indexer.pipeline:index_repo · src/cartograph/indexer/pipeline.py:42 · 0.45 (name-only)
  • cartograph.service:Cartograph.load · src/cartograph/service.py:84 · 0.45 (name-only)
  • tests.test_cli:test_opening_an_index_writes_nothing · tests/test_cli.py:201 · 0.45 (name-only)
  • tests.test_cli:test_reading_while_indexing_does_not_lock · tests/test_cli.py:185 · 0.45 (name-only)
  • tests.test_cli:test_schema_is_not_recreated_on_reopen · tests/test_cli.py:220 · 0.45 (name-only)
  • tests.test_incremental:test_fts_rows_do_not_leak_after_reindex · tests/test_incremental.py:189 · 0.45 (name-only)
  • tests.test_resolver:test_deleting_a_target_removes_its_edges · tests/test_resolver.py:334 · 0.45 (name-only)
  • tests.test_traversal:test_find_symbol_exact_flag_excludes_partial_matches · tests/test_traversal.py:42 · 0.45 (name-only)
  • cartograph.cli:search_cmd · src/cartograph/cli.py:175 · 0.40 (ambiguous)
  • cartograph.mcp_server.server:search_code · src/cartograph/mcp_server/server.py:131 · 0.40 (ambiguous)

Recall-first view (min_confidence 0.3): treat as a review checklist, not a proof of impact.

Architecture, extracted not documented

Module graph, layering, cycles and hotspots — computed from imports and the call graph. Edge labels are import counts; PageRank picks out the symbols that carry the codebase.

flowchart LR
  m0["cartograph.indexer.extract"] -- 5 --> m1["cartograph.models"]
  m2["cartograph.mcp_server.server"] -- 5 --> m3["cartograph.service"]
  m4["cartograph.cli"] -- 4 --> m3["cartograph.service"]
  m5["cartograph.graph.store"] -- 4 --> m1["cartograph.models"]
  m6["cartograph.views"] -- 4 --> m3["cartograph.service"]
  m7["cartograph.graph.resolver"] -- 3 --> m5["cartograph.graph.store"]
  m3["cartograph.service"] -- 3 --> m8["cartograph.graph.algorithms"]
  m3["cartograph.service"] -- 3 --> m1["cartograph.models"]
  m6["cartograph.views"] -- 3 --> m1["cartograph.models"]
  m9["scripts.build_docs"] -- 3 --> m3["cartograph.service"]
  m4["cartograph.cli"] -- 2 --> m10["cartograph.indexer.pipeline"]
  m7["cartograph.graph.resolver"] -- 2 --> m11["cartograph.indexer.languages"]
  m12["cartograph.indexer.bindings"] -- 2 --> m11["cartograph.indexer.languages"]
  m0["cartograph.indexer.extract"] -- 2 --> m11["cartograph.indexer.languages"]
  m10["cartograph.indexer.pipeline"] -- 2 --> m5["cartograph.graph.store"]
  m10["cartograph.indexer.pipeline"] -- 2 --> m1["cartograph.models"]
  m3["cartograph.service"] -- 2 --> m5["cartograph.graph.store"]
  m4["cartograph.cli"] --> m13["cartograph"]
  %% showing 18 of 40 module edges by weight

Hotspots — highest PageRank

SymbolLocationFan-in Rank
cartograph.indexer.pipeline:index_reposrc/cartograph/indexer/pipeline.py:42470.0266
cartograph.models:Symbolsrc/cartograph/models.py:9520.0236
cartograph.graph.store:_to_symbolsrc/cartograph/graph/store.py:72190.0231
cartograph.views:TokenBudget.addsrc/cartograph/views.py:41940.0213
tests.test_extract:parsetests/test_extract.py:14210.0142
cartograph.indexer.languages:adapter_for_pathsrc/cartograph/indexer/languages.py:46160.0127
tests.test_resolver:edges_oftests/test_resolver.py:21380.0117
cartograph.indexer.extract:extract_filesrc/cartograph/indexer/extract.py:12340.0117
tests.test_mcp:calltests/test_mcp.py:43150.0113
tests.test_resolver:buildtests/test_resolver.py:36370.0109

Confidence is a first-class column

Without a type checker you cannot know that store.who_calls() means GraphStore.who_calls. You can only rank hypotheses. So every edge records the rule that produced it and a confidence, and callers choose their own operating point: who_calls is precision-first (≥0.5), blast_radius is recall-first (≥0.3).

RuleConfidenceIntuition Edges here
same-file0.95the definition is right there in scope411
import0.90the file explicitly imported this name185
receiver-type0.96Foo.bar() where Foo is a known container35
same-module0.75sibling file in the same package1
unique-global0.60one repo symbol has this name, bare call site6
name-only0.45one match, but on an untyped receiver69
ambiguous0.40*N candidates, kept as N edges at 1/N each59
external0.40*rooted at a third-party or stdlib import355
unresolved0.00genuinely unknown (dynamic, or a typed method)389

The name-only tier exists because of a real bug: seen.add(...) on a builtin set was resolving to a repo class’s add method purely because the name was unique — and surfacing as a confident caller. A method name on a receiver you cannot type is not evidence, so it now sits below the precision line.

Local type inference, and where it refuses

The largest class of unresolvable call is receiver.method(): you cannot know that self.conn.execute(...) means sqlite3.Connection.execute without knowing what conn is. Cartograph works it out from three shallow sources — annotations (def f(store: GraphStore), s *Store), construction (GraphStore(...), new Engine(), &Engine{}) and aliases (self.store = store, chased transitively) — across Python, TypeScript, JavaScript and Go. Attribute chains resolve one hop, so s.conn.Exec() is a Conn.Exec when s is a Store.

What it deliberately will not do is return-type propagation, generics or cross-module dataflow. A wrong inferred type produces a confident edge pointing at the wrong function, and an agent acts on it — strictly worse than an honest unresolved. So a variable reassigned to a second type keeps its first binding, and a real union binds nothing.

Inheritance is answered from the inherits edges the graph already stored: once a rule names the receiver's owner and finds no member on it directly, its declared bases are walked breadth-first. That is worth +6.9 points on django and +5.8 on flask — the mirror image of inference, which pays on annotated code and does almost nothing for class-heavy dynamic Python.

Once a pinned owner's whole ancestry has been searched, where that ancestry ends is itself the answer. Entirely inside the repo, and the member is genuinely not ours. At a third-party base, and the member is unittest’s — django’s self.assertEqual reaches unittest.TestCase through several in-repo base classes, and 15,360 call sites move from “no idea” to a correct, specific answer. At a name we cannot account for, absence proves nothing and the cascade goes on.

Locality is not evidence about ownership either. Same-package name matches used to be broken by a deterministic sort, which produced confident but arbitrary edges — on gin, w.Header().Get resolved to Params.Get. Measured against the edges inference now types at 0.97: where a method name had exactly one owner the old rules were right 100% of the time (n=658), and where it had more than one, 66% (n=296). Multi-owner ties are now reported as ambiguous.

The tool surface

Ten MCP tools, two resources and a prompt. Descriptions are written as routing guidance, because the description is the prompt the model uses to choose.

find_symbol

Where is X defined? Ranked by structural importance.

search_code

Full-text over names, signatures and docstrings (BM25).

get_symbol

One symbol: signature, doc, members, callers, callees, source.

who_calls

Reverse call tree — before you change a signature.

what_it_calls

Forward call tree — understand code without reading every file.

blast_radius

What a change could break, and which tests to run.

related_symbols

“What else should I read?” via personalized PageRank.

file_summary

What a file defines, imports, and who imports it.

architecture_overview

Modules, layering, import cycles, hotspots, entry points.

index_stats

Index health and the edge-resolution breakdown by rule.

Benchmarks

Real repositories, single process, M-series laptop. Cold = full index; warm = no-op reindex.

RepoFilesKLOC SymbolsEdgesCold WarmInternal resolution
django2,97753445,508241,27812.5s1.82s42.7%
gin (Go)98241,61010,0060.45s0.04s74.0%
flask83181,6244,1750.25s0.04s37.3%

Warm reindex is fast because resolution is skipped only when both its inputs are provably unchanged. Resolution is counted per call site, and only a above-threshold edge counts — an ambiguous site emits one edge per candidate, so counting edges scored six guesses as six successes. Reproduce with scripts/bench.py --clone.

Try it

uv tool install cartograph-mcp

cartograph index ~/code/my-repo
cartograph arch
cartograph blast src/auth/token.py

# wire it into an agent
claude mcp add cartograph -- cartograph serve ~/code/my-repo

Limitations

Stated plainly, because a code-intelligence tool that oversells its precision is worse than useless.