The problem
Give a coding agent a large unfamiliar repo and watch it work: grep, read a
file, grep again. It burns context rebuilding structure a parser could have
supplied in one call — and still misses the caller three modules away that its change just
broke.
The reflex is RAG: embed the code, retrieve “similar” chunks. But “who calls this function?” is not a similarity question. It has an exact answer, and that answer lives in the call graph.
Live numbers from this repo
Impact analysis, before the edit
The query that earns the whole project. One call names the dependent files and the tests to run.
$ cartograph blast src/cartograph/graph/store.py
Blast radius — file src/cartograph/graph/store.py
17 dependent file(s), 20 affected symbol(s), 7 test file(s).
Tests to run first
tests/test_cli.pytests/test_docs.pytests/test_incremental.pytests/test_mcp.pytests/test_resolver.pytests/test_traversal.pytests/test_views.py
Dependent files (by import distance)
tests/test_cli.py· d1 · testtests/test_incremental.py· d1 · testtests/test_resolver.py· d1 · testsrc/cartograph/graph/resolver.py· d1src/cartograph/indexer/pipeline.py· d1src/cartograph/service.py· d1tests/test_mcp.py· d2 · testtests/test_traversal.py· d2 · testtests/test_views.py· d2 · testscripts/bench.py· d2scripts/build_docs.py· d2scripts/smoke_mcp.py· d2src/cartograph/cli.py· d2src/cartograph/mcp_server/server.py· d2src/cartograph/views.py· d2tests/conftest.py· d2tests/test_docs.py· d3 · test
Affected symbols
cartograph.graph.resolver:resolve_imports· src/cartograph/graph/resolver.py:107 · 0.97 (inferred-type)cartograph.graph.resolver:resolve_references· src/cartograph/graph/resolver.py:135 · 0.97 (inferred-type)cartograph.service:Cartograph._one· src/cartograph/service.py:118 · 0.97 (inferred-type)cartograph.service:Cartograph.candidates· src/cartograph/service.py:124 · 0.97 (inferred-type)cartograph.service:Cartograph.close· src/cartograph/service.py:94 · 0.97 (inferred-type)cartograph.service:Cartograph.find_symbol· src/cartograph/service.py:105 · 0.97 (inferred-type)cartograph.service:Cartograph.related· src/cartograph/service.py:181 · 0.97 (inferred-type)cartograph.service:Cartograph.search· src/cartograph/service.py:115 · 0.97 (inferred-type)cartograph.service:Cartograph.stats· src/cartograph/service.py:293 · 0.97 (inferred-type)cartograph.service:Cartograph.symbol_detail· src/cartograph/service.py:127 · 0.97 (inferred-type)cartograph.indexer.pipeline:index_repo· src/cartograph/indexer/pipeline.py:42 · 0.45 (name-only)cartograph.service:Cartograph.load· src/cartograph/service.py:84 · 0.45 (name-only)tests.test_cli:test_opening_an_index_writes_nothing· tests/test_cli.py:201 · 0.45 (name-only)tests.test_cli:test_reading_while_indexing_does_not_lock· tests/test_cli.py:185 · 0.45 (name-only)tests.test_cli:test_schema_is_not_recreated_on_reopen· tests/test_cli.py:220 · 0.45 (name-only)tests.test_incremental:test_fts_rows_do_not_leak_after_reindex· tests/test_incremental.py:189 · 0.45 (name-only)tests.test_resolver:test_deleting_a_target_removes_its_edges· tests/test_resolver.py:334 · 0.45 (name-only)tests.test_traversal:test_find_symbol_exact_flag_excludes_partial_matches· tests/test_traversal.py:42 · 0.45 (name-only)cartograph.cli:search_cmd· src/cartograph/cli.py:175 · 0.40 (ambiguous)cartograph.mcp_server.server:search_code· src/cartograph/mcp_server/server.py:131 · 0.40 (ambiguous)
Recall-first view (min_confidence 0.3): treat as a review checklist, not a proof of impact.
Architecture, extracted not documented
Module graph, layering, cycles and hotspots — computed from imports and the call graph. Edge labels are import counts; PageRank picks out the symbols that carry the codebase.
flowchart LR m0["cartograph.indexer.extract"] -- 5 --> m1["cartograph.models"] m2["cartograph.mcp_server.server"] -- 5 --> m3["cartograph.service"] m4["cartograph.cli"] -- 4 --> m3["cartograph.service"] m5["cartograph.graph.store"] -- 4 --> m1["cartograph.models"] m6["cartograph.views"] -- 4 --> m3["cartograph.service"] m7["cartograph.graph.resolver"] -- 3 --> m5["cartograph.graph.store"] m3["cartograph.service"] -- 3 --> m8["cartograph.graph.algorithms"] m3["cartograph.service"] -- 3 --> m1["cartograph.models"] m6["cartograph.views"] -- 3 --> m1["cartograph.models"] m9["scripts.build_docs"] -- 3 --> m3["cartograph.service"] m4["cartograph.cli"] -- 2 --> m10["cartograph.indexer.pipeline"] m7["cartograph.graph.resolver"] -- 2 --> m11["cartograph.indexer.languages"] m12["cartograph.indexer.bindings"] -- 2 --> m11["cartograph.indexer.languages"] m0["cartograph.indexer.extract"] -- 2 --> m11["cartograph.indexer.languages"] m10["cartograph.indexer.pipeline"] -- 2 --> m5["cartograph.graph.store"] m10["cartograph.indexer.pipeline"] -- 2 --> m1["cartograph.models"] m3["cartograph.service"] -- 2 --> m5["cartograph.graph.store"] m4["cartograph.cli"] --> m13["cartograph"] %% showing 18 of 40 module edges by weight
Hotspots — highest PageRank
| Symbol | Location | Fan-in | Rank |
|---|---|---|---|
cartograph.indexer.pipeline:index_repo | src/cartograph/indexer/pipeline.py:42 | 47 | 0.0266 |
cartograph.models:Symbol | src/cartograph/models.py:95 | 2 | 0.0236 |
cartograph.graph.store:_to_symbol | src/cartograph/graph/store.py:721 | 9 | 0.0231 |
cartograph.views:TokenBudget.add | src/cartograph/views.py:41 | 94 | 0.0213 |
tests.test_extract:parse | tests/test_extract.py:14 | 21 | 0.0142 |
cartograph.indexer.languages:adapter_for_path | src/cartograph/indexer/languages.py:461 | 6 | 0.0127 |
tests.test_resolver:edges_of | tests/test_resolver.py:21 | 38 | 0.0117 |
cartograph.indexer.extract:extract_file | src/cartograph/indexer/extract.py:123 | 4 | 0.0117 |
tests.test_mcp:call | tests/test_mcp.py:43 | 15 | 0.0113 |
tests.test_resolver:build | tests/test_resolver.py:36 | 37 | 0.0109 |
Confidence is a first-class column
Without a type checker you cannot know that store.who_calls() means
GraphStore.who_calls. You can only rank hypotheses. So every edge records the rule
that produced it and a confidence, and callers choose their own operating point:
who_calls is precision-first (≥0.5), blast_radius is
recall-first (≥0.3).
| Rule | Confidence | Intuition | Edges here |
|---|---|---|---|
same-file | 0.95 | the definition is right there in scope | 411 |
import | 0.90 | the file explicitly imported this name | 185 |
receiver-type | 0.96 | Foo.bar() where Foo is a known container | 35 |
same-module | 0.75 | sibling file in the same package | 1 |
unique-global | 0.60 | one repo symbol has this name, bare call site | 6 |
name-only | 0.45 | one match, but on an untyped receiver | 69 |
ambiguous | 0.40* | N candidates, kept as N edges at 1/N each | 59 |
external | 0.40* | rooted at a third-party or stdlib import | 355 |
unresolved | 0.00 | genuinely unknown (dynamic, or a typed method) | 389 |
The name-only tier exists because of a real bug:
seen.add(...) on a builtin set was resolving to a repo class’s
add method purely because the name was unique — and surfacing as a confident
caller. A method name on a receiver you cannot type is not evidence, so it now sits below the
precision line.
Local type inference, and where it refuses
The largest class of unresolvable call is receiver.method(): you cannot know
that self.conn.execute(...) means sqlite3.Connection.execute without
knowing what conn is. Cartograph works it out from three shallow sources —
annotations (def f(store: GraphStore), s *Store), construction
(GraphStore(...), new Engine(), &Engine{}) and
aliases (self.store = store, chased transitively) — across Python, TypeScript,
JavaScript and Go. Attribute chains resolve one hop, so s.conn.Exec() is a
Conn.Exec when s is a Store.
What it deliberately will not do is return-type propagation, generics or cross-module
dataflow. A wrong inferred type produces a confident edge pointing at the wrong function, and
an agent acts on it — strictly worse than an honest unresolved. So a variable
reassigned to a second type keeps its first binding, and a real union binds nothing.
Inheritance is answered from the inherits edges the graph already stored:
once a rule names the receiver's owner and finds no member on it directly, its declared bases
are walked breadth-first. That is worth +6.9 points on django and
+5.8 on flask — the mirror image of inference, which pays on annotated code
and does almost nothing for class-heavy dynamic Python.
Once a pinned owner's whole ancestry has been searched, where that ancestry ends
is itself the answer. Entirely inside the repo, and the member is genuinely not ours. At a
third-party base, and the member is unittest’s — django’s
self.assertEqual reaches unittest.TestCase through several in-repo
base classes, and 15,360 call sites move from “no idea” to a correct, specific
answer. At a name we cannot account for, absence proves nothing and the cascade goes on.
Locality is not evidence about ownership either. Same-package name matches
used to be broken by a deterministic sort, which produced confident but arbitrary edges — on
gin, w.Header().Get resolved to Params.Get. Measured against the
edges inference now types at 0.97: where a method name had exactly one owner the old rules
were right 100% of the time (n=658), and where it had more than one,
66% (n=296). Multi-owner ties are now reported as ambiguous.
The tool surface
Ten MCP tools, two resources and a prompt. Descriptions are written as routing guidance, because the description is the prompt the model uses to choose.
find_symbol
Where is X defined? Ranked by structural importance.
search_code
Full-text over names, signatures and docstrings (BM25).
get_symbol
One symbol: signature, doc, members, callers, callees, source.
who_calls
Reverse call tree — before you change a signature.
what_it_calls
Forward call tree — understand code without reading every file.
blast_radius
What a change could break, and which tests to run.
related_symbols
“What else should I read?” via personalized PageRank.
file_summary
What a file defines, imports, and who imports it.
architecture_overview
Modules, layering, import cycles, hotspots, entry points.
index_stats
Index health and the edge-resolution breakdown by rule.
Benchmarks
Real repositories, single process, M-series laptop. Cold = full index; warm = no-op reindex.
| Repo | Files | KLOC | Symbols | Edges | Cold | Warm | Internal resolution |
|---|---|---|---|---|---|---|---|
| django | 2,977 | 534 | 45,508 | 241,278 | 12.5s | 1.82s | 42.7% |
| gin (Go) | 98 | 24 | 1,610 | 10,006 | 0.45s | 0.04s | 74.0% |
| flask | 83 | 18 | 1,624 | 4,175 | 0.25s | 0.04s | 37.3% |
Warm reindex is fast because resolution is skipped only when both its inputs
are provably unchanged. Resolution is counted per call site, and only a
above-threshold edge counts — an ambiguous site emits one edge per candidate, so
counting edges scored six guesses as six successes. Reproduce with
scripts/bench.py --clone.
Try it
uv tool install cartograph-mcp
cartograph index ~/code/my-repo
cartograph arch
cartograph blast src/auth/token.py
# wire it into an agent
claude mcp add cartograph -- cartograph serve ~/code/my-repo
Limitations
Stated plainly, because a code-intelligence tool that oversells its precision is worse than useless.
- Type inference is local only. Annotations, constructors and aliases,
one attribute hop, within a file. A receiver whose type is never written down —
app.config.get(...)— stays unresolved, which is most of what remains on untyped Python. - Inheritance is walked, but not linearised. Bases are searched breadth-first, which matches C3 for single inheritance but can pick a different branch of a diamond than the interpreter would.
- Interface dispatch is not modelled. A call through a Go interface or an abstract base resolves to no single implementation.
- Dynamic dispatch is invisible.
getattr(obj, name)(), decorator registries and DI containers produce no edges. - No cross-language edges. A TypeScript frontend calling a Python endpoint is two disconnected subgraphs.