flw · research skill

Code Graphs for flw Research

What exists for building a queryable map of a repo, what it costs, and what I'd actually use.

Short answer: yes, and there's proven prior art. One caveat.

A ranked symbol map is a real win. Aider has shipped one for three years and fits a whole repo in a 1,000-token default budget. That's the cheap high-level read you described.

The caveat: a reference graph is precise where it's precise and silently blind where it isn't. I proved that on this repo in three tool calls.

The caveat, measured

I asked your host's LSP for every reference to run_one in core/scripts/gates.py.

LSP findReferences → 1 reference: core/scripts/gates.py:100:5 grep -rn "run_one" → 8 call sites: cli/flw.py:799 + 7 in tests

Zero of eight real callers. The cause: cli/flw.py reaches gates through sys.path.insert then import gates. Pyright can't resolve that, so the edge doesn't exist in the graph.

A second one in the same session: workspaceSymbol("collect") returned four hits, all from plugins/flow/ — a different project. The workspace root was a parent directory.

What this means. Dynamic imports, plugin registries, DI containers, string-keyed dispatch, anything resolved at runtime — those edges are missing and nothing tells you. Big enterprise Python is full of them. A graph that returns zero looks the same as a symbol nobody calls.

It doesn't kill the idea. It sets the rule: a graph is a fast first pass, not an authority. Any "nobody uses this" conclusion needs a grep to confirm.

Two layers, not one

The mistake would be building one thing. A whole-repo symbol graph is hundreds of thousands of tokens — you never want it in context. The high-level skeleton is a page and you always want it.

Layer 1 The skeleton — write it down Packages, entry points, the most-referenced symbols, ranked. ~1k tokens for a whole repo. Durable enough to commit. Gives an agent the nouns.
Layer 2 The graph — query it, don't store it Who calls what, who implements what, where a type is defined. Huge, rots fast, answered on demand.

They depend on each other. findReferences needs a symbol name, so you can't ask "who calls X" until you know X exists. Layer 1 tells you. Layer 2 answers.

The landscape

ToolGives youLanguagesInstallLicenseState
Serena Symbols, references, type hierarchy, symbol-level edits. Real LSPs underneath. 40+ uv tool install serena-agent MIT active 24.8k★, Jun 2026
Host LSP tool Same operations. Already in Claude Code. whatever's configured none — tested works, with the holes above
multilspy Python library wrapping LSP servers. You build the tool. 8 pip install multilspy MIT research code Microsoft Research
tree-sitter + tags Definitions and references per file. No cross-file resolution. 130+ pip install tree-sitter MIT active what aider uses
SCIP Portable index file. Defs + refs. expt-convert dumps to SQLite. ~10 via indexers brew / GitHub releases Apache-2 CLI stale last release Jun 2024
universal-ctags Definitions only, JSON output. No edges. ~150 brew install universal-ctags GPL-2 active
ast-grep Structural search and rewrite. A query tool, not an index. tree-sitter's brew install ast-grep MIT active v0.43
Joern Code property graph — AST + control flow + data flow. Query language. C/C++, Java, JS, Py, Kotlin, binaries JVM Apache-2 heavy security-oriented
CodeQL Relational database of the code, queried in QL. many gh restricted Free for public repos only; private needs a paid GitHub licence

CodeQL is out for many teams. The licence blocks private commercial repos without a per-committer seat.

Build once, query many — the CLI model

An index on disk queried by a CLI is a better fit for flw than an MCP server. MCP is a running process with per-host config. flw is host-agnostic and its whole model is a file plus a command. An index file is that shape.

Three of the four things you'd expect from it hold up. One doesn't.

ExpectationVerdictWhy
Reusable yes It's a file. Commit it, diff it, share it, query it from CI or any tool. A warm server can't be any of those.
Faster across sessions SQLite answers cold in milliseconds. Pyright takes minutes to index a monorepo before it answers anything. Within one long session they converge.
Cheaper in tokens marginally The saving comes from not reading files, and both models give you that. The real CLI edge is that you can pipe — | head -20, | jq — and control result size exactly.
More precise no, a wash Static indexers are the language servers. scip-python is pyright. rust-analyzer --output scip is rust-analyzer. Same analyzer, same blind spots. Tree-sitter indexers are less precise — no type resolution.

The build-then-query options

ToolIndexQueryPrecisionState
GNU GLOBAL GTAGS / GRTAGS global -r sym, global -x Parser-level. Tracks references, which ctags does not. stable since 1996, in every package manager, GPL
CodeGraph SQLite + FTS5 callers, callees, impact, query Tree-sitter level young MIT, 67.6k★, but created Jan 2026
SCIP + indexers protobuf → SQLite via expt-convert Write the SQL yourself Compiler-grade stale CLI last release Jun 2024; converter marked EXPERIMENTAL
ctags + readtags tags file readtags Definitions only, no edges stable ~150 languages
Joern Code property graph, persisted joern --script, CPGQL Deepest open one — AST + control flow + data flow heavy Apache-2, JVM

On CodeGraph. Verified through the GitHub API rather than the page: 67,590 stars, 4,288 forks, created 2026-01-18, last push 2026-08-20, MIT. Real growth — but a seven-month-old project with 439 open issues and 157 watchers. A 430:1 star-to-watcher ratio is a hype signature, not a stability one. Its token-saving claims are self-reported. npm-distributed, which may matter in a locked-down environment.

GNU GLOBAL is the unhyped one that does exactly what you described. gtags builds, global -r symbol queries references. Thirty years of doing one thing.

The infra tier — what Meta actually had

You're remembering Glean. Meta open-sourced it in 2021 and it's still maintained. It's the real thing: typed, schema-defined facts about source code in a queryable database, with Angle, a Datalog-style query language.

An Angle query looks like this:

FunctionDeclaration { name = "parseJson", namespace = "folly" }

It answers questions like "all classes inheriting from exception that have a what method overriding a base" in milliseconds. Facts live in RocksDB. Incremental indexing is O(files changed), not O(repository).

The numbers are genuinely good

Simon Marlow — a Glean author — indexed all of Hackage outside Meta and published the comparison against hiedb:

Gleanhiedb
Index time470s1,021s
Database size0.8 GB5.2 GB
Find references0.03s2.3s
And it's the wrong tool for one repo on a Mac

Linux only. The docs say the build is "only tested on Linux so far." macOS is absent.

Haskell, GHC 9.6.7, plus folly and RocksDB built from source because the distro packages are too old. Marlow's own words: "Glean is not the easiest thing in the world to build."

He positions his own experiment as "an exploratory project demonstrating potential rather than production-ready technology" for outside-Meta use.

Why it felt like infra is that it is infra. Glean works at Meta because Meta's build system emits index data as a by-product of compilation, across a monorepo, for thousands of engineers. Outside that, you hand-roll the pipeline that fed it.

Google's equivalent, Kythe, is the same story with a worse ending.

The part worth stealing

Glean ingests SCIP and LSIF. That's the tell. The valuable half of "AST infra" is the indexer — the thing that turns code into accurate facts. The Datalog engine is what makes it feel like infra, and it earns its keep at billions of facts and thousands of concurrent users.

At one-repo scale, with an agent asking a handful of questions, SQL over SQLite answers the same questions. Same facts, from the same indexers, minus the Haskell.

What I'd pick

Layer 2 — the graph

GNU GLOBAL first, because it's boring, installs anywhere, and is exactly build-then-query. CodeGraph if you want the agent-shaped one and can accept a seven-month-old dependency.

Serena stays the answer if you decide MCP is acceptable after all — MIT, 24.8k stars, 40+ languages, real language servers. But it's MCP-only, and that's a genuine mismatch with how flw is built.

Layer 1 — the skeleton

Parse with tree-sitter, extract definitions and references, rank by how often a symbol is referenced elsewhere, print the top N within a token budget. A function called by 20 others outranks a private helper called once. This is aider's approach and it's the part worth writing to disk.

Test before choosing

There's a ground-truth case in this repo already. run_one has 8 real callers. The host LSP found 1.

Any candidate that returns 8 is decisively better. One that returns 1 carries the same blind spot under a different name. That's a twenty-minute test and it beats any table on this page.

What flw itself ships

None of it. flw stays stdlib-only, zero dependencies, language-agnostic. Bundling tree-sitter or an LSP client breaks all three, and it would put flw in the business of maintaining an indexer.

The dependency belongs in your repo and your environment, not in flw. Which also means your environment decides what's possible, not us.

So research does three things:

  1. Probe what's available — Serena, host LSP, ctags, an existing index.
  2. Verify it. Pick one known call edge, ask the graph, confirm with grep. Record whether it agreed.
  3. Record the working recipe and its blind spots in .flw/extensions/flw-spec.md.

Step 2 is the one I'd have skipped a week ago. Today's test is why it's in the list.

What I could not confirm

Sources