flw · research dossier 21 Aug 2026 4 agents · 15 threads

Scouting Unknown Code

What it takes to hand an agent the shape of a repository it has never opened — and why almost everything sold for the job falls over.

Build it. Adopt nothing. Change what you write to disk.

No off-the-shelf tool does build a graph, rank it, summarise it. Every one we examined has a documented scale failure, no independent evidence, or a popularity signal that can't be checked. Meanwhile structural retrieval beats dense embeddings in every published head-to-head we found, and a 120-line stdlib prototype produced usable orientation to repos nobody had opened, in under a second.

The one real surprise: the single published A/B test of repository overviews found they don't help. That changes the artifact, not the plan.

How to read the evidence

The hardest part of this research wasn't finding tools. It was separating measurement from marketing. A large share of what search returns for this category is AI-generated content restating vendor numbers as independent review.

One claim made it into an earlier draft of this research as independent before being caught: grepai's "97% fewer input tokens" is hosted on the tool author's own documentation site. Every claim below carries its provenance.

peer-reviewed accepted at a venue independent third party, no stake vendor author measuring themselves none found no evidence exists

The scale wall

Every claim in this section comes from a GitHub issue with numbers in it, not from a review.

Time to index · lower is better
Serena — LLVM, 7,749 files≈ 11 hours
Reporter states the machine was not compute-bound. independent
Our scout — same file count, extrapolated≈ 41 seconds
Measured 9.4s at 1,780 files, roughly linear. Not like-for-like — see below.

That bar is honest but the comparison is not like-for-like, and it matters. Serena builds a full semantic symbol index supporting precise navigation and refactoring. The scout builds an import graph supporting orientation. Serena is paying for something real — it just isn't the thing being asked for here.

ToolWhat brokeEvidence
Serena ~30GB RAM across three incidents, machine frozen; suspected unbounded cache independent
Serena A TypeScript lib file with 2,660 symbols crashed the terminal at 82%: "I have to reboot windows" independent
CodeGraph Out-of-memory at 44% on a 137,699-file repo — Fatal process out of memory: Zone independent
CodeGraph Vendor's own changelog: after 80 commits of incremental sync, 5.7% of graph edges were wrong; now 1.3% vendor
aider --map-tokens 1024 produced 16,419 actual tokens on a 552-file repo — 16× over independent
grepai No reports at any scale. Every source is the vendor or a content farm restating it none found

That CodeGraph staleness figure is the most credible number in the whole dig, precisely because it's an admission against interest. Even after the fix, 1.3% of edges are wrong after 80 commits.

Four findings that changed the plan

Ordered by how much each moved the decision.

1

Repository overviews may not help at all peer-reviewed

Evaluating AGENTS.md, N = 438 tasks, A/B-tested static prose repo overviews against no overview. They did not reduce steps-to-first-relevant-file and did not improve resolution rate. LLM-generated overviews hurt resolution by 0.5–2%.

It tested prose, not ranked structural maps — so our artifact type is untested rather than disproven. But the failure mechanism transfers: an overview costs context on every request whether or not it's relevant. Which is exactly what aider users report.

What it changes: cache the command, not the map. The scout runs in a third of a second. Persisting its output buys staleness and a permanent context tax; regenerating on demand costs nothing worth counting.

2

Structure beats embeddings, everywhere it's been measured peer-reviewed

This began as an unjustified omission — the research went straight to symbol graphs without arguing against semantic retrieval. The head-to-heads all point one way.

DevEval Pass@1 · dense embeddings vs graph
Graph-only43.06%
Pure dense embeddings (AlignCoder)23.60%
DyCoder, ASE 2026. Dense retrieval is nearly 2× worse at scale.
StudyResult
LocAgent, ACL 2025+10pp file-level · +19pp function-level
RANGER / DyCoder / SpIDER / RepoGraphhybrid beats either alone by 9–30% relative
Agent Retrieval Benchtested as separate arms, no fusion: no single family dominates

Caveat that applies to all of it: this is targeted retrieval against a known query. Query-free scouting has no dedicated benchmark at all.

3

Barrel files are unsolved in every existing tool independent

When index.ts re-exports everything, every import resolves to the barrel and a file-level ranking collapses onto it. No import-graph tool sees through it — not dependency-cruiser, not skott, not grimp or pydeps on the Python side. CPython's semantics guarantee the distortion: from pkg import name always edges to pkg first.

The algorithm exists, in a bundler. Next.js's optimize_barrel transform walks re-export chains, bails at the first file doing real work, and reattributes. Verified across 10,000 modules; cut one bundle from 552kB to 64kB. Nobody has published it applied to ranking.

4

We reimplemented aider exactly — and nobody has ever evaluated it none found

Source-verified against repomap.py: every constant matches verbatim. The prototype is aider's own cold-start configuration, independently re-derived from its blog post.

The 2023 post contains zero benchmark numbers. No paper evaluates the technique; those that mention it cite it only as related work.

No other agent in the corpus uses graph-theoretic relevance ranking.

Inside the Scaffold — a 13-scaffold survey, arXiv:2604.03515

So the technique is simultaneously unvalidated and unique to aider. There is no state of the art to be behind.

The industry disagrees with itself

The two most-used coding agents reached opposite conclusions, and both are still shipping their answer.

Claude CodeCursor
ApproachRemoved its RAG pipeline, May 2025. Grep-based agentic searchCustom-trained embedding model over Turbopuffer
Claimed result"it outperformed everything, by a lot"+12.5% avg accuracy · +2.6% retention on 1000+ file repos
Evidencevendorvendor

And the most uncomfortable number in the dig

Code Isn't Memory found that adding an index gave +7.9pp resolve rate over no index at all. But measured against a strong grep-agent comparator, the gain was not statistically significant — p = 0.087.

Which is to say: the case that any of this beats a competent agent with grep is not yet made.

What we measured ourselves

A ~120-line stdlib Python prototype, plus a Node equivalent that loads typescript out of the target repo's own node_modules so it installs nothing.

RepositoryFilesTime
Two-project Python tree1080.34s
Plugin codebase770.23s
Python standard library1,7809.4s
TypeScript plugin150.15s

Rank over imports, never over names

The first attempt ranked by how often a defined name appeared anywhere. The top results were a pytest fixture with 313 hits, then close(), then _utcnow(). Useless — because .get() on a dictionary is not a reference to your class's get method, and no weight multiplier repairs a parsing problem.

Switching the edges to imports — explicit, unambiguous, resolvable — surfaced the real architecture instead: SearchConfig, StorageDB, RawListing, ProfileSource.

This is the one place the prototype diverges from aider deliberately. Aider ranks over name references because tree-sitter across 130 languages cannot resolve imports uniformly. Python's ast and TypeScript's compiler API both can — so copying aider here would import a workaround for a problem we don't have.

Vendored code, fixed with a borrowed list

On one real repo, half the top ten was a vendored copy of tomlkit — its modules import each other heavily, which is indistinguishable from a well-factored core. Applying GitHub linguist's vendor.yml patterns removed all six entries and promoted the actual abstractions in their place. Borrowed, not invented: that list runs against every repository on GitHub.

The decision

Build the scout. Ship nothing else.

Four changes from what was about to be written, each traceable to a finding above.

  1. Generate on demand. Persist the recipe for orientation, never the orientation itself.
  2. Rank over imports, with linguist's vendor exclusions applied before ranking.
  3. Add the missing top rung. A study of professional auditors found they want global → local: what the project is for, then structure, then symbols. The current output has the bottom two.
  4. Use ts.resolveModuleName with a resolution cache for TypeScript aliases. The cache is mandatory at scale, not an optimisation.

Deliberately not built

Barrel collapse. Unsolved everywhere, and the algorithm to port is known — but a check across the actual target repositories found no TypeScript barrels at all and single-line __init__.py re-exports. Recorded as a known limitation rather than engineered against a problem that isn't there.

Monorepo machinery. Workspaces, project references, cross-package aliases. Not the deployment shape in question.

Four research agents across fifteen threads. One thread returned nothing and is recorded as returning nothing. Two claims made it into interim drafts before being caught: a fabricated arXiv identifier, and a vendor benchmark relayed as independent — both corrected in the record.

The full sourced report this summarises, with every citation, is held outside this repository.