The common answer is embed everything into vectors and rank by cosine distance.
OKTS takes a different route: portable text descriptors + a derived graph.
Pick a query, hit play, and watch both pipelines run the same request side by side.
catalog size:
🧭 Vector-embedding search
the common way — “RAG over tools”
needs an embedding model / API
embedding space (random, for illustration)
embed API calls
0
query latency
0 ms
re-index on edit
—
Top-k by cosine distance
🗂️ OKTS descriptors + graph
rank on prose · prefilter by hierarchy · expand the graph
zero required API calls · works offline
catalog as a graph (hierarchy + alternatives)
embed API calls
0
query latency
0 ms
re-index on edit
—
Top-k · hybrid + graph-expanded
10embed API calls to build the index · re-embed on every edit
vs
0embed calls · index is text · no re-embed round-trips at any scale
Dimension
Vector-embedding search
OKTS descriptors + graph
Index artifact
opaque float vectors
human-readable OKT markdown, git-diffable
To add / edit a tool
re-embed (model or API call)
edit text → re-rank instantly, no model
Runtime dependency
embedding model must be reachable
lexical fallback runs fully offline
Knows about alternatives?
no — distance only
yes — alternatives edges expand results
Knows about categories?
no structure
hierarchy prefilter narrows the field
Why a match ranked
unexplainable cosine number
matched terms + edges you can read
The honest nuance: OKTS isn’t “anti-embedding.” Its retriever is hybrid — it can use dense
vectors too. The difference is they’re optional: the portable descriptor and its derived graph are the
backbone, so retrieval degrades gracefully to lexical BM25 when no embedding model is present — and you never
pay the re-embed / API tax just to add a tool. Embeddings become an upgrade, not a hard dependency.
And to be precise: OKTS query latency isn’t zero — both engines do an O(N) scan and top-k at query time.
What OKTS avoids is the per-query embedding round-trip layered on top of that scan, plus the index-time
re-embedding — which is where the “query latency” and “cost at scale” numbers above diverge.
(Those millisecond figures are illustrative, not benchmarks: they all derive from one assumed
~85 ms embedding round-trip, and real index builds batch and parallelize — hence the “if sequential” caveat.)