OKTS
← Playground
60-second explainer · press play

How do you find the right tool out of hundreds?

The common answer is embed everything into vectors and rank by cosine distance. OKTS takes a different route: portable text descriptors + a derived graph. Pick a query, hit play, and watch both pipelines run the same request side by side.

catalog size:

🧭 Vector-embedding search

the common way — “RAG over tools”
needs an embedding model / API
embedding space (random, for illustration)
embed API calls
0
query latency
0 ms
re-index on edit
—

🗂️ OKTS descriptors + graph

rank on prose · prefilter by hierarchy · expand the graph
zero required API calls · works offline
catalog as a graph (hierarchy + alternatives)
embed API calls
0
query latency
0 ms
re-index on edit
—
10embed API calls to build the index · re-embed on every edit
vs
0embed calls · index is text · no re-embed round-trips at any scale
DimensionVector-embedding searchOKTS descriptors + graph
Index artifactopaque float vectorshuman-readable OKT markdown, git-diffable
To add / edit a toolre-embed (model or API call)edit text → re-rank instantly, no model
Runtime dependencyembedding model must be reachablelexical fallback runs fully offline
Knows about alternatives?no — distance onlyyes — alternatives edges expand results
Knows about categories?no structurehierarchy prefilter narrows the field
Why a match rankedunexplainable cosine numbermatched terms + edges you can read
The honest nuance: OKTS isn’t “anti-embedding.” Its retriever is hybrid — it can use dense vectors too. The difference is they’re optional: the portable descriptor and its derived graph are the backbone, so retrieval degrades gracefully to lexical BM25 when no embedding model is present — and you never pay the re-embed / API tax just to add a tool. Embeddings become an upgrade, not a hard dependency. And to be precise: OKTS query latency isn’t zero — both engines do an O(N) scan and top-k at query time. What OKTS avoids is the per-query embedding round-trip layered on top of that scan, plus the index-time re-embedding — which is where the “query latency” and “cost at scale” numbers above diverge. (Those millisecond figures are illustrative, not benchmarks: they all derive from one assumed ~85 ms embedding round-trip, and real index builds batch and parallelize — hence the “if sequential” caveat.)