Open Knowledge Format, explained — and whether to pour your DBA knowledge base into it

OKF is Google's draft answer to one question: how should AI agents read a body of knowledge without you dumping everything into context or hiding it in an opaque vector index? Its answer is deliberately boring — a folder of markdown files with YAML frontmatter. That boringness is the whole point.

One-paragraph verdict

Adopt the format, not (yet) the ecosystem. OKF is a thin, sensible spec and your DBA knowledge base is already ~80% of the way there — you have a markdown mirror and a clean organized/ tree. Adding OKF frontmatter + index.md/log.md + cross-links is cheap and fully reversible (it stays plain markdown). The real payoff is on the consumption side: one MCP server / Claude Project over all six courses with progressive disclosure and citations, plus a cross-course concept graph. But the spec is v0.1 draft, single-vendor, weeks old, and the third-party tools around it are hobby-grade — so build your own generator with the printing-press pattern and don't depend on anyone else's OKF software.

1 · What OKF actually is (and is not)

Three things ship in the GoogleCloudPlatform/knowledge-catalog repo, and it's worth separating them because people conflate them:

ThingWhat it isStatus
The OKF specA ~450-line markdown document (SPEC.md) defining the format: a directory of markdown + YAML frontmatter. This is the actual contribution.v0.1 DRAFT
The enrichment agentA reference producer — a Python agent on Google's ADK + Gemini that reads BigQuery metadata (+ crawls seed docs) and writes a bundle. Explicitly labelled a "proof of concept."PoC only
The viewer (viz.html)A reference consumer — a self-contained force-directed graph viewer (Cytoscape + marked, one HTML file, no backend).PoC only

What it is not: not a Google product, not a hosted service, not a schema registry, not a database. There is no central authority, no required SDK, no required tooling. In the spec's own words: "If you can cat a file, you can read OKF; if you can git clone a repo, you can ship it."

Its design center is data-catalog metadata (describing BigQuery tables, datasets, metrics) — but the format itself is domain-agnostic. A "concept" can be a table, an API, a metric, a playbook… or a lecture session. That generality is what makes it relevant to your use case.

2 · The problem it solves

When you want an agent to use a knowledge corpus today, you usually pick a bad option:

Option A — paste everything

Dump all transcripts/docs into context. Burns tokens, blows the window, and the model drowns in irrelevant material.

Option B — hidden vector index

Embed everything into a RAG store. Now retrieval is a black box you can't read, diff, audit, or version. You trust a similarity score you can't see.

OKF is the third option: structured, navigable, plain-text knowledge the agent traverses on demand — open the index, read only the relevant concept, follow its links to neighbours, cite the source. Everything stays human-readable and git-diffable. The bet is that established, accessible formats beat bespoke ones: markdown + YAML are already understood by humans, agents, and half the tools you own.

The four properties it optimises for, verbatim from the spec: Readable (no tooling), Parseable (no SDK), Diffable (version control), Portable (across tools, orgs, time).

3 · Anatomy of a bundle

A bundle is one directory tree. A concept is one markdown file inside it. Each concept has a YAML frontmatter block (the small, queryable layer) and a markdown body (the prose humans and LLMs actually read). Two reserved filenames carry meaning at any level: index.md (a directory listing for progressive disclosure) and log.md (an update history). Concepts link to each other with ordinary markdown links — which quietly turns the tree into a graph.

OKF bundle anatomy — tree on the left, a single concept document's frontmatter and body on the right
Diagram 1 — A bundle is a tree of markdown files (left). Zooming into one concept (right): required type + recommended fields in frontmatter, structural markdown body, and a cross-link that makes the tree a graph. Conformance rules are intentionally tiny (dark box).

4 · Produce → bundle → consume

The format's leverage is that it sits between many producers and many consumers as a neutral interchange layer. Nothing in the middle is proprietary, so any tool that can write markdown can produce it, and any tool that can read markdown can consume it. The reference agent and viewer are just one shape at each end.

OKF ecosystem — producers on the left feed one bundle in the centre, consumers on the right read it
Diagram 2 — Many producers (humans, enrichment agents, catalog exporters, your own CLIs) write into one plain-text bundle; many consumers (LLMs, MCP servers, Obsidian/Notion, the graph viewer, plain file servers) read it. The OKF spec is the only contract.

5 · The entire spec, in one table

OKF is "minimally opinionated" — you can hold the whole thing in your head. Here it is:

ElementRuleRequired?
type: frontmatterShort string identifying the kind of concept (BigQuery Table, Playbook, Session…). Not registered centrally; consumers must tolerate unknown types.YES — the only hard requirement
title / description / resource / tags / timestampRecommended frontmatter. resource = canonical URI of the underlying asset; absent for abstract ideas.Optional
Extra frontmatter keysAny producer-defined keys allowed; consumers preserve them and must not reject on unknown keys.Allowed
BodyStandard markdown. Conventional headings: # Schema, # Examples, # Citations. No required sections.Free-form
Cross-linksStandard markdown links. /absolute = bundle-relative (recommended, stable). Relationship meaning lives in the prose, not the link. Broken links are legal (= not-yet-written knowledge).Optional
index.mdReserved. A grouped listing of a directory's contents for progressive disclosure. No frontmatter (except optional okf_version at the root).Optional
log.mdReserved. Date-grouped change history, newest first, ISO-8601 dates.Optional

Conformance (v0.1): a bundle conforms if every non-reserved .md has parseable YAML frontmatter, every frontmatter has a non-empty type, and index.md/log.md follow their shape when present. Consumers must not reject a bundle over missing optional fields, unknown types, unknown keys, broken links, or a missing index. This permissiveness is deliberate — bundles are expected to be partially agent-generated and constantly refactored.

6 · Applied to YOUR knowledge base

Analysis Your DBA corpus is an unusually good fit because it's already structured. You have ~/DBA/.../organized/ with C1…C6 courses, each Session_NN_YYYY-MM-DD/ with video/ audio/ transcripts/ slides/ chat/ misc/, and a parallel markdown/ mirror. OKF doesn't ask you to move the heavy binaries — it asks you to lay a thin markdown layer beside them, where each .md carries frontmatter and a resource: pointer to the real asset.

Mapping the current DBA folder tree on the left to an OKF bundle on the right, with use cases unlocked
Diagram 3 — Left: what your CLIs already produce (binaries + flat text, no metadata or relationships). Right: the same content as an OKF bundle — a Course concept, a Session concept with frontmatter + asset links, a Transcript concept, and an abstract concepts/bias.md linked across C2/C3/C6. Binaries stay put; the .md points at them.

The mapping

Today (on disk)OKF concepttype:
A course folder (C6_Responsible_AI/)courses/c6-responsible-ai.mdCourse
A Session_NN_DATE/ foldersessions/c6-session-01.md (frontmatter: timestamp, resource: youtube://…, tags; body: summary + asset links + citations)Session
A transcript .txttranscripts/c6-session-01.md (body = plaintext, chunked by headings)Transcript
Slides / chat / case-study PDFsconcept docs with resource: → the fileSlides, Reference
(new) cross-cutting topicsconcepts/bias.md, concepts/transformer.md — abstract ideas, no resource, linked from every session that touches themConcept

7 · Use cases this unlocks

One agent over all 6 courses

Point a markdown MCP server (or a Claude Project) at the bundle. The agent searches the index, opens only the relevant session, and cites it — instead of you pasting a 4,000-line transcript.

Cross-course concept graph

Define bias, transformer, evaluation once and link them from every session. Now "where did the program cover bias mitigation?" traverses C2, C3 and C6 in one hop.

Your study log, for free

The bundle lives in git. git diff is your revision history; you can literally PR-review your own notes. log.md records each new download.

One source, many front-ends

The same bundle feeds Obsidian (graph + backlinks), the OKF viz.html force-graph, NotebookLM, and Claude — no re-export, no per-tool conversion.

Append-only growth

Future course downloads just drop new concept files and one log.md line. Your printing-press CLI already produces the raw material; it only needs a frontmatter-emitting step.

Auditable retrieval

Unlike a vector store, every retrieval is a file you can open and verify. For a doctoral programme where provenance matters, that's not cosmetic.

8 · Should you adopt it?

Structured so you can see what's verifiable vs. my read vs. my advice.

Facts

Analysis

The spec is genuinely well-judged: it standardises the minimum needed for interop and leaves everything else open. For you the value is lopsided — almost all of it is on the consumption side (agent retrieval, cross-course graph), and very little depends on OKF-the-spec specifically. You could get ~90% of the benefit by simply adding frontmatter and pointing an MCP markdown server at your existing tree, without conforming to any spec. What conformance buys you on top is interop with other people's OKF tools — and your need for cross-organisation exchange is currently zero. Confidence: Likely.

Strongest argument against adopting

OKF is a brand-new, single-vendor draft with a thin, hobby-grade tool ecosystem, and its design center is data catalogs — not lecture transcripts. The resource:/type:/# Schema conventions fit a BigQuery table cleanly and a 90-minute video awkwardly. Betting your study workflow on its tooling now is premature; the spec could change under you, and the consumer tools you'd want (a good MCP server, the viewer) are either generic-markdown tools that don't need OKF or PoCs that may never harden. A reasonable person concludes: add frontmatter for your own benefit, ignore the OKF brand entirely, and revisit in 6–12 months if the ecosystem proves itself. This is a serious position, not a strawman.

Recommendation

Adopt the format, decline the dependency. Concretely: extend your existing printing-press CLI to emit an OKF-conformant bundle from organized/ — it's a natural fit for the pattern you already use, and conforming costs almost nothing over "just adding frontmatter" while keeping the door open to OKF tooling if it matures. Then drive consumption with tools you already trust (Claude Project / a generic markdown MCP server), not third-party OKF software. You keep every upside (structure, graph, auditability, portability) and take on essentially no lock-in, because the worst case is "you have a folder of nicely-tagged markdown." Confidence: Likely. The one thing that would change my mind: if you actually need to exchange these bundles with classmates or upGrad systems — then strict conformance and the broader ecosystem suddenly matter, and the calculus shifts toward going all-in on the spec.

9 · A concrete migration path (if you proceed)

1
Generator step in the CLI. Add an okf subcommand that walks organized/C*/Session_*/ and emits one .md per course/session/transcript with frontmatter (type, title, timestamp from the folder date, resource = YouTube id / file path, tags).
2
Auto-build index.md + log.md. Generate a root index grouping by course, a per-course index grouping by session, and append a log.md line on every run. Put okf_version: "0.1" in the root index frontmatter.
3
Mint a few concepts/. Hand-write (or have an enrichment agent draft) abstract concept files for the recurring themes — bias, transformer, evaluation, rag — and link them from the sessions that cover them. This is where the cross-course graph comes alive.
4
Optional enrichment pass. Run an LLM over each transcript to write a # Summary and extract # Key concepts links into the session body. (This is exactly what Google's reference agent does, minus BigQuery.)
5
Consume. Point a generic markdown MCP server / Claude Project at the bundle. Optionally generate a one-file viz.html graph to browse the whole programme. Commit the bundle to git so diffs become your study log.
The whole thing is additive — it never touches your existing videos, audio, or transcripts. Worst case you delete the .md files and you're back to today.

Sources

Diagrams generated with the draw.io skill · editable .drawio + .svg + .png in ~/okf-research/diagrams/. This page: ~/okf-research/okf-explained.html.