Turn DBA course material into an open, AI-augmented knowledge base.
A pipeline that ingests raw DBA content — transcripts, slides, videos, references — enriches and structures it in the Open Knowledge Format (OKF), and publishes it to a public GitHub organization as a navigable knowledge base for DBA cohorts.
A three-stage pipeline
Raw material flows from source to a published knowledge base. Feedback loops run back across every stage.
Source
Ingest raw material from UpGrad & GitHub.
- Transcripts
- Slides & PDFs
- YouTube videos
- References
- Sources & processed folders
Processing
Enrich, de-duplicate, structure.
- Dedupe
- Orchestration loop
- Enrichment (cohort / slot)
- OKF + AI layer
- Quality: evals, UI/UX
Destination
Publish & engage on a GitHub org.
- Courses → Sessions
- Concepts → References
- Assignments → Submissions
- Public GitHub organization
- Marketing & engagement
AI + DBA Knowledge Base, in OKF
The core asset is the DBA corpus expressed in the Open Knowledge Format — a stable, portable, git-friendly schema — with an AI layer for enrichment, search, and tutoring on top.
OKF corpus
118 nodes across 6 courses — courses · sessions · concepts · references — with cross-links and enriched bodies.
Knowledge-map viewer
Interactive graph (v2) with progressive disclosure, full rendered content, and math — the public-facing UX foundation.
AI layer
Evals + AI-assisted enrichment and Q&A over the corpus. Quality-gated before publication.
Six workstreams
The sketch maps onto six durable workstreams. Explorations are tracked separately as a parked backlog.
Ingestion
Pull and normalize raw content from UpGrad & GitHub.
Enrichment & Processing
Dedupe, orchestrate the loop, enrich per cohort/slot.
Knowledge Base (OKF)
The canonical OKF schema, AI layer, and content structure.
Quality & UX
Evals on content; the public viewer and site experience.
Community
GitHub org, contributors, and feedback loops.
Growth
Marketing, engagement, teaching materials, audio.
Two personas
User
A DBA student who studies the knowledge base — reads, searches, and navigates concepts, sessions, and references to learn faster.
Contributor
A community member who adds, corrects, and enriches content through the open-source GitHub project and its feedback loops.
Shipping roadmap
Focused on one goal: a public OKF knowledge base for DBA cohorts. Explorations stay parked until after launch.
Foundation
OKF DBA bundle (118 nodes) and the v2 knowledge-map viewer.
Harden the knowledge base
Complete courses → sessions → concepts → references → assignments → submissions; dedupe; close content gaps (e.g. C6/C7); enrich Cohort-8 slots.
Quality & UX
Run evals on content; polish the viewer and site for public use.
Publish
Stand up the GitHub organization + repo layout, README, ARCHITECTURE, CONTRIBUTING, license, and OKF spec docs.
Public launch
Site live for DBA cohorts; engagement & marketing; feedback loops via issues and PRs.
Parked explorations
Real ideas, deliberately out of the committed roadmap. Revisit after public launch.
Open questions
Items illegible in the source sketch or under-specified — confirm to lock the capture.
- 5th content type — what is it? (transcripts · slides · videos · references · ?)
- Enrichment grid — the faint table: what does it track (cohort × slot × asset)?
- Feedback loops — what mechanism (issues/PRs · email · in-app)?
- "AIR" & "Hyperframes" — what are these under Explorations / deliverables?
- Prompt injection — a threat to defend against, or a technique to use?
- "Train" — train what (a model · embeddings · fine-tune)?