Product Vision

Turn DBA course material into an open, AI-augmented knowledge base.

A pipeline that ingests raw DBA content — transcripts, slides, videos, references — enriches and structures it in the Open Knowledge Format (OKF), and publishes it to a public GitHub organization as a navigable knowledge base for DBA cohorts.

Objective: Ship the public KB Audience: Public / DBA cohorts Explorations: Parked Artifact: This page + roadmap
The shape of the system

A three-stage pipeline

Raw material flows from source to a published knowledge base. Feedback loops run back across every stage.

Stage 01

Source

Ingest raw material from UpGrad & GitHub.

  • Transcripts
  • Slides & PDFs
  • YouTube videos
  • References
  • Sources & processed folders
Stage 02

Processing

Enrich, de-duplicate, structure.

  • Dedupe
  • Orchestration loop
  • Enrichment (cohort / slot)
  • OKF + AI layer
  • Quality: evals, UI/UX
Stage 03

Destination

Publish & engage on a GitHub org.

  • Courses → Sessions
  • Concepts → References
  • Assignments → Submissions
  • Public GitHub organization
  • Marketing & engagement
Source ──▶ Processing ──▶ Destination · with Feedback Loops closing back across all three
At the centre

AI + DBA Knowledge Base, in OKF

The core asset is the DBA corpus expressed in the Open Knowledge Format — a stable, portable, git-friendly schema — with an AI layer for enrichment, search, and tutoring on top.

Built

OKF corpus

118 nodes across 6 courses — courses · sessions · concepts · references — with cross-links and enriched bodies.

Built

Knowledge-map viewer

Interactive graph (v2) with progressive disclosure, full rendered content, and math — the public-facing UX foundation.

Next

AI layer

Evals + AI-assisted enrichment and Q&A over the corpus. Quality-gated before publication.

How the work is organized

Six workstreams

The sketch maps onto six durable workstreams. Explorations are tracked separately as a parked backlog.

W1

Ingestion

Pull and normalize raw content from UpGrad & GitHub.

Owns: source types · raw → processed folders
W2

Enrichment & Processing

Dedupe, orchestrate the loop, enrich per cohort/slot.

Owns: dedupe · orchestration · assets
W3

Knowledge Base (OKF)

The canonical OKF schema, AI layer, and content structure.

Owns: OKF · structure · AI core
W4

Quality & UX

Evals on content; the public viewer and site experience.

Owns: evals · UI/UX · viz
W5

Community

GitHub org, contributors, and feedback loops.

Owns: collaboration · feedback loops
W6

Growth

Marketing, engagement, teaching materials, audio.

Owns: marketing · teaching · ElevenLabs
Who it serves

Two personas

User

A DBA student who studies the knowledge base — reads, searches, and navigates concepts, sessions, and references to learn faster.

Contributor

A community member who adds, corrects, and enriches content through the open-source GitHub project and its feedback loops.

Path to launch

Shipping roadmap

Focused on one goal: a public OKF knowledge base for DBA cohorts. Explorations stay parked until after launch.

Phase 0Done

Foundation

OKF DBA bundle (118 nodes) and the v2 knowledge-map viewer.

Phase 1Now

Harden the knowledge base

Complete courses → sessions → concepts → references → assignments → submissions; dedupe; close content gaps (e.g. C6/C7); enrich Cohort-8 slots.

Phase 2Next

Quality & UX

Run evals on content; polish the viewer and site for public use.

Phase 3Next

Publish

Stand up the GitHub organization + repo layout, README, ARCHITECTURE, CONTRIBUTING, license, and OKF spec docs.

Phase 4Next

Public launch

Site live for DBA cohorts; engagement & marketing; feedback loops via issues and PRs.

Not now

Parked explorations

Real ideas, deliberately out of the committed roadmap. Revisit after public launch.

Google Agents Web-agents AIR ? Interactive website ElevenLabs audio Hyperframes ? Teaching materials Prompt-injection handling ?
To confirm

Open questions

Items illegible in the source sketch or under-specified — confirm to lock the capture.

  • 5th content type — what is it? (transcripts · slides · videos · references · ?)
  • Enrichment grid — the faint table: what does it track (cohort × slot × asset)?
  • Feedback loops — what mechanism (issues/PRs · email · in-app)?
  • "AIR" & "Hyperframes" — what are these under Explorations / deliverables?
  • Prompt injection — a threat to defend against, or a technique to use?
  • "Train" — train what (a model · embeddings · fine-tune)?