LUCIANO CARRIZO
YouTube Knowledge Pipeline
All projects
Completed

YouTube Knowledge Pipeline

A local pipeline that turns YouTube videos into verifiable knowledge packages and indexes them to deliver broad, deduplicated and cited context to agents. Two repositories—an Agent Skill and a RAG system with no internal LLM—joined by an explicit evidence and provenance contract.

PythonTypeScriptNode.jsSQLiteMarkdown
The system

From opaque video to auditable knowledge

A video concentrates transcript, images, sequences, and claims whose origin is easily lost when it is summarized. I built this pipeline so another agent can reason over a collection without reopening every video or separating an answer from the evidence supporting it.

The experience spans two public repositories. youtube-video-context is the Agent Skill that downloads the material, inspects visual evidence, and produces a validated package. auto-youtube-rag reads those packages without modifying them, indexes them, and assembles cited context for the agent that will perform the final interpretation.

Pipeline from a YouTube video through a skill, a validated knowledge package, a local RAG system, cited context, and an agent
Two systems with their own technical lifecycles, connected by a verifiable data contract.
The skill

Downloading captions and summarizing is not enough

The skill is defined against capabilities—shell, reading, writing, search, and image inspection—so its procedure does not depend on one provider. For every video, it preserves the source, transcript, metadata, and visual coverage; before authoring, it builds a ledger that relates every claim to its evidence, classification, and limits.

01

Baseline and adaptive coverage

It extracts exactly 20 uniform frames from 0% through 95%, then complements them with scene changes, temporal anchors, dense sequences, and supplementary captures when motion or an omitted scene changes the meaning.

02

Evidence with an explicit class

It keeps six classes separate: direct source, visual confirmation, time-sensitive claim, unverified claim, analyst judgment, and recommendation. The dossier does not present every observation with the same level of certainty.

03

Artifacts for people and machines

It delivers context.md for reading and analysis.json schema 2.0 for structured consumption, together with transcript, metadata, frames, and a visual coverage record.

04

Validation before publication

The validator checks structure, identity, keys, depth, cited frames, and declared coverage. Semantic review and the language contract remain in an explicit manual checklist.

That last boundary matters: validate-dossier.py is a package validator, not a code test suite. It can prove that canonical artifacts exist and are consistent, but it cannot decide whether the analysis understood the video correctly. Automated extraction and semantic review are recorded separately in visual/coverage.json.

The contract

Two repositories, one explicit data boundary

The skill can finish its work without the RAG system; the RAG system, however, was designed as a consumer of that content. The boundary is expressed in files rather than an implicit integration: manifest.json declares the collection and its resources; context.md preserves the readable dossier; analysis.json schema 2.0 provides the structured representation that the parser validates before it enters the application.

The source stays still

sync only reads packages. The database, model, and results live in the RAG system's local library, so indexing or rebuilding never rewrites the evidence produced by the skill.

The contract can evolve

The current reader understands analysis.json 2.0 and retains compatibility with rules.json 1.0 for historical collections. The two formats are mutually exclusive within a package.

This decision allows Python and TypeScript to remain in separate repositories without turning them into two disconnected experiences in the portfolio.

The RAG system

Incremental indexing, isolated domain, replaceable infrastructure

auto-youtube-rag uses a domain-centered ports-and-adapters architecture. The domain keeps identities, entities, and rules without knowing SQLite, the embedding model, or the CLI; the application defines use cases and ports; infrastructure implements filesystem, persistence, embeddings, and search; main composes the concrete pieces.

01

Incremental synchronization

Per-source hashes make indexing incremental and idempotent. The document → section → unit hierarchy and provenance are preserved without touching the original package.

02

Local persistence

SQLite and FTS5 store the catalog, text, and index state. The library can be rebuilt from immutable sources if an infrastructure decision changes.

03

Embeddings on the machine

E5 Small generates local vectors, and exact in-memory search covers the semantic path selected for the MVP. The network is only involved when the model is downloaded during initialization.

04

Retrieval, not generation

The system contains no LLM and does not produce the final answer. Its responsibility ends when it delivers context.md and result.json with evidence and provenance to another agent.

Keeping those boundaries made infrastructure details replaceable and kept core behavior testable without launching a human interface or depending on a remote service during use.

The evidence

Broad retrieval without losing provenance

A query combines FTS5 and vector search, merges both paths through weighted Reciprocal Rank Fusion, expands ancestors, deduplicates, and diversifies before assembling the context budget. The focused, balanced, and deep modes change how much material enters, but they do not remove the reference connecting each fragment to its source.

Evaluation

Citation integrity: 24/24

The August 12, 2026 report covers 8 queries at 3 depths. All 24 bundles kept their [S0N] citations aligned with result.json, while a SHA-256 digest verified that the collection did not change during sync and retrieval.

Benchmark

31.71 ms versus 156.60 ms

Using the same 50,000-vector fixture, exact in-memory search recorded a 31.71 ms p50 and sqlite-vec 156.60 ms, with 100% matching top-k sets. This is a measurement from one documented machine, not a universal speed claim.

The automated evidence belongs to the RAG system: it has 65 *.test.ts files, plus type-test, smoke, and E2E coverage. Tests cross the domain, application with fakes, shared adapter contracts, real SQLite, CLI, and integration; the CI pipeline runs checks and the build. The file count does not prove quality by itself, but it makes it possible to trace that retrieval, persistence, and assembly decisions are exercised beyond the README.

Closing

Technical separation supports one experience

Local-first, without a generative brain

The skill preserves evidence and the RAG system makes it queryable on the machine. There is no human UI or internal LLM: the consulting agent interprets the context and writes the answer.

Separated for maintenance, united in use

The skill is portable and can be used by itself. The RAG system depends on its package contract. Both repositories keep their own technical lifecycles while the portfolio presents one functional unit.

The result does not try to hide that asymmetry. Generating verifiable knowledge and retrieving it are different problems; the explicit contract makes it possible to solve them separately without losing the experience that motivated the project: reasoning over a video collection without reopening it or giving up provenance.

Next step

The case shows the code. Let’s talk about what comes next.

YouTube Knowledge Pipeline is documented end to end: architecture, decisions, and what remains unfinished. If you have questions about the approach, write to me.

Let’s talk

or directly → LuchoC.dev@gmail.com