Eval-First RAG
A learning project: retrieval-augmented generation built from scratch, with evals to show where it fails.
Python · FastAPI · Claude API · sentence-transformers · ChromaDB · Jupyter
The problem
RAG systems fail quietly. When an answer is wrong, was it retrieval or generation? Frameworks abstract away exactly the parts you need to understand to answer that question.
The approach
I built a RAG pipeline over the Kubernetes documentation with no frameworks. Chunking, embeddings, a ChromaDB vector index, retrieval and Claude-based generation each became an explicit, inspectable step, and then I wrote an evaluation suite that scores retrieval and generation separately.
The result
A hands-on reference that makes RAG legible: a notebook-by-notebook progression from raw documents to an evaluated pipeline, plus a FastAPI service. The evals tell you where the system is failing, not just that it failed.
Highlights
- No frameworks, so every RAG component is explicit
- Separate evals for retrieval vs generation quality
- Notebook progression from documents to a served pipeline
Want to talk through any of this? Get in touch.