A small RAG (Retrieval-Augmented Generation) pipeline that ingests PDF documents, splits them into chunks, embeds them, and stores them in a local Chroma vector database.
This project uses uv for dependency management.
uv syncPlace PDF files in the documents/ directory, then run:
uv run main.pyThe chroma/ database directory is not checked into git, but you do not need to create it — Chroma creates it on first use. A fresh clone works straight away:
Hello from RAG stuff!
Loaded 37 documents from the directory.
Successfully added 72 chunks to Chroma
Search: 'sailboat instruments' -> 3 results
Ingestion runs before the search within the same main() call, so the database is already populated by the time the smoke test queries it. The first run is slower than later ones because the embedding model has to be downloaded and cached.
- Load documents —
PyPDFDirectoryLoaderreads every PDF in thedocuments/directory and loads it into a list of LangChainDocumentobjects (one per page). - Split into chunks —
split_documents(in document_utils.py) uses aRecursiveCharacterTextSplitterto break each document into chunks of 1000 characters with 200 characters of overlap. - Add to Chroma —
add_to_chroma(in vector_db_utils.py) writes the chunks to a local Chroma database persisted in thechroma/directory:- Each chunk gets a deterministic ID built from
source:page:chunk_index, so the same chunk always maps to the same ID across runs. - Before inserting, any existing chunks belonging to the same source file are deleted, so re-running
main()on an updated document replaces its old chunks instead of duplicating them. - Chunks are embedded using a HuggingFace sentence-transformer model (
get_embedding_function, defaulting to the"mini"model —all-MiniLM-L6-v2) and added to the database.
- Each chunk gets a deterministic ID built from
- Verify the database is searchable —
search_chromaruns a similarity search against the freshly written database usingSMOKE_TEST_QUERY, andprint_search_resultsprints the top 3 hits with their source, page, score, and a text preview. This confirms the ingestion actually produced a queryable index rather than just reporting success.
The end result is a persisted, queryable vector store in chroma/ containing embeddings for every chunk of every PDF in documents/.
search_chroma(query, k=5) returns a list of (Document, score) tuples ordered by relevance. The score is a distance, so lower means a closer match.
from vector_db_utils import search_chroma, print_search_results
results = search_chroma("sailboat instruments", k=3)
print_search_results("sailboat instruments", results)uv run pytestTests live in tests/, mirroring the module layout of the project root.