Skip to content

Repository files navigation

ragstuff

A small RAG (Retrieval-Augmented Generation) pipeline that ingests PDF documents, splits them into chunks, embeds them, and stores them in a local Chroma vector database.

Setup

This project uses uv for dependency management.

uv sync

Usage

Place PDF files in the documents/ directory, then run:

uv run main.py

The chroma/ database directory is not checked into git, but you do not need to create it — Chroma creates it on first use. A fresh clone works straight away:

Hello from RAG stuff!
Loaded 37 documents from the directory.
Successfully added 72 chunks to Chroma

Search: 'sailboat instruments' -> 3 results

Ingestion runs before the search within the same main() call, so the database is already populated by the time the smoke test queries it. The first run is slower than later ones because the embedding model has to be downloaded and cached.

What happens when main() runs

  1. Load documents — PyPDFDirectoryLoader reads every PDF in the documents/ directory and loads it into a list of LangChain Document objects (one per page).
  2. Split into chunks — split_documents (in document_utils.py) uses a RecursiveCharacterTextSplitter to break each document into chunks of 1000 characters with 200 characters of overlap.
  3. Add to Chroma — add_to_chroma (in vector_db_utils.py) writes the chunks to a local Chroma database persisted in the chroma/ directory:
    • Each chunk gets a deterministic ID built from source:page:chunk_index, so the same chunk always maps to the same ID across runs.
    • Before inserting, any existing chunks belonging to the same source file are deleted, so re-running main() on an updated document replaces its old chunks instead of duplicating them.
    • Chunks are embedded using a HuggingFace sentence-transformer model (get_embedding_function, defaulting to the "mini" model — all-MiniLM-L6-v2) and added to the database.
  4. Verify the database is searchable — search_chroma runs a similarity search against the freshly written database using SMOKE_TEST_QUERY, and print_search_results prints the top 3 hits with their source, page, score, and a text preview. This confirms the ingestion actually produced a queryable index rather than just reporting success.

The end result is a persisted, queryable vector store in chroma/ containing embeddings for every chunk of every PDF in documents/.

Searching

search_chroma(query, k=5) returns a list of (Document, score) tuples ordered by relevance. The score is a distance, so lower means a closer match.

from vector_db_utils import search_chroma, print_search_results

results = search_chroma("sailboat instruments", k=3)
print_search_results("sailboat instruments", results)

Tests

uv run pytest

Tests live in tests/, mirroring the module layout of the project root.

About

A small RAG pipeline: ingest PDFs, chunk and embed them, store in a local Chroma vector database

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages