CapybaraDB
A vector database built from scratch in Python, implementing indexing, similarity search, persistence, and the full query pipeline.
- Python
- Pytorch
- RAG
- Vector Search
- <10ms Query latency
- 100% Recall @ k=5
CapybaraDB is a lightweight vector database built from scratch in Python to make the full semantic-search pipeline inspectable: text ingestion, token-aware chunking, embedding generation, vector storage, exact similarity retrieval, and persistence. It deliberately avoids FAISS and other ANN libraries in the core path, using NumPy and PyTorch instead.
The goal is not to replace a production distributed vector database. It is to expose the mechanics hidden by one: how documents become vectors, what is persisted with those vectors, how precision and device choices affect the representation, and where end-to-end query time is actually spent.
Technical deep dive: Building a Vector Database from Scratch - CapybaraDB.
API and data flow
The primary interface is CapybaraDB. A caller selects a named collection and configures whether documents are chunked, the chunk size, numeric precision, and execution device:
from capybaradb.main import CapybaraDB
db = CapybaraDB(
collection="my_docs",
chunking=True,
chunk_size=512,
precision="float32",
device="cuda",
)
db.add_document("Capybaras are the largest rodents in the world.")
results = db.search("biggest rodent", top_k=3)
db.save()
Indexing follows a linear pipeline:
source document
-> optional file extraction
-> optional token chunking
-> embedding generation
-> vector + text + document metadata
-> named collection storage
A query takes the inverse path: the query text is embedded with the same representation, compared with the collection vectors, ranked by similarity, and returned as text, document metadata, and score. Exact search means every indexed vector participates in ranking. The implementation makes this cost explicit rather than obscuring it behind an approximate index.
Ingestion and chunking
CapybaraDB accepts text directly and can extract text from PDF, DOCX, and TXT inputs, including OCR support where needed. Extracted text enters the same add_document pipeline as a string supplied by the caller, keeping file format handling at the ingestion boundary rather than in retrieval.
Chunking is optional. When enabled, token-based chunking with tiktoken breaks a long source into smaller retrieval units. Each chunk is embedded and stored with enough document association to return the originating document in search results. This is the key recall-versus-context trade-off in the system:
- Larger chunks preserve more surrounding context and produce fewer vectors.
- Smaller chunks improve passage-level targeting but increase embedding work, storage, and the number of vectors examined per exact query.
- Disabling chunking makes a document the retrieval unit, which is simpler but too coarse for long documents.
The default demonstration uses 512-token chunks. The benchmark’s source documents are mostly shorter than that, so it measures document-level retrieval in practice. That caveat matters: strong metrics on short documents should not be generalized to passage retrieval in long corpora without a new evaluation.
Vector representation and storage
Embeddings are generated with sentence-transformers, then retained alongside the source text and collection metadata. The database supports in-memory operation for experimentation and persistent storage through serialization for reuse across processes. Named collections provide separate logical indexes without requiring a server-side multi-tenant control plane.
The caller can choose float32, float16, or binary precision and select CPU or CUDA execution. These are explicit trade-offs, not hidden defaults:
float32preserves the standard dense representation at the highest memory cost.float16reduces vector memory and can improve accelerator throughput at reduced numerical precision.- Binary representation minimizes storage further but changes the fidelity of similarity computation.
- CUDA can accelerate embedding and vector operations when compatible hardware is available; CPU keeps the system portable.
The project persists the index rather than only a model cache, so search does not require re-ingesting the collection. save, load, and clear define the lifecycle of a collection; get_document exposes the original stored text by document ID.
Exact similarity retrieval
Search embeds the query, computes similarity against the indexed collection, and returns the highest-scoring top_k matches. The core retrieval operation is vectorized rather than implemented as a Python loop, which keeps the per-vector comparison in optimized numerical kernels.
Exact search has an important and intentional complexity property: retrieval work grows linearly with the number of vectors. There is no graph traversal, quantization probe, or probabilistic candidate selection. At a larger corpus size this becomes the ceiling, but at small to medium collection sizes it provides predictable behavior and no approximation-induced recall loss.
The benchmarks show why the simple design is useful within the tested range. Across 100 to 5,000 vectors, average query latency ranged from 7.54ms to 9.10ms, with throughput remaining above 109 queries per second. p50 latency ranged from 7.45ms to 8.79ms; p95 was 10.09ms to 12.01ms, and p99 was 11.80ms to 16.39ms.
The time budget is split roughly evenly between query embedding and retrieval: embedding took about 3.87–4.53ms and vector retrieval about 3.50–4.57ms. This is useful operationally because replacing the index alone would not halve end-to-end latency; query embedding is an equally material part of the path.

Indexing cost and footprint
Indexing measures extraction and chunking preparation, embedding generation, and storage. The benchmark processed 10, 50, 100, 500, and 1,000 documents. Total indexing time grew from 0.138 seconds at 10 documents to 76.331 seconds at 1,000. The average cost per document rose from 13.8ms to 76.3ms as batches grew and memory pressure appeared.
Embedding dominates this work. At 1,000 documents, storage took approximately 0.122 seconds of a 76-second indexing run, while the serialized index occupied 1.715MB. Peak memory ranged from roughly 2.2MB to 65.4MB across the tested scales. The result is a compact index with a clear bottleneck: improving ingestion throughput requires reducing or batching embedding work before optimizing persistence.

Retrieval quality
Quality was measured on a synthetic dataset with 512-token chunks. At k=1, precision and nDCG were both 1.00, showing that the highest-ranked result was relevant for the benchmark queries. Recall reached approximately 0.956 at k=3 and 1.00 at k=5; nDCG@5 and nDCG@10 were 0.979.
Precision necessarily declines as k expands because the result set includes more candidates: P@3 was about 0.756, P@5 about 0.480, and P@10 about 0.240. The useful conclusion is not that all top-10 results are precise; it is that the expected relevant material appeared by five results in this synthetic, short-document setup.

The benchmark also reported a negative min_latency_ms value at the 500-vector size. That is a measurement artifact and should not be interpreted as a real latency result. The percentile measurements and averages are the reliable basis for the performance conclusions above.
What the project demonstrates
CapybaraDB makes several engineering trade-offs tangible:
- Semantic retrieval is an end-to-end pipeline, not only a nearest-neighbor operation.
- Exact vector search can provide stable sub-10ms retrieval at thousands of vectors when implemented with vectorized numerical operations.
- Embedding time is as important as similarity computation in the observed query budget.
- Chunk size changes both retrieval quality and the size of the exact-search problem.
- Benchmark quality depends on the corpus shape; document-level synthetic results are not evidence of long-document passage retrieval quality.
The natural next scaling step is not to prematurely add an ANN dependency. It is to validate smaller chunk sizes and longer corpora, measure the real collection size where exact search stops meeting latency targets, and only then compare an approximate index against the measured recall, memory, and operational costs.