RAG as a service
RAG that's already built
Most teams spend two months assembling chunking, embedding, retrieval, and reranking — then another three tuning it. Zenmem ships that stack hardened, self-hosted, and ready for production traffic.
The problem
Retrieval is easy to demo and hard to ship
A prototype takes an afternoon: split the text, embed it, cosine search, done. The gap between that and production is where the months go. Naive character splitting cuts a function in half and returns something the model can't use. Pure vector search misses exact identifiers — the error code, the class name, the SKU — because embeddings smooth away the specificity that made the query useful. Without a second-stage reranker, the top result is often the fourth-best answer. And every one of those problems only shows up under real queries, after you've shipped. You end up owning a retrieval stack you didn't set out to build, and tuning it forever.
What you get
The whole retrieval path, already solved
Structure-aware chunking
Code is split at class and method boundaries using tree- sitter, not at character counts. Documents are split at section boundaries with the heading hierarchy carried into each chunk, so a paragraph retrieved from page 40 still knows which chapter it belongs to.
Hybrid search, every tier
Dense embeddings find meaning; sparse retrieval finds exact terms. Zenmem runs both and fuses the results with reciprocal rank fusion. You don't configure it and you don't pay more for it — it's on for every account.
Reranking on by default
A cross-encoder scores the candidate pool against the actual query before results are returned. This is the single highest-leverage step in retrieval quality, and it's the one most stacks skip because it's fiddly to wire. Here it's the default path.
Retrieval that learns
Zenmem records which retrieved memories the model actually used and feeds that signal back into a contextual bandit. Retrieval quality improves with traffic instead of decaying as your corpus grows.
How it fits
Two calls to integrate
Write memories as they're created. Read them back when the agent needs context. Everything between — chunking, embedding, indexing, hybrid search, reranking — happens inside Zenmem.
client = Client(Config(vectorDbUrl="http://localhost:6636", companyCode="ACME")
client.addMemory("Auth tokens expire after 15 minutes", tags={"service": "auth"
mem = client.fetchMemory("how long do tokens last?") print(mem.memoryText)
Self-hosted with one command. Your vectors never leave your infrastructure.
from zenmem import Client, ConfigWhy not build it
The build-versus-buy math
Hand-rolling this is roughly 2–4 engineer-months to first production deployment, and the maintenance never ends — embeddings models change, retrieval quality drifts, and eval tooling becomes its own project. Zenmem is a day to integrate and stays flat-priced whether you serve a hundred queries or a million. There's no per-call metering and no retrieval feature held back for a higher tier. CTA Compare all options →
FAQ
Do I have to migrate my existing vector database?
No. Zenmem runs its own Qdrant- backed store, so you can point it at a subset of your data and run both in parallel until you're satisfied with retrieval quality.
Which embedding model does it use?
A local sentence-transformers model by default, which means no per-embedding API cost and no data leaving your network. You can send pre-computed embeddings instead if you already have a model you trust.
Is this hosted or self-hosted?
Self-hosted by design — one Docker command brings up the full stack on your own infrastructure. A managed option is available if you'd rather not run it yourself.
Stop building retrieval
One command to install, two functions to integrate.