AI Systems / RAG
Secure Multi-Tenant RAG
A production-oriented retrieval-augmented generation platform with per-tenant isolation and access control.

Chatting with documents and attempting to access information from another department.
Showing image 1 of 4
01
Most RAG tutorials assume a single user and a single document set. Real organizations need multiple tenants and users, each with different documents and permissions, and they need confidence that one tenant's data can never leak into another tenant's answers — while retrieval quality still has to hold up against messy, real-world documents.
02
The system ingests documents per tenant (including direct Google Drive ingestion), indexes them for both keyword and semantic search, and enforces tenant and role boundaries at the data-access layer rather than only in the prompt. Retrieval combines BM25 keyword search with vector search, merges the results, and reranks them before they ever reach the language model, so answers are grounded in the right, permitted documents.
Challenges & decisions
Isolating tenants without duplicating infrastructure
Rather than standing up separate infrastructure per tenant, tenant and role scoping is enforced at the data-access layer so a single deployment can safely serve multiple organizations.
Balancing keyword and semantic retrieval
Pure vector search misses exact terms and identifiers; pure keyword search misses paraphrased questions. Combining BM25 and vector search, then reranking the merged results, gave more reliable grounding than either approach alone.
Treating the LLM as an untrusted boundary
User input and retrieved content are both treated as potentially adversarial. An AI firewall layer screens for prompt injection and out-of-scope requests before they reach the model.
03
Secure Multi-Tenant RAG is a backend-first AI system designed to let multiple organizations query their own knowledge bases through a single deployment, without ever mixing data between tenants. It focuses on the parts of RAG that most demos skip: tenant isolation, access control, retrieval quality, and guarding the system against malicious or unsafe inputs.
How it works
A FastAPI backend exposes ingestion and query APIs behind role-based access control (RBAC), with every request scoped to a tenant and a user role. Documents are ingested from sources like Google Drive, chunked, embedded, and written to a vector store alongside a BM25 index for hybrid search. At query time, candidates from both retrieval paths are merged and passed through a reranking step before being sent to the LLM for a grounded answer. An AI firewall layer sits in front of the model calls to screen for prompt injection and unsafe or out-of-scope requests. A frontend client provides a chat-style interface on top of the API, and the whole stack is containerized for deployment.
Key features
- test Hybrid retrieval combining BM25 keyword search with vector similarity search
- Reranking step to improve the relevance of retrieved context before generation
- Multi-tenant data isolation enforced at the access layer
- Role-based access control (RBAC) for users within each tenant
- Google Drive ingestion pipeline for real-world document sources
- AI firewall to screen for prompt injection and unsafe queries
- Web frontend for querying and reviewing grounded answers
- Containerized for deployment
04
Technologies & Deployment
- RAG
- AI Security
- FastAPI
- Vector Search
A working end-to-end RAG pipeline: ingestion, hybrid retrieval, reranking, and grounded generation Tenant and role isolation enforced consistently across the API A deployable, containerized system with a functioning frontend client