SecureAgentRAG
Privacy-first English/Arabic agentic RAG where access control, citations and privacy are enforced by the system, not requested in a prompt.

696 test functions (718 collected), plus live-Qdrant integration tests in CI.
Every significant call is recorded as an ADR in the repository.
MS-MARCO hold-out, with a small in-domain gain too. Shipped opt-in, because a small gain didn't justify a ~2 GB dependency for every user.
Bring-your-own-key public demo on free tiers. The local-only privacy guarantee applies when self-hosted, and the demo says so.
Problem
Company document assistants fail in two dangerous ways. They show people documents they shouldn't see, and they answer confidently when the documents don't support the answer.
Approach
A nine-node LangGraph pipeline. A router and guardrails come first. Retrieval fuses BGE-M3 dense vectors with Qdrant-native sparse vectors by reciprocal rank, and a cross-encoder reranks the result. A grader decides whether the context is good enough, and if it isn't, a rewriter tries again (corrective retrieval) before the system admits it found nothing.
Access control lives in the retrieval filter itself. Organization, clearance level and roles are applied to dense and sparse search under one filter, so a cross-tenant read is structurally blocked rather than checked afterwards. A sensitivity router keeps high-sensitivity work on a local model when self-hosted, and every operation lands in a SHA-256 hash-chained audit log.
Challenges
Being honest about the demo. A free public demo has no local GPU, so the "high-sensitivity stays local" promise can't hold there. The demo surfaces a sensitivity badge and documents the trade-off instead of silently breaking the guarantee. The fine-tuned reranker raised NDCG@10 from 0.774 to 0.790, real but small, so it ships as an opt-in checkpoint rather than a default that costs every user a ~2 GB dependency.
Outcome
Open source, with about 700 automated tests, 41 ADRs, Prometheus/Grafana observability for self-hosted runs and a bring-your-own-key public demo. Arabic questions retrieve, cite and answer end to end over an illustrative Egyptian corpus.
What I'd change
The generation side has stronger evaluation than retrieval. Next is a retrieval benchmark on in-domain questions, so reranker and chunking changes are judged by numbers rather than spot checks.
