Architecture
The problem
RAG is easy to demo and hard to trust. Retrieval quality drifts, prompts regress, and there's rarely a way to tell whether a change made the system better or worse. Teams ship on vibes because the measurement layer doesn't exist.
What I built
- An ingestion pipeline that chunks, embeds and indexes documents into PostgreSQL/pgvector.
- Multiple retrieval strategies behind one interface, so they can be swapped and compared.
- A chat UI for interactive use, backed by a FastAPI service.
- An offline evaluation harness computing Hit@K and MRR over labelled sets, plus LLM-judge scoring for answer quality.
- An observability layer (structured logging, metrics, health checks) surfacing what the system actually did in production.
- Azure CI/CD so evaluation runs are part of the deploy path, not a manual chore.
Why it matters
It turns RAG from a black box into something you can reason about. Every retrieval and generation decision is measurable and observable, so changes are validated against numbers instead of intuition.
System notes
- Retrieval strategies share a common contract, which keeps the eval harness strategy-agnostic.
- Offline eval and live observability use the same logging schema, so production traffic can feed back into evaluation sets.
- pgvector keeps retrieval and application state in one database, simplifying the operational surface.
Key decisions
- PostgreSQL/pgvector over a dedicated vector DB
- One datastore for vectors and relational state cuts operational complexity for a single-owner system without sacrificing retrieval quality at this scale.
- LLM-judge alongside ranking metrics
- Hit@K and MRR measure retrieval; they don't measure whether the final answer is good. The judge closes that gap.
- Evaluation in CI/CD
- Putting metrics in the deploy path makes regressions visible before they reach users.