Skip to content
All work

RAG Eval & Observability

A full-stack platform for building and evaluating RAG systems, with the eval and observability rails built in.

Next.jsFastAPIPostgreSQL / pgvectorAzure

Architecture

DocumentsChunk + embedpgvectorRetrieveLLM answerEVALUATIONHit@K · MRRretrieval metricsLLM-judgeanswer qualityOBSERVABILITYLoggingMetricsHealth↺ runs in Azure CI/CD

The problem

RAG is easy to demo and hard to trust. Retrieval quality drifts, prompts regress, and there's rarely a way to tell whether a change made the system better or worse. Teams ship on vibes because the measurement layer doesn't exist.

What I built

  • An ingestion pipeline that chunks, embeds and indexes documents into PostgreSQL/pgvector.
  • Multiple retrieval strategies behind one interface, so they can be swapped and compared.
  • A chat UI for interactive use, backed by a FastAPI service.
  • An offline evaluation harness computing Hit@K and MRR over labelled sets, plus LLM-judge scoring for answer quality.
  • An observability layer (structured logging, metrics, health checks) surfacing what the system actually did in production.
  • Azure CI/CD so evaluation runs are part of the deploy path, not a manual chore.

Why it matters

It turns RAG from a black box into something you can reason about. Every retrieval and generation decision is measurable and observable, so changes are validated against numbers instead of intuition.

System notes

  • Retrieval strategies share a common contract, which keeps the eval harness strategy-agnostic.
  • Offline eval and live observability use the same logging schema, so production traffic can feed back into evaluation sets.
  • pgvector keeps retrieval and application state in one database, simplifying the operational surface.

Key decisions

PostgreSQL/pgvector over a dedicated vector DB
One datastore for vectors and relational state cuts operational complexity for a single-owner system without sacrificing retrieval quality at this scale.
LLM-judge alongside ranking metrics
Hit@K and MRR measure retrieval; they don't measure whether the final answer is good. The judge closes that gap.
Evaluation in CI/CD
Putting metrics in the deploy path makes regressions visible before they reach users.