How I Built an Evaluation Harness for My RAG Chatbot
AIA walkthrough of puppy-evals, an evaluation harness for the RAG chatbot at windsordevelopmentstudio.io. Covers the three-judge architecture, a golden set of 60 test questions, and the results of two experiments that improved retrieval from 86.7% to 100% and persona consistency scores by 0.66 points.
5 min read


