How are you versioning evaluation datasets for retrieval-augmented generation apps?
Our RAG evaluation set changes as product documentation and expected answers evolve. We want to compare prompt, retriever, and model changes without quietly changing the benchmark underneath the experiment. How are you versioning the questions, reference context, grading rubric, and human review notes so results remain reproducible?
0
