Your RAG "feels accurate" in the demo and is quietly wrong at scale. Evaluate it as two systems: retrieval metrics + faithfulness/groundedness, in a golden-set harness in CI.