Got tired of trusting vibes on whether a retrieval change helped or hurt, so I built a small harness: 40 hand-written questions against our internal docs, context precision and recall scored per run, diffed against the last commit. Not Ragas, not fancy, just enough to stop shipping regressions blind. Already caught one bad re-ranker change before it went out.
Related in this Cluster
- BuildsRunSarcasm GPS Tracker
A custom hardware/software project building an aggressively sarcastic GPS tracker for marathon training. Uses ESP32 and an e-ink…
- BuildsMigrating this site to a block theme, from scratch
Rebuilt koushik.site as a minimal custom FSE block theme instead of a page builder. Five categories, one query…
- BuildsTraining data dashboard, mostly to stop checking four apps
Pulled GPS, pace, and heart rate data into one local dashboard instead of checking Garmin Connect, Strava, and…
