No Model Passed the Bake-Off. I Shipped One Anyway.
Choosing a groundedness metric when no leaderboard covers your corpus, every candidate fails your gates, and the bias you ship has to be documented instead of hidden.
Choosing a groundedness metric when no leaderboard covers your corpus, every candidate fails your gates, and the bias you ship has to be documented instead of hidden.
I was two days from rewriting a working system to fix a deficit that did not exist. Cosine similarity was scoring vocabulary reuse, not truth.
A mocked unit test passed on every commit while the live Cohere call failed on every run. What that gap cost, and the three-layer methodology that closes it.
One day of evaluation produced three null results. Every one was a measurement artifact, and each nearly became a documented fact.
The same vector store comparison gave opposite answers two weeks apart. The decision wasn't really about the vectors. It was about how long the system needed to live.
A grid search across 15 RAG configs revealed that chunk size matters more than embedding model, overlap is not optional, and bigger parameters don't mean better recall.
Grid search across 16 RAG configurations reveals embedding model selection drives 26% more retrieval quality than chunk tuning.
Why I am building 9 AI systems from scratch while working full-time as an Engineering Manager. The portfolio, the progression, and what I have learned so far.