Your Vector Search Got Faster and Recall Fell Off a Cliff — Nobody Got Paged
The latency graph looked like a win. Someone had tuned the vector index over the weekend, and p99 for semantic search dropped from around 200ms to 25ms. The dashboard was green, the change shipped, everyone moved on. Two weeks later a different graph started climbing: support tickets. Search "felt dumb." The RAG assistant kept missing answers that were obviously in the knowledge base — type in almost the exact wording of a document and it still wouldn't surface. Nothing had errored. Nothing had paged. The index was returning ten results for every query, fast, every time. They were just increasingly the wrong ten. Why approximate search fails quietly Any production vector search runs on an approximate nearest neighbor (ANN) index, because exact nearest-neighbor over millions of high-dimensional embeddings means scanning every row. ANN indexes buy their speed by not looking at most of your data on each query — and every one of them exposes a knob that controls exact...