All posts

Semantic search was off in production. Keyword search made it look healthy.

Semantic retrieval was broken in production, but keyword fallback kept answers flowing and hid the failure. We fixed the credential, backfilled current embeddings, and made keyword-only mode visible instead of treating any answer as proof of retrieval health.

September 30, 20265 min readSaytu team

On this page

Semantic retrieval had been off in production for as long as it had existed, and the application still looked like it worked.

Shoppers could ask questions. The assistant still returned answers. Knowledge search still produced results. There was no outage page and no obvious exception pointing at retrieval.

The system had a fallback, and the fallback was doing exactly what it was built to do. Keyword search kept answering. That made the failure harder to see than a crash.

The embedding key was not a key

We found the problem while running an embeddings backfill, not while investigating a shopper complaint. Our environment file set the embeddings credential by referring to the OpenRouter credential already used elsewhere. Next.js expands that kind of reference when it loads an environment file. Docker Compose's env_file does not.

Inside the container, the value therefore arrived as the literal variable reference rather than the credential it referred to. The embeddings endpoint rejected it.

The application did not fall over because semantic retrieval had been designed to fail soft. If embeddings were unavailable, keyword retrieval and the rest of the search path stayed behind it. That behaviour was intentional. A temporary embeddings-provider problem should not make a shop assistant stop answering altogether.

The unexpected consequence was that a permanent configuration failure looked almost identical to a working system with weaker retrieval.

Graceful degradation can hide permanent degradation

Fallbacks are usually discussed as reliability features. That is what this one was. No embeddings key, a provider problem or knowledge indexed before semantic retrieval existed should not turn every shopper request into an error. The system should still do the best it can with the paths that remain.

The problem is that a fallback answers a different question from observability. A fallback asks: can we still return something useful? Observability asks: is the system running the way we think it is?

We had answered the first question and not the second. Because keyword search continued to produce plausible results, the semantic lane could disappear without producing the kind of symptom engineers normally associate with a broken dependency. There was no blank response to trace back to the embeddings call. There was simply a keyword result in a place where a semantic result might have been better.

From the outside, that looks like retrieval quality. From the inside, an entire retrieval path is missing.

The model was never going to reveal the problem

This class of failure is easy to misdiagnose as a model-quality problem. A shopper asks something the system ought to know. The answer is vague, misses the better source or says it cannot find enough information. The visible component is the AI reply, so the natural place to investigate is the model or its prompt.

But the model can only work with the context it receives. If semantic retrieval silently stops contributing candidates, rewriting the prompt does not restore those candidates. Moving to a stronger model does not restore them either. The model may become more articulate about incomplete context, but the missing retrieval lane remains missing.

That was the useful part of finding this through the backfill. The embeddings endpoint failed directly there, without the successful keyword path standing in front of it. A problem that looked like search quality in the full product became a configuration error when one layer was exercised on its own.

Fixing the credential was only half the repair

Resolving the environment reference restored the ability to create and query embeddings, but existing vectors still had to be treated carefully. A stored vector only belongs to the embedding model that created it. Rows written before vectors existed, or rows carrying vectors from another embedding model, cannot simply be assumed to participate correctly in the current semantic path.

So the same work added a backfill that rewrites the knowledge that needs current embeddings while skipping rows that are already current. The operation can be run again without treating every existing row as new work.

The larger point was that enabling a semantic path is not just a runtime configuration change. There is stored state behind it. The configuration that asks the question and the vectors that answer it have to agree. Otherwise a green provider call can still sit in front of an index that is not actually usable by that provider.

We stopped treating fallback as proof of health

The original bug taught us that a successful answer was too high-level a health signal for retrieval. Later, we made the missing semantic lane visible in the retrieval console. A missing resolvable key is now a headline condition rather than something inferred from the quality of the final answer. A stored embedding model that no longer matches the configured one is surfaced too, because that condition also demotes retrieval to keywords.

That distinction matters operationally. Keyword fallback is still useful. We did not remove it just to make failures louder. A shopper should not lose the whole assistant because one retrieval method is unavailable.

What changed was our interpretation of success. "An answer was returned" means the product survived the request. It does not mean every retrieval lane ran.

A healthy fallback needs an unhealthy signal

The uncomfortable part of graceful degradation is that the better it works, the easier it is to forget that something degraded. If the fallback returns nothing, the failure is obvious. If it returns a plausible answer, the system can spend a long time operating below the quality level everyone believes it has.

That is especially dangerous in retrieval systems because quality failures rarely announce their cause. A missing semantic match, a stale vector and genuinely thin knowledge can all produce the same shopper-facing sentence: the assistant did not find the answer. Those states should not look identical to the people operating the system.

Our keyword path did its job. It kept the assistant useful while semantic search was unavailable. The mistake was treating that usefulness as evidence that semantic search was healthy.

We now want both properties at once: a fallback that protects the shopper and a signal that tells us the fallback had to take over. Production AI needs graceful degradation. It also needs degradation that refuses to stay invisible.

Keep reading

All posts