RAG in practice, or why vector search alone isn't enough
Created: June 12, 2025 Updated: Sept. 29, 2026
I'm building a chatbot for medical device support. It worked great in the demo, less so with real questions.
The idea was simple, someone asks about a specific device, the bot looks through the documentation and answers. Classic RAG, we split documents into chunks, compute embeddings, put them into Azure AI Search and, when a question comes in, pull the closest chunks into the prompt.
There are several thousand devices and many of them have almost identical names, differing by one letter or a version number, and to vector search that's almost the same thing. The bot could answer very confidently, only based on the manual for a different model, and a mixed-up manual for a medical device is a serious matter.
The fix turned out to be rather down to earth, before we even start searching the documentation, we first work out which device it's about. Exact match on code and name first, then fuzzy matching, embeddings only at the end, and if the product can't be identified unambiguously, the bot asks instead of guessing.
def resolve_product(question: str) -> Product | None: candidates = exact_match(question) or fuzzy_match(question) if not candidates: candidates = vector_match(question, top_k=5) if len(candidates) == 1: return candidates[0] return None # the bot asks which device the user means
Only once the product is known do we search for chunks, and only in its documentation. Noticeably fewer mistakes straight away, and the answers are shorter and more to the point.
Then came tools the model chooses from itself: step-by-step troubleshooting, documentation search, looking things up in the product catalogue system and in the system with issues reported by users, data from a Postgres database, all on LangChain and LangGraph. It works, but every new tool is one more place where the model can take a wrong turn.
Questions come in many languages while the documentation is mostly in English. Embeddings cope with the language itself, product names less so, because those have to match exactly whatever language the question is in, which is one more reason to keep product recognition as a separate step.
What I'd do from day one
A test set from the very first day, real questions with the expected device and document, run in CI on every change to the prompt or the chunking. Without it every fix is a bit of a guess whether we've just broken something else.
Machine-translated from Polish (original).