Your RAG Pipeline Is Vector-Only. Hybrid Search and a Reranker Are the Cheapest Quality Win You Are Skipping.
Most marketing RAG systems fail on the same query. Someone asks for the "SOC 2 Type II" one-pager, or the pricing sheet for a SKU like "PRO-24-ANNUAL", and the assistant confidently returns a general security overview or the wrong plan. The model is fine. The embedding model is fine. The problem is that you built retrieval on dense vectors alone, and dense vectors are bad at exact strings.
If your team has already fixed chunking and picked a vector database, this is the next lever, and it costs far less than another model upgrade. Run keyword search and vector search together, fuse the results, and rerank the top candidates before the LLM sees anything. That is hybrid search plus a reranker, and it should be the default architecture for any marketing knowledge base.
Why Vector Search Alone Misses the Queries That Matter
Dense embeddings map meaning into a shared space. They are excellent at paraphrase: "how do we cut churn" finds a document titled "retention playbook." They are weak at anything where the literal token is the point. Product codes, competitor names, campaign IDs, regulation names, error strings, and acronyms all get smeared into their semantic neighborhood, so a near neighbor can outrank the exact match.
Keyword search with BM25 has the opposite profile. It nails the literal match and fails on vocabulary mismatch. Marketing content is full of both failure modes at once: your sales team searches by SKU and contract term, while your content team searches by concept. A single retriever cannot serve both, so stop asking it to.
Three Retrieval Setups, Compared
| Setup | Strong At | Fails At | Best For |
|---|---|---|---|
| Vector only | Paraphrase, concept matching, vague questions | Exact codes, names, acronyms, rare terms | Small corpora of long-form prose |
| BM25 only | Literal matches, SKUs, legal and product terms | Synonyms, reworded questions | Catalogs and structured references |
| Hybrid plus reranker | Both, with a final relevance check | Adds a pipeline stage and some latency | Any mixed marketing knowledge base |
How to Build It Without Rewriting Your Stack
The build has three parts, and none of them requires replacing what you have.
First, run both retrievers in parallel against the same corpus. Most platforms already ship both: Elasticsearch and OpenSearch, Weaviate, Qdrant, and Postgres with pgvector alongside its built-in full-text search. If your vector store cannot do keyword search, put a lexical index next to it rather than migrating.
Second, fuse the two ranked lists with Reciprocal Rank Fusion. RRF scores each document by the sum of 1 divided by (k plus its rank) across the lists, with k commonly set to 60. It uses ranks instead of raw scores, which matters because BM25 scores and cosine similarities live on incompatible scales. You do not have to normalize anything or tune weights on day one. That is why RRF is the right starting point.
Third, pull the top 30 to 50 fused candidates and pass them through a cross-encoder reranker. Unlike an embedding model, which scores query and document separately, a cross-encoder reads them together and judges actual relevance. Hosted options include Cohere Rerank and Voyage's rerankers, and open-weight options like the BGE rerankers run on your own infrastructure. Keep the top 5 to 8 for the prompt. Fewer, better chunks also means a smaller context bill and less noise for the model to wander through.
Measure It Before You Ship It
The reranker adds a network hop and inference time, so do not adopt it on faith. Build a small evaluation set from real questions your team asks: 50 to 100 queries, each labeled with the document that should come back. Then compare three configurations on recall at 5 and on p95 latency: vector only, hybrid, and hybrid plus reranker.
Expect the biggest jump on your exact-match queries, the SKU and acronym lookups, and a smaller one on conceptual queries. If hybrid alone gets you most of the gain, ship that and hold the reranker for the query classes that still miss. If you already run evals on your marketing AI, this slots into the same harness. If you do not, this is the cheapest place to start, because retrieval failures are visible and countable in a way that answer quality is not.
Also log every retrieval in production: the query, the returned chunk IDs, and which retriever surfaced each one. When a wrong answer reaches a customer, you want to know in one lookup whether retrieval or generation caused it.
Hybrid Retrieval Rollout Checklist
- Collect 50 to 100 real queries from sales, support, and content teams, each labeled with the correct source document
- Confirm your platform supports keyword and vector search on the same index, or add a lexical index beside your vector store
- Fuse results with Reciprocal Rank Fusion at k equal to 60 before touching any weights
- Rerank the top 30 to 50 candidates and send only the top 5 to 8 chunks to the model
- Compare recall at 5 and p95 latency across vector only, hybrid, and hybrid plus reranker
- Log query, chunk IDs, and source retriever for every production call so bad answers can be traced
The Takeaway
Swapping in a bigger model will not fix a retriever that cannot find the SKU. Hybrid search and reranking fix the retriever, they use tools you likely already pay for, and they can be tested in a week. Do that before the next model migration, before the next fine-tuning discussion, and before you blame the LLM for an answer it never had the right document to give.
Tags
LETSGROW Dev Team
Marketing Technology Experts
Ready to Apply This Insight?
Schedule a strategy call to map these ideas to your architecture, data, and operating model.
Schedule Strategy Call