A Local Hybrid Retrieval Pattern for Health Records
Vector search is fantastic when it comes to discovering related language, but it struggles with specific identifiers, rare drug names, or clinical codes. On the other hand, full-text search excels at literal matches, yet it might overlook paraphrases. To retrieve health records effectively, we need a balanced approach that combines both methods—while understanding that their scoring systems are not interchangeable.
Part 5 focuses on developing the retrieval component of an on-premises RAG (Retrieval-Augmented Generation) system. This includes features like conversation-aware rewrites, local BGE-M3 embeddings, simultaneous vector and lexical searches in DocumentDB, reciprocal rank fusion implemented in Rust, row-level deduplication, adaptive cross-encoder reranking, a calibrated operational gate, and accurate citations integrated into local generation.
Retrieval is an area where unintentional cloud dependencies can sneak in. Embedding and reranking processes generally rely on APIs, and if a hosted reranker accesses retrieved text passages, this can be problematic, especially since these passages often contain sensitive information. Certain countries, like Kenya—which is where this research originated—mandate health data processing to occur within their borders, requiring stringent data safeguards. Thus, to adhere to these regulations, the BGE-M3 and cross-encoder reranking systems are localised and operate on the same machine where generation occurs. The principles guiding this setup are explored in Part 1.
This prototype utilises synthetic records. It’s important to note this does not substitute for clinical advice, compliance certifications, or validate score thresholds for other data sets.
While Rust is the chosen language for implementation in this series, you can achieve similar results using the Foundry Local SDK in C#, JavaScript, or Python, with links available in Part 2.
Neither vector search nor full-text search on their own suffice for satisfying health record inquiries. Dense retrieval can overlook precise codes or uncommon drug names, while lexical retrieval might miss paraphrased terms conveying the same idea. Opting for just one approach leads to predictable gaps.
Combining the two in a simplistic way can introduce new issues. For instance, scores based on cosine similarity and text relevance aren’t directly comparable. Overlapping query variants can skew record rewards, and if weak candidates are fed into the generation process, it might produce results that appear confident but are actually misleading.
The retrieval framework must maintain both exact and semantic recognitions while refraining from establishing a common scoring scale. It needs to merge rankings responsibly, eliminate duplicates at the row level, rerank a carefully selected set of candidates locally, and impose an operational relevance gate before any evidence reaches Foundry Local. Most critically, all processing must occur without transferring data outside the facility, as any segments sent to an offsite reranker would carry strict residency obligations.
- Both exact terms and their paraphrases should have their own retrieval paths.
- Vector and lexical results should be merged based on rank, not raw-score calculations.
- Duplicate expressions and overlapping sections must not inflate evidence from a single row.
- A local cross-encoder must rerank a diverse set of candidates.
- Weak evidence should result in no relevant records instead of proceeding to generation.
- Accepted passages must keep consistent citations relating to their original records.
PostgreSQL, MySQL, and SQL Server are classified as External Clinical Databases. They remain operational systems where records are maintained. Selected and permitted fields are copied into a separate Internal DocumentDB Hybrid Store, which consists of documents, vectors, metadata, and application states. Retrieval actions occur specifically on these copied versions, excluding vector searches against the original databases.
During generative phases—such as rewriting and generating answers—Foundry Local comes into play. Meanwhile, the embedding and reranking operations are handled by fastembed. It’s vital to differentiate this in operational terms: changing the chat model does not affect the contractual obligations concerning the vector index.
Questions like “What happened after that?” become inadequate without conversational context. The server maintains a limited, operational memory—consisting of a rolling summary and a chronological log—and instructs a local model to generate a standalone rewrite. When expansion might be helpful, the same local request produces a few variant responses; three is the typical default.
However, expansion is not always a beneficial step. Short queries, quoted phrases, and identifier-type questions bypass it, as rephrasing could eliminate critical details. Should the preparation phase falter or yield inconsistent results, retrieval defaults back to the original question rather than concluding the request as unsuccessful.
In a conceptual framework:
let (standalone, variants) = prepare_queries(
foundry,
&memory.rewrite_turns(),
question,
).await;
let queries = if should_expand(&standalone) {
unique_bounded(variants, 3)
} else {
vec![standalone.clone()]
};
Remember, the rewrite serves as an aid for recall, not as concrete evidence. While conversational memory clarifies pronouns, only records found through retrieval become citable.
Every variant of a query is locally embedded using fastembed BGE-M3. The current framework involves:
- 1,024 dimensions;
- the same prefix-free encoder for both records and queries;
- normalized query text used as a cache key; and
- a process-local LRU cache limited to 1,024 entries.
Batching becomes crucial; all cache misses for the variants are submitted in one embedding batch, staying within a single task boundary. Fastembed sessions work synchronously, running outside the Tokio async reactor and are protected by a process-level mutex, ensuring accuracy, though concurrent requests might stack behind the shared model.
Batch processing is essential since we want to keep the overhead of model setup minimal during expansion. Caching also plays a key role, as follow-up interactions and repeated queries often rely on normalised text. Neither the cache retains final answers nor prompts.
Hybrid mode creates two distinct searches for each query variant:
- Vector: DocumentDB cosmosSearch focuses on contentVector, applying an IVF cosine index.
- Lexical: DocumentDB $text searches over the designated text field.
Both operations filter for active: true, ensuring that incomplete ingest generations are not visible. Each variant-side operation runs at the same time and yields a ranked list.
for (query, vector) in queries.iter().zip(query_vectors) { searches.push(vector_search(db.clone(), vector, per_side));if mode == RetrievalMode::Hybrid { let lexical = enhance_exact_terms(query); searches.push(text_search(db.clone(), lexical, per_side)); }}
let ranked_lists = futures::future::join_all(searches).await;
The lexical enhancement features exact tokens from ICD-like terms and capitalised drug names. Importantly, this does not replace the user’s query but offers exact tokens another chance to rank effectively.
Introducing concurrency reduces the overall time for one request, but it increases the workload on DocumentDB. Consequently, a retrieval admission semaphore is included in the methodology, to prevent overloading as excessive parallel processing could lead to failures.
It’s important to note that cosine scores and full-text relevance scores do not represent probabilities and are not directly comparable. A weighted average such as 0.7 * cosine + 0.3 * textScore seems scientific, yet lacks validity without meticulous normalisation and calibration to the corpus.
Reciprocal rank fusion (RRF) steers clear of raw scores and instead combines ranks:
In this context, ri(d) represents the rank of document d in list i. The default k=60 counteracts spikes from single lists and rewards evidence appearing across various search modes or variants.
The Rust implementation has been designed to be concise:
pub fn fuse(rankings: &[Vec], k: f64) -> Vec<(String, f64)> { let mut scores = HashMap::::new(); for ranking in rankings { for (index, id) in ranking.iter().enumerate() { *scores.entry(id.clone()).or_default() += 1.0 / (k + index as f64 + 1.0); } } let mut fused: Vec<_> = scores.into_iter().collect(); fused.sort_by(|a, b| b.1.total_cmp(&a.1).then_with(|| a.0.cmp(&b.0))); fused }A stable ID tie-breaker makes repeated results predictable when the fused scores align. RRF operates in application code, as it efficiently uses cosmosSearch and $text within DocumentDB, though the tested path doesn’t support a native rank-fusion stage suitable for this architecture.
Ingestion divides long rows into overlapping chunks. If row-level deduplication is absent, multiple adjacent chunks from one record could dominate the most viable candidate slots. Therefore, only the strongest chunk per (source_id, table, row_pk) is kept before limiting to the rerank candidate set.
This sequencing is crucial. Deduplication after limiting to top_k could yield fewer than top_k unique rows when a broader range of candidates exists just beneath the threshold. A second protective deduplication occurs post reranking.
This architecture prioritises diversity over reconstructing the original data. The chosen chunk carries duplicated row attributes for citation use, but the current setup does not acquire sibling chunks, nor does it rebuild the entire row. This absence of parent expansion is an acknowledged limitation, not an assumed capability.
While RRF offers a commendable candidate ranking, it does not correlate the question and passage jointly. A local bge-reranker-v2-m3 cross-encoder evaluates the primary standalone question against a select group of candidates.
This prefix is adaptive, initially comprising at least max(top_k, 8) candidates and may cut off at a point where fused scores clearly drop, never going beyond the defined rerank maximum. This strategy mitigates costly pairwise evaluations when the fused list’s ranking is evident.
The reranker’s raw logits undergo a sigmoid transformation:
This process provides a stable operational range of (0, 1), but it does not transform the score into a calibrated probability of clinical relevance, an important distinction that helps guard against common documentation mistakes.
Following the reranking stage, less convincing elements can be trimmed if they fall below half of the predefined answer gate, though at least one passage must remain. The default final context cap is set to six passages, keeping in mind row and total approximate word constraints.
A local model should not attempt to produce a grounded answer unless retrieval yields sufficient evidence. For specific questions, the current default gate will decline when the top sigmoid-transformed reranker score is lower than 0.30. This threshold can be adjusted or disabled as needed.
On the other hand, broad inquiries—like “list recent records” or “provide an overview”—are handled differently. They may intentionally bypass reranking, relying instead on RRF scores at the top. When comparing these to a reranker threshold, one would inadvertently mix scales. Thus, broad requests are only rejected when retrieval yields zero results.
fn should_refuse(question: &str, top_score: Option, gate: Option) -> bool { match top_score { None => true, Some(_) if is_broad_question(question) => false, Some(score) => gate.is_some_and(|floor| score < floor), } }This gate serves as a mechanism for risk control, not a guarantee of accuracy. It requires calibration using a diverse set of answerable and unanswered questions relevant to the target corpus. Microsoft’s architectural guidance also stresses the need to assess retrieval and generation stages independently, as well as the overall performance from end to end.
Accepted passages are transformed into numbered context. The system’s prompt allows citations solely for these records, while conversational memory aids in continuity but is not considered evidence. Citations appear before answer tokens, and the successful assistant message retains the precise information referenced.
Foundry Local generates answers locally via its embedded Rust SDK, without relying on any cloud model as a fallback. Should the local model be unavailable, it will return an error response instead of broadening trust boundaries.
An optional verifier can run after streaming. Initially, it verifies citation indexes methodically, and then it may engage a local verification model to assess claims as supported, partial, or unsupported compared to the provided passages. Verification is turned off by default, adds latency post-visible answer, and defaults to “skipped” should it encounter issues. This does not retroactively secure generation.
When evaluating architecture, three gaps should be acknowledged:
Lack of general metadata filters: There’s currently no standard allow-listed source/table/patient/date filter applied to both vector and lexical searches. The outlined RetrievalFilter is proposed; until it is established and confirmed, claiming patient-safe prefiltering is inappropriate.
Absence of parent expansion: The system preserves only the strongest chunk of each row, yet does not reconstruct sibling chunks after ranking.
Approximate chunking: The 384/64 windows and prompt budgets rely on whitespace counts rather than the tokenizer of the embedding model.
The retrieval router can categorise hybrid cohort-plus-narrative questions but lacks the capacity to execute a structured cohort and inject its row keys into both retrieval avenues. Therefore, such requests revert to standard semantic retrieval.
Hybrid search enhances recall but doubles the search operations needed for each variant. Expansion amplifies this multiplier even further. It’s crucial to set limits on both the number of queries and per-side depth, and to ensure admission control is enforced.
Reranking enhances precision but requires a local model that is serialized. The adaptive prefix manages costs, while broad queries avoid the complexities of applying a targeted scorer where it’s not suitable.
Here are a few anti-patterns to dodge:
- Avoid comparing vector and text scores directly.
- Don’t deduplicate after applying the top_k limit.
- Beware of treating sigmoid outputs as calibrated probabilities.
- Don’t use conversational memory as citable evidence.
- Avoid generating outputs when retrieval is weak or absent.
- Don’t assume metadata filtering is in effect just because metadata is stored.
- Don’t presume “hybrid” entails a single atomic database operation.
Current implementations include: rewriting and three-query expansion; BGE-M3 for query batching and caching; simultaneous cosmosSearch and $text operations; Rust RRF with k=60; pre-truncation row deduplication; adaptive bge-reranker-v2-m3; sigmoid operational scores; a configurable default 0.30 gate for pointed questions; limited context; reliable citations; local grounded generation; and optional post-stream verification.
Future enhancements will focus on: calibrating corpus-specific gates; developing general metadata filters; enabling parent/sibling expansion; refining tokenizer-accurate chunking; and executing a comprehensive structured cohort-to-hybrid-search pipeline. The designs for scoped-retrieval and cohort features are outlined in plans/new but are yet to be put into practice. Validation relies on synthetic health data and does not confirm clinical safety or adherence to regulations.
- Utilize both dense and lexical retrieval to cover complementary failure modes.
- Combine rank positions using RRF rather than blending incompatible scores.
- Deduplicate source rows before truncating candidate lists.
- Allocate resources for cross-encoder work on an adaptive bounded prefix.
- Implement a score gate on an appropriate scale and calibrate it based on empirical data.
- Ensure only retrieved passages are citable; memory should be for continuity, not for evidence.
The code underpinning this series is open source: github.com/kevin-gatimu/onprem-health-rag. It’s a prototype built for synthetic data research, and not intended for clinical or compliance-ready applications—feedback and improvements are welcome.
Share this content:
Discover more from Qureshi
Subscribe to get the latest posts sent to your email.