Search, explained: four retrieval arms, one deterministic ranking
Ask most memory systems how they retrieve and the answer is one word: embeddings. Vector similarity is a good signal, but it is one signal. It misses exact identifiers, misses structure, and has no idea what "last sprint" means. Search in Illumina treats retrieval as a fusion problem instead.
Four arms, one query
Every search fans out across four retrieval arms in parallel:
- Semantic. pgvector HNSW search over 1024-dimension Cohere Embed v4 embeddings. Catches paraphrases and conceptual matches.
- Lexical. BM25 over native PostgreSQL tsvector. Catches the things embeddings fumble: ticket numbers, function names, exact phrases.
- Graph. Link expansion through the knowledge graph, pulling in facts connected to the candidates by entity and causal edges.
- Temporal. Joins only when the query carries temporal intent. "What changed after the March incident" activates it; "how does auth work" does not.
Each arm returns a ranked candidate list. Reciprocal Rank Fusion merges them with k=60 and a locked arm order, so a fact that ranks well across several arms beats a fact that ranks first in one and nowhere else.
When the query reads as thematic rather than specific ("what have we learned about onboarding", not "what is the retry timeout") search also retrieves over community summaries: the report Illumina writes for each cluster it finds in the entity graph. Ten near-duplicate snippets are the wrong answer to a question about a theme, and the community summaries are what turn that into one.
Rerank, then score
The fused candidates go through a reranker, Cohere Rerank 3.5, which reads the query and the candidate together rather than comparing pre-computed vectors. That is where most of the precision comes from.
The final score multiplies normalized relevance by three boosts: recency, temporal proximity, and corroboration. Corroboration matters more than it sounds. A fact backed by more independent sources ranks higher than a one-off mention. Each boost is designed so that a missing signal is neutral, never distorting: a fact with no temporal data is not penalized on a query with no temporal intent. Standing decisions never decay below neutral recency, so a policy set two years ago still ranks like it was said yesterday.
Deterministic, on purpose
Run the same search twice and you get byte-identical output. Final ordering uses a total-order float comparison with a stable id tie-break, so there is no dependence on insertion order, parallel scheduling, or floating-point drift between runs.
If you are building agents on top of a memory system, this is not a nicety. Deterministic retrieval means reproducible agent behavior, debuggable failures, and evals that measure your changes instead of retrieval noise.
{
"query": "why did we move ingestion off the shared queue",
"budget": "mid",
"trace": true
}
Set budget to low, mid, or high to trade latency for depth, and max_tokens to cap the result set so it drops straight into a prompt. Pass trace: true and the response carries the scoring detail for every result.
Failure is a mode, not an outage
Retrieval infrastructure fails in pieces. If the embedding service is down, search does not fail with it. It skips the semantic and temporal arms, answers from BM25 plus graph expansion off the top lexical hits, and flags the response as degraded. If the reranker is down instead, the fused RRF ordering stands in with neutral scores. Either way your agent keeps working, and it knows the answer came from a reduced pipeline.
Every arm runs inside the same PostgreSQL database. No separate vector store, no graph engine, no fan-out across services that can half-fail. One query planner, one transaction boundary, one thing to operate.
Search is available through the REST API, the Python and TypeScript SDKs, and as an MCP tool your agent can call directly.