AI App Development Day 4: Building a Minimal RAG Retriever from Keyword Matching to TopK
This technical walkthrough, adapted from Juejin developer community, demonstrates the evolution path of the Retrieve phase in RAG (Retrieval-Augmented Generation). The core contribution is implementing weighted keyword scoring and TopK ranking to overcome the semantic limitations of naive contains() string matching.
Core Technical Evolution
RAG consists of three stages: Retrieve, Augment, and Generate. This article focuses on practical implementation upgrades in retrieval:
- Starting point: Direct
contains()keyword matching in Java - Intermediate step: Decoupling knowledge base from retrieval logic with dictionary mapping
- Critical upgrade: Introducing
KeywordScoreweighted scoring—multiple matching synonyms increment the score - Final architecture: Iterate library→calculate relevance→sort descending→truncate TopK
Key insight: Keyword hits ≠ semantic similarity. For instance, the question “Why is this in-memory database so fast?” fails contains("Redis") since the word “Redis” does not appear. This exposes the fundamental limitation of pure string matching.
From Single to Multi-Knowledge召回 with TopK
Three practical implementation challenges arise:
Structured knowledge storage: Use
Map<String, List<String>> semanticMapto map entities to related terms. Example: Redis → [“redis”, “cache”, “in-memory database”, “high performance”, “key-value”]Multi-knowledge simultaneous召回: Questions like “What is the difference between Redis and MySQL?” require matching both entities, not stopping at the first hit
TopK ranking mechanism: Set constant
TOP_K = 3to limit returned items, preventing Prompt context overflow in production
Test results show: For “Redis and MySQL有什么区别?”, MySQL scores 2 (hits [“mysql”, “sql”]), Redis scores 1 (hits [“redis”]), returned in descending order.
Architecture Design
The complete retrieval process involves five steps:
- Convert user input to lowercase (e.g., “Redis为什么这么快?” → “redis为什么这么快?”)
- Iterate semantic dictionaries per knowledge item, counting matches as raw scores
- Wrap results in
KnowledgeScoreDTO containing text and score - Sort descending by score, truncate at TopK
- Extract knowledge texts for Prompt injection under “Known context”
Notable Bug case: contains("sql") falsely matches “mysql” due to substring containment, illustrating spurious correlation in keyword-based schemes.
Practical Recommendations
- Ready to implement now: Small-to-medium knowledge bases (< 1000 items), semantically stable vertical domains (product docs, FAQ systems)
- Wait longer: Multilingual support, broad synonym expansion, or long-context dependency scenarios should adopt Embedding + vector database approaches
Final Thoughts
This keyword-based retrieval is essentially a simplified version of “relevance scoring.” Its architectural skeleton—“candidates→score→rank→truncate”—remains identical to vector检索 solutions. The distinction lies solely in score computation: replacing simple string-hit counting with semantic vector cosine similarity. This provides a natural bridge to learning Embedding techniques next.
