Retrieval
Retrieval is the process of finding the most relevant information from enterprise knowledge and providing it to an AI system during execution.
This is what allows AI agents to answer questions using your organization’s data instead of relying solely on model training data.
Without retrieval:
User Question ↓LLM ↓Answer Based On Training DataWith retrieval:
User Question ↓Semantic Search ↓Relevant Enterprise Knowledge ↓LLM ↓Grounded ResponseThis is the foundation of Retrieval-Augmented Generation (RAG).
Why Retrieval Matters
Section titled “Why Retrieval Matters”Consider the question:
My VPN access expired. Can you restore it?
The language model itself does not know:
- Your company’s VPN policy
- Approval requirements
- Security procedures
- Employee access rules
Without retrieval, the model will guess.
With retrieval, the model receives relevant company context before generating a response.
The Retrieval Pipeline
Section titled “The Retrieval Pipeline”When an agent receives a request, Karyam executes the following process:
User Request ↓Query Embedding ↓Vector Similarity Search ↓Relevant Chunks Retrieved ↓Context Added To Prompt ↓Model Generates ResponseThis entire process typically happens in milliseconds.
Step 1 — User Request
Section titled “Step 1 — User Request”The retrieval process begins with a question.
Example:
Can you restore my VPN access?
Step 2 — Generate Query Embedding
Section titled “Step 2 — Generate Query Embedding”The configured embedding model converts the request into a vector.
Example:
Can you restore my VPN access?
↓
[0.182, -0.734, 0.219, ...]This vector represents the semantic meaning of the request.
Step 3 — Search The Vector Database
Section titled “Step 3 — Search The Vector Database”The query embedding is compared against all stored embeddings.
Example:
Vector Database│├── VPN credentials require manager approval│├── Password resets are automated│├── Contractors require security review│└── Employee handbook introductionThe retrieval engine identifies which chunks are most semantically similar.
Step 4 — Retrieve Relevant Chunks
Section titled “Step 4 — Retrieve Relevant Chunks”Example retrieval result:
VPN credentials expire every 90 days.
VPN renewals require manager approval.
Contractor renewals require additional security review.Only the most relevant chunks are selected.
Step 5 — Augment The Prompt
Section titled “Step 5 — Augment The Prompt”The retrieved context is automatically injected into the model prompt.
Example:
User Question:Can you restore my VPN access?
Retrieved Context:VPN credentials expire every 90 days and require manager approval for renewal.
System Instructions:You are an IT Support Assistant.This process is called prompt augmentation.
Step 6 — Generate Response
Section titled “Step 6 — Generate Response”The model now responds using company knowledge rather than assumptions.
Example:
Your VPN access appears to require manager approval before renewal. I’ve initiated the approval workflow and notified your manager.
Similarity Search
Section titled “Similarity Search”Retrieval uses vector similarity instead of exact keyword matching.
Traditional search:
VPN access expired≠VPN credentials renewalSemantic search:
VPN access expired≈VPN credentials renewalThis allows AI systems to understand meaning rather than wording.
Top-K Retrieval
Section titled “Top-K Retrieval”Most retrieval systems do not return every matching chunk.
Instead, they return the most relevant results.
Example:
Top 5 ResultsTop 10 ResultsTop 20 ResultsReturning too much information can reduce answer quality.
The goal is to provide:
- Enough context
- Minimal noise
- Low token usage
Metadata Filtering
Section titled “Metadata Filtering”Retrieval can be constrained using metadata.
Examples:
Search only:
- HR policies
- Finance documents
- IT documentation
- Workspace-specific content
Example:
Workspace = EngineeringCategory = SecuritySource = VPN PoliciesThis improves precision and governance.
Retrieval In Karyam
Section titled “Retrieval In Karyam”Karyam retrieval combines:
Embedding Model +Vector Database +Metadata Filters +AI AgentThis architecture allows retrieval pipelines to scale across multiple business domains.
Retrieval During Agent Execution
Section titled “Retrieval During Agent Execution”Example:
Employee:"My VPN access expired."
Agent:↓Generate Query Embedding↓Search Vector Database↓Retrieve VPN Policy↓Check HR Status↓Trigger Approval Flow↓Create TicketRetrieval becomes one step within a larger AI system.
Observability
Section titled “Observability”Retrieval operations are fully observable.
You can inspect:
- Retrieved chunks
- Similarity scores
- Source documents
- Token usage
- Retrieval latency
This makes production RAG systems debuggable.
Example Retrieval Trace
Section titled “Example Retrieval Trace”10:01 Query Received10:01 Query Embedding Generated10:01 Vector Search Completed10:01 5 Chunks Retrieved10:01 Context Added To Prompt10:02 Response GeneratedCommon Retrieval Problems
Section titled “Common Retrieval Problems”Poor Chunking
Section titled “Poor Chunking”Large chunks reduce retrieval precision.
Weak Embedding Models
Section titled “Weak Embedding Models”Low-quality embeddings reduce semantic understanding.
Missing Metadata
Section titled “Missing Metadata”Without metadata filtering, irrelevant information may be retrieved.
Insufficient Context
Section titled “Insufficient Context”Retrieving too few chunks may omit important information.
Best Practices
Section titled “Best Practices”Keep Knowledge Organized
Section titled “Keep Knowledge Organized”Use folders and listeners to separate domains:
/hr/finance/legal/engineeringUse Metadata
Section titled “Use Metadata”Metadata improves both retrieval quality and observability.
Monitor Retrieval Performance
Section titled “Monitor Retrieval Performance”Use RAG Runs to investigate:
- Retrieval latency
- Missing results
- Incorrect matches
- Ingestion failures
The Karyam Model
Section titled “The Karyam Model”User Question ↓Embedding Model ↓Vector Search ↓Retrieved Context ↓AI Agent ↓ResponseRetrieval is what transforms AI from a chatbot into an enterprise system that understands your organization.
Next Steps
Section titled “Next Steps”Continue with:
➡️ RAG Runs
Learn how Karyam provides complete observability into ingestion and retrieval pipelines.
