Skip to content
Karyam

Retrieval

Retrieval is the process of finding the most relevant information from enterprise knowledge and providing it to an AI system during execution.

This is what allows AI agents to answer questions using your organization’s data instead of relying solely on model training data.

Without retrieval:

User Question
↓
LLM
↓
Answer Based On Training Data

With retrieval:

User Question
↓
Semantic Search
↓
Relevant Enterprise Knowledge
↓
LLM
↓
Grounded Response

This is the foundation of Retrieval-Augmented Generation (RAG).


Consider the question:

My VPN access expired. Can you restore it?

The language model itself does not know:

  • Your company’s VPN policy
  • Approval requirements
  • Security procedures
  • Employee access rules

Without retrieval, the model will guess.

With retrieval, the model receives relevant company context before generating a response.


When an agent receives a request, Karyam executes the following process:

User Request
↓
Query Embedding
↓
Vector Similarity Search
↓
Relevant Chunks Retrieved
↓
Context Added To Prompt
↓
Model Generates Response

This entire process typically happens in milliseconds.


The retrieval process begins with a question.

Example:

Can you restore my VPN access?


The configured embedding model converts the request into a vector.

Example:

Can you restore my VPN access?
↓
[0.182, -0.734, 0.219, ...]

This vector represents the semantic meaning of the request.


The query embedding is compared against all stored embeddings.

Example:

Vector Database
│
├── VPN credentials require manager approval
│
├── Password resets are automated
│
├── Contractors require security review
│
└── Employee handbook introduction

The retrieval engine identifies which chunks are most semantically similar.


Example retrieval result:

VPN credentials expire every 90 days.
VPN renewals require manager approval.
Contractor renewals require additional security review.

Only the most relevant chunks are selected.


The retrieved context is automatically injected into the model prompt.

Example:

User Question:
Can you restore my VPN access?
Retrieved Context:
VPN credentials expire every 90 days and require manager approval for renewal.
System Instructions:
You are an IT Support Assistant.

This process is called prompt augmentation.


The model now responds using company knowledge rather than assumptions.

Example:

Your VPN access appears to require manager approval before renewal. I’ve initiated the approval workflow and notified your manager.


Retrieval uses vector similarity instead of exact keyword matching.

Traditional search:

VPN access expired
≠
VPN credentials renewal

Semantic search:

VPN access expired
≈
VPN credentials renewal

This allows AI systems to understand meaning rather than wording.


Most retrieval systems do not return every matching chunk.

Instead, they return the most relevant results.

Example:

Top 5 Results
Top 10 Results
Top 20 Results

Returning too much information can reduce answer quality.

The goal is to provide:

  • Enough context
  • Minimal noise
  • Low token usage

Retrieval can be constrained using metadata.

Examples:

Search only:

  • HR policies
  • Finance documents
  • IT documentation
  • Workspace-specific content

Example:

Workspace = Engineering
Category = Security
Source = VPN Policies

This improves precision and governance.


Karyam retrieval combines:

Embedding Model
+
Vector Database
+
Metadata Filters
+
AI Agent

This architecture allows retrieval pipelines to scale across multiple business domains.


Example:

Employee:
"My VPN access expired."
Agent:
↓
Generate Query Embedding
↓
Search Vector Database
↓
Retrieve VPN Policy
↓
Check HR Status
↓
Trigger Approval Flow
↓
Create Ticket

Retrieval becomes one step within a larger AI system.


Retrieval operations are fully observable.

You can inspect:

  • Retrieved chunks
  • Similarity scores
  • Source documents
  • Token usage
  • Retrieval latency

This makes production RAG systems debuggable.


10:01 Query Received
10:01 Query Embedding Generated
10:01 Vector Search Completed
10:01 5 Chunks Retrieved
10:01 Context Added To Prompt
10:02 Response Generated

Large chunks reduce retrieval precision.


Low-quality embeddings reduce semantic understanding.


Without metadata filtering, irrelevant information may be retrieved.


Retrieving too few chunks may omit important information.


Use folders and listeners to separate domains:

/hr
/finance
/legal
/engineering

Metadata improves both retrieval quality and observability.


Use RAG Runs to investigate:

  • Retrieval latency
  • Missing results
  • Incorrect matches
  • Ingestion failures

User Question
↓
Embedding Model
↓
Vector Search
↓
Retrieved Context
↓
AI Agent
↓
Response

Retrieval is what transforms AI from a chatbot into an enterprise system that understands your organization.


Continue with:

➡️ RAG Runs

Learn how Karyam provides complete observability into ingestion and retrieval pipelines.