Skip to content
Karyam

Embeddings

Embeddings are the foundation of semantic search and Retrieval-Augmented Generation (RAG).

They transform text, documents, images, and other content into numerical representations that AI systems can search and understand based on meaning rather than exact keywords.

Without embeddings, AI systems cannot effectively retrieve enterprise knowledge.


Consider the following question:

My VPN access expired. Can you restore it?

The relevant company document may contain:

VPN credentials must be renewed every 90 days and require manager approval.

A traditional keyword search may struggle because:

  • “expired” ≠ “renewed”
  • “restore” ≠ “approval”
  • “VPN access” ≠ “VPN credentials”

Humans understand these concepts are related.

Embeddings allow AI systems to understand this relationship as well.


An embedding is a numerical representation of content.

For example:

"VPN credentials require manager approval"
↓
[0.284, -0.193, 0.729, 0.041, ...]

This list of numbers represents the semantic meaning of the text.

Documents with similar meanings produce vectors that are close together in vector space.


Similar Meaning Produces Similar Embeddings

Section titled “Similar Meaning Produces Similar Embeddings”

Example:

Text Semantic Similarity
VPN access expired High
VPN credentials require renewal High
Reset employee password Medium
Company quarterly revenue report Low

The embedding model learns these relationships automatically.


When a file is uploaded, Karyam executes the following process:

File Upload
↓
Embedding Listener Triggered
↓
Document Parsing
↓
Chunk Generation
↓
Embedding Model
↓
Vector Creation
↓
Vector Database Storage

The generated embeddings are then available for retrieval by AI agents and workflows.


Embeddings are generated for both:

Example:

VPN credentials require approval.

↓

[0.284, -0.193, 0.729, ...]

Example:

Can you restore my VPN access?

↓

[0.276, -0.201, 0.741, ...]

The vector database compares embeddings and retrieves the closest matches.

User Question
↓
Query Embedding
↓
Vector Similarity Search
↓
Relevant Chunks Retrieved
↓
Context Added To Prompt

This process happens automatically during retrieval.


Embedding models are different from chat models.

Their only job is to convert content into vectors.

Examples include:

  • Gemini Embeddings
  • Voyage AI
  • Ollama Embeddings

Capability Chat Model Embedding Model
Generate responses Yes No
Hold conversations Yes No
Generate embeddings No Yes
Semantic search No Yes

Both model types are required for Retrieval-Augmented Generation.


Karyam supports multiple embedding providers.

Examples:

Examples:

  • gemini-embedding-2

Examples:

  • voyage-multimodal-3.5

Examples:

  • qwen3-embedding:8b

Ideal for on-premises deployments.


Each embedding model produces vectors of a fixed size.

Examples:

Model Dimensions
gemini-embedding-2 3072
voyage-multimodal-3.5 1024
qwen3-embedding:8b 4096

Vector databases must support the configured dimensions.


In Karyam, users do not manually generate embeddings.

Embeddings are created automatically when:

  • Files are uploaded
  • Embedding listeners are triggered
  • Chunking is completed

This makes knowledge ingestion fully automated.


Embedding operations are fully observable through:

  • RAG Runs
  • Chunk Metadata
  • Embedding Status
  • Error Tracking

This allows teams to monitor and debug ingestion pipelines.


vpn-policy.pdf uploaded
↓
Listener triggered
↓
35 chunks generated
↓
35 embeddings created
↓
Stored in Qdrant
↓
RAG Run completed

Use A Single Embedding Model Per Vector Database

Section titled “Use A Single Embedding Model Per Vector Database”

Changing dimensions requires rebuilding storage.


Always choose the models based on the tasks to be performed


Poor retrieval quality is often caused by:

  • Weak embedding models
  • Poor chunking
  • Missing metadata
  • Incorrect vector dimensions

Files
↓
Chunking
↓
Embeddings
↓
Vector Database
↓
Retrieval
↓
Agents

Embeddings are the bridge between enterprise knowledge and AI reasoning.


Continue with:

➡️ File Uploads

Learn how Karyam ingests enterprise content and automatically triggers embedding pipelines.