Embeddings
Embeddings are the foundation of semantic search and Retrieval-Augmented Generation (RAG).
They transform text, documents, images, and other content into numerical representations that AI systems can search and understand based on meaning rather than exact keywords.
Without embeddings, AI systems cannot effectively retrieve enterprise knowledge.
Why Embeddings Matter
Section titled “Why Embeddings Matter”Consider the following question:
My VPN access expired. Can you restore it?
The relevant company document may contain:
VPN credentials must be renewed every 90 days and require manager approval.
A traditional keyword search may struggle because:
- “expired” ≠ “renewed”
- “restore” ≠ “approval”
- “VPN access” ≠ “VPN credentials”
Humans understand these concepts are related.
Embeddings allow AI systems to understand this relationship as well.
What Is An Embedding?
Section titled “What Is An Embedding?”An embedding is a numerical representation of content.
For example:
"VPN credentials require manager approval"
↓
[0.284, -0.193, 0.729, 0.041, ...]This list of numbers represents the semantic meaning of the text.
Documents with similar meanings produce vectors that are close together in vector space.
Similar Meaning Produces Similar Embeddings
Section titled “Similar Meaning Produces Similar Embeddings”Example:
| Text | Semantic Similarity |
|---|---|
| VPN access expired | High |
| VPN credentials require renewal | High |
| Reset employee password | Medium |
| Company quarterly revenue report | Low |
The embedding model learns these relationships automatically.
How Embeddings Work In Karyam
Section titled “How Embeddings Work In Karyam”When a file is uploaded, Karyam executes the following process:
File Upload ↓Embedding Listener Triggered ↓Document Parsing ↓Chunk Generation ↓Embedding Model ↓Vector Creation ↓Vector Database StorageThe generated embeddings are then available for retrieval by AI agents and workflows.
Query Embeddings
Section titled “Query Embeddings”Embeddings are generated for both:
Documents
Section titled “Documents”Example:
VPN credentials require approval.↓
[0.284, -0.193, 0.729, ...]User Questions
Section titled “User Questions”Example:
Can you restore my VPN access?↓
[0.276, -0.201, 0.741, ...]Similarity Search
Section titled “Similarity Search”The vector database compares embeddings and retrieves the closest matches.
User Question ↓Query Embedding ↓Vector Similarity Search ↓Relevant Chunks Retrieved ↓Context Added To PromptThis process happens automatically during retrieval.
Embedding Models
Section titled “Embedding Models”Embedding models are different from chat models.
Their only job is to convert content into vectors.
Examples include:
- Gemini Embeddings
- Voyage AI
- Ollama Embeddings
Chat Models vs Embedding Models
Section titled “Chat Models vs Embedding Models”| Capability | Chat Model | Embedding Model |
|---|---|---|
| Generate responses | Yes | No |
| Hold conversations | Yes | No |
| Generate embeddings | No | Yes |
| Semantic search | No | Yes |
Both model types are required for Retrieval-Augmented Generation.
Supported Embedding Providers
Section titled “Supported Embedding Providers”Karyam supports multiple embedding providers.
Examples:
Google Gemini
Section titled “Google Gemini”Examples:
- gemini-embedding-2
Voyage AI
Section titled “Voyage AI”Examples:
- voyage-multimodal-3.5
Ollama
Section titled “Ollama”Examples:
- qwen3-embedding:8b
Ideal for on-premises deployments.
Embedding Dimensions
Section titled “Embedding Dimensions”Each embedding model produces vectors of a fixed size.
Examples:
| Model | Dimensions |
|---|---|
| gemini-embedding-2 | 3072 |
| voyage-multimodal-3.5 | 1024 |
| qwen3-embedding:8b | 4096 |
Vector databases must support the configured dimensions.
Embeddings Are Generated Automatically
Section titled “Embeddings Are Generated Automatically”In Karyam, users do not manually generate embeddings.
Embeddings are created automatically when:
- Files are uploaded
- Embedding listeners are triggered
- Chunking is completed
This makes knowledge ingestion fully automated.
Observability
Section titled “Observability”Embedding operations are fully observable through:
- RAG Runs
- Chunk Metadata
- Embedding Status
- Error Tracking
This allows teams to monitor and debug ingestion pipelines.
Example Execution
Section titled “Example Execution”vpn-policy.pdf uploaded ↓Listener triggered ↓35 chunks generated ↓35 embeddings created ↓Stored in Qdrant ↓RAG Run completedBest Practices
Section titled “Best Practices”Use A Single Embedding Model Per Vector Database
Section titled “Use A Single Embedding Model Per Vector Database”Changing dimensions requires rebuilding storage.
Choose Models Based On Workload
Section titled “Choose Models Based On Workload”Always choose the models based on the tasks to be performed
Monitor Retrieval Quality
Section titled “Monitor Retrieval Quality”Poor retrieval quality is often caused by:
- Weak embedding models
- Poor chunking
- Missing metadata
- Incorrect vector dimensions
The Karyam Model
Section titled “The Karyam Model”Files ↓Chunking ↓Embeddings ↓Vector Database ↓Retrieval ↓AgentsEmbeddings are the bridge between enterprise knowledge and AI reasoning.
Next Steps
Section titled “Next Steps”Continue with:
➡️ File Uploads
Learn how Karyam ingests enterprise content and automatically triggers embedding pipelines.
