Vector Databases
Vector Databases are the knowledge storage layer of Karyam.
Instead of storing documents as plain text, they store vector embeddings—numerical representations of content that allow AI systems to search by meaning rather than exact words.
This enables Retrieval-Augmented Generation (RAG), where AI systems retrieve relevant business knowledge before generating a response.
Why Vector Databases Exist
Section titled “Why Vector Databases Exist”Large Language Models do not automatically know your organization’s information.
They cannot answer questions about:
- Internal documentation
- Company policies
- Product specifications
- Customer knowledge
- Technical manuals
- Private business data
Vector Databases bridge this gap by making organizational knowledge searchable using semantic similarity.
How It Works
Section titled “How It Works”When knowledge is added to Karyam, it passes through an ingestion pipeline.
Document ↓Chunking ↓Embedding Model ↓Vector Embeddings ↓Vector DatabaseLater, when an Agent receives a request:
User Question ↓Embedding Model ↓Semantic Search ↓Relevant Chunks Retrieved ↓LLM Generates ResponseInstead of searching for matching keywords, the system searches for the most semantically relevant information.
What is Stored?
Section titled “What is Stored?”A Vector Database stores more than embeddings.
Each stored record typically contains:
- Vector embedding
- Original text chunk
- Source document
- File path
- Metadata
- Workspace information
This allows Karyam to retrieve both the relevant content and its associated context.
Semantic Search
Section titled “Semantic Search”Unlike traditional databases, Vector Databases search by meaning.
For example:
Stored document
Employees must use VPN when working remotely.A user asks:
How do I securely connect from home?Although neither sentence uses the exact same words, semantic search identifies that they express the same concept and retrieves the correct document.
Supported Providers
Section titled “Supported Providers”Karyam supports multiple vector database providers.
Current providers include:
- Qdrant
- pgvector
Additional providers can be integrated as the platform evolves.
Where Vector Databases Are Used
Section titled “Where Vector Databases Are Used”Vector Databases are shared infrastructure.
They support:
- Knowledge retrieval
- AI Agents
- AI Flows
- Enterprise search
- Document understanding
- RAG pipelines
A single Vector Database can serve multiple AI systems across a workspace.
Knowledge Ingestion
Section titled “Knowledge Ingestion”Vector Databases are populated through Embedding Listeners.
The ingestion workflow is:
File Upload ↓Embedding Listener ↓Artifacts Generated ↓Chunking ↓Embedding Model ↓Vector DatabaseOnce stored, the knowledge becomes immediately available for semantic retrieval.
Retrieval
Section titled “Retrieval”During execution, Agents and AI Flows retrieve relevant knowledge before generating a response.
Question ↓Embedding ↓Vector Search ↓Relevant Chunks ↓Model ↓Final ResponseThis process improves response quality by grounding the model in your organization’s data.
Vector Databases and RAG Runs
Section titled “Vector Databases and RAG Runs”Every ingestion and retrieval operation is tracked in RAG Runs.
RAG Runs provide visibility into:
- Artifact generation
- Chunk creation
- Embedding generation
- Vector storage
- Retrieval operations
- Search performance
- Processing failures
This makes the knowledge pipeline fully observable.
Relationship to Other Concepts
Section titled “Relationship to Other Concepts”Files ↓Embedding Listener ↓Embedding Model ↓Vector Database ↓Retrieval ↓Agent / AI Flow ↓RunVector Databases form the persistent knowledge layer that powers Retrieval-Augmented Generation across the platform.
The Karyam Philosophy
Section titled “The Karyam Philosophy”Enterprise AI should answer questions using your organization’s knowledge—not just what the model learned during training.
Vector Databases make private knowledge searchable, reusable, and available to every AI system built on Karyam.
