Skip to content
Karyam

Vector Databases

Vector Databases are the knowledge storage layer of Karyam.

Instead of storing documents as plain text, they store vector embeddings—numerical representations of content that allow AI systems to search by meaning rather than exact words.

This enables Retrieval-Augmented Generation (RAG), where AI systems retrieve relevant business knowledge before generating a response.


Large Language Models do not automatically know your organization’s information.

They cannot answer questions about:

  • Internal documentation
  • Company policies
  • Product specifications
  • Customer knowledge
  • Technical manuals
  • Private business data

Vector Databases bridge this gap by making organizational knowledge searchable using semantic similarity.


When knowledge is added to Karyam, it passes through an ingestion pipeline.

Document
↓
Chunking
↓
Embedding Model
↓
Vector Embeddings
↓
Vector Database

Later, when an Agent receives a request:

User Question
↓
Embedding Model
↓
Semantic Search
↓
Relevant Chunks Retrieved
↓
LLM Generates Response

Instead of searching for matching keywords, the system searches for the most semantically relevant information.


A Vector Database stores more than embeddings.

Each stored record typically contains:

  • Vector embedding
  • Original text chunk
  • Source document
  • File path
  • Metadata
  • Workspace information

This allows Karyam to retrieve both the relevant content and its associated context.


Unlike traditional databases, Vector Databases search by meaning.

For example:

Stored document

Employees must use VPN when working remotely.

A user asks:

How do I securely connect from home?

Although neither sentence uses the exact same words, semantic search identifies that they express the same concept and retrieves the correct document.


Karyam supports multiple vector database providers.

Current providers include:

  • Qdrant
  • pgvector

Additional providers can be integrated as the platform evolves.


Vector Databases are shared infrastructure.

They support:

  • Knowledge retrieval
  • AI Agents
  • AI Flows
  • Enterprise search
  • Document understanding
  • RAG pipelines

A single Vector Database can serve multiple AI systems across a workspace.


Vector Databases are populated through Embedding Listeners.

The ingestion workflow is:

File Upload
↓
Embedding Listener
↓
Artifacts Generated
↓
Chunking
↓
Embedding Model
↓
Vector Database

Once stored, the knowledge becomes immediately available for semantic retrieval.


During execution, Agents and AI Flows retrieve relevant knowledge before generating a response.

Question
↓
Embedding
↓
Vector Search
↓
Relevant Chunks
↓
Model
↓
Final Response

This process improves response quality by grounding the model in your organization’s data.


Every ingestion and retrieval operation is tracked in RAG Runs.

RAG Runs provide visibility into:

  • Artifact generation
  • Chunk creation
  • Embedding generation
  • Vector storage
  • Retrieval operations
  • Search performance
  • Processing failures

This makes the knowledge pipeline fully observable.


Files
↓
Embedding Listener
↓
Embedding Model
↓
Vector Database
↓
Retrieval
↓
Agent / AI Flow
↓
Run

Vector Databases form the persistent knowledge layer that powers Retrieval-Augmented Generation across the platform.


Enterprise AI should answer questions using your organization’s knowledge—not just what the model learned during training.

Vector Databases make private knowledge searchable, reusable, and available to every AI system built on Karyam.