Skip to content
Karyam

Models

Models are the reasoning engine behind every AI system in Karyam.

Whenever an Agent answers a question, an AI Flow executes, or a document is embedded for retrieval, a Model performs the underlying computation.

Models provide the intelligence that enables AI systems to understand, reason, and generate responses.


Every AI capability depends on a model.

Different models are optimized for different tasks, including:

  • Conversational AI
  • Reasoning
  • Code generation
  • Embedding generation
  • Vision understanding
  • Audio processing

Karyam allows you to configure and manage these models independently from the applications that use them.


Karyam supports two primary model types.

Chat models are used by:

  • Agents
  • AI Flows
  • API executions
  • Embedded chat widgets

These models perform reasoning, planning, and response generation.

Examples include:

  • GPT-4o
  • Claude Sonnet
  • Gemini
  • Bedrock Models
  • Ollama Models

Embedding models convert text into numerical vectors.

These vectors are stored inside Vector Databases and enable semantic search.

Embedding models are used by:

  • Embedding Listeners
  • Knowledge ingestion
  • Retrieval-Augmented Generation (RAG)

Unlike chat models, embedding models do not generate responses.

They generate vector representations of content.


Karyam supports models from multiple providers.

Examples include:

  • ChatGPT
  • Google Gemini
  • Anthropic
  • Perplexity
  • AWS Bedrock
  • OpenAI Compat
  • Cloudflare
  • Ollama
  • Voyage AI

This allows organizations to choose the models that best meet their performance, cost, and compliance requirements.


Each model is configured with provider-specific information such as:

  • Model name
  • Provider
  • API credentials
  • Endpoint configuration
  • Token pricing
  • Spending limits

These settings determine how the model behaves within the platform.


Models can define pricing information for cost tracking.

Typical configuration includes:

Setting Description
Input Token Cost Cost per one million input tokens
Output Token Cost Cost per one million output tokens
Spending Cap Maximum allowed spend (0 = unlimited)

This information powers Karyam’s token usage and cost analytics.


Models are shared infrastructure.

A single model may be used by multiple components.

Model
├── Agent
├── AI Flow
├── Embedding Listener
└── RAG Retrieval

This allows organizations to centrally manage AI providers while reusing them across multiple AI systems.


Models
↓
Agents
↓
AI Flows
↓
Runs
Embedding Models
↓
Embedding Listeners
↓
Vector Databases
↓
RAG Runs

Models provide the intelligence layer for both conversational AI and enterprise knowledge retrieval.


AI systems should not be tightly coupled to a single provider.

Karyam separates AI infrastructure from business logic, allowing organizations to evolve their model strategy without rebuilding their AI systems.