Skip to content
Karyam

Token Usage

Large Language Models consume tokens for every request they process.

Understanding token usage is essential for:

  • Cost optimization
  • Capacity planning
  • Budget forecasting
  • Model selection
  • Production operations

Karyam provides detailed token and cost visibility across the entire AI lifecycle.

Unlike traditional AI systems where usage remains hidden behind provider invoices, Karyam exposes token consumption at the run, model, AI flow, and workspace levels.


Token usage information is available from multiple locations within Karyam.

OPERATE
└── Runs
└── Info
└── Usage

RAG executions also expose token and cost information.

Navigate to:

AI Infra
└── RAG Runs
└── Info
└── Usage

This includes:

  • Embedding model usage
  • Retrieval context size
  • Retrieval token usage
  • Completion tokens
  • Retrieval costs

Karyam provides workspace-level token analytics directly on the Dashboard.


Every execution exposes detailed token information.

The maximum number of reasoning steps allowed during execution.

Example:

Max Steps
5

The number of tokens sent to the model.

This includes:

  • User messages
  • System instructions
  • Retrieved context
  • Tool responses
  • Conversation history
  • AI flow state

Example:

Input Tokens
2412

The number of tokens generated by the model.

Example:

Output Tokens
584

The total token consumption for the run.

Total Tokens
=
Input Tokens
+
Output Tokens

Example:

Total Tokens
2996

Karyam calculates costs based on model pricing and provider rates.

Cost associated with input tokens.

Example:

Input Cost
$0.0124

Cost associated with output tokens.

Example:

Output Cost
$0.0042

Additional costs incurred during execution.

Examples include:

  • Provider processing fees
  • Infrastructure overhead
  • Future platform services

Example:

System Cost
$0.0011

Usage
──────────────
Max Steps 5
Input Tokens 2,412
Output Tokens 584
Total Tokens 2,996
Cost
──────────────
Input Cost $0.0124
Output Cost $0.0042
System Cost $0.0011

Each model in Karyam can define pricing and spending controls.

Navigate to:

AI Infra
└── Models
└── Model Configuration

Defines the cost per 1 million input tokens.

Example:

Input Token Cost ($)
0.15

Meaning:

$0.15 per 1,000,000 input tokens

Defines the cost per 1 million output tokens.

Example:

Output Token Cost ($)
0.60

Meaning:

$0.60 per 1,000,000 output tokens

Defines the maximum amount that can be spent using the model.

Example:

Spending Cap ($)
100

This would limit total spending on the model to:

$100

Setting:

Spending Cap ($)
0

means:

Unlimited spending

This is useful for:

  • Development environments
  • Internal testing
  • Dedicated enterprise deployments

Spending caps help organizations:

  • Prevent unexpected costs
  • Enforce budgets
  • Control experimentation
  • Limit runaway AI flows
  • Improve governance

Karyam provides organization-wide visibility into token consumption and model costs.


Identify AI flows consuming the highest number of tokens.

Examples:

  • Long-running flows
  • Large context windows
  • Multi-step reasoning pipelines

This helps identify optimization opportunities.


Identify AI flows generating the highest costs.

This helps organizations:

  • Optimize prompts
  • Reduce context size
  • Improve model selection

Track token usage across all configured models.

Examples:

  • GPT-4o
  • Claude Sonnet
  • Gemini
  • Llama
  • Bedrock Models

Monitor how model costs evolve over time.

Examples:

  • Spending spikes
  • Usage trends
  • Provider distribution

Provides a consolidated workspace-wide token usage overview.

Useful for:

  • Capacity planning
  • Budget forecasting
  • Chargeback reporting

Track overall model spending across providers and teams.

Examples:

  • OpenAI costs
  • Anthropic costs
  • Google costs
  • AWS Bedrock costs

Monitor daily token usage trends.

Useful for detecting:

  • Usage spikes
  • Adoption growth
  • Operational anomalies

Track daily model spending across the organization.

Useful for:

  • Budget monitoring
  • Cost control
  • Forecasting

Token visibility allows organizations to:

  • Optimize AI costs
  • Select appropriate models
  • Forecast budgets
  • Monitor adoption
  • Understand AI utilization patterns

Production AI systems should be observable not only from an execution perspective, but also from a financial perspective.


Review AI flows with unusually high token usage.


Large prompts increase both latency and costs.


Not every workload requires the most expensive model.


Protect against unexpected costs by setting model-level spending limits.


Monitor token and cost trends over time to identify growth patterns and optimization opportunities.


Token Usage works alongside:

  • Agent Logs
  • AI flow Logs
  • RAG Logs
  • Cost Analytics

Together they provide complete operational visibility into AI systems.


Continue with:

➡️ Cost Analytics

Learn how Karyam provides financial visibility and spending insights across your AI infrastructure.