Token Usage
Large Language Models consume tokens for every request they process.
Understanding token usage is essential for:
- Cost optimization
- Capacity planning
- Budget forecasting
- Model selection
- Production operations
Karyam provides detailed token and cost visibility across the entire AI lifecycle.
Unlike traditional AI systems where usage remains hidden behind provider invoices, Karyam exposes token consumption at the run, model, AI flow, and workspace levels.
Where to Find Token Usage
Section titled “Where to Find Token Usage”Token usage information is available from multiple locations within Karyam.
AI flow Runs
Section titled “AI flow Runs”OPERATE └── Runs └── Info └── UsageRAG Runs
Section titled “RAG Runs”RAG executions also expose token and cost information.
Navigate to:
AI Infra └── RAG Runs └── Info └── UsageThis includes:
- Embedding model usage
- Retrieval context size
- Retrieval token usage
- Completion tokens
- Retrieval costs
Dashboard Analytics
Section titled “Dashboard Analytics”Karyam provides workspace-level token analytics directly on the Dashboard.
Run-Level Usage Metrics
Section titled “Run-Level Usage Metrics”Every execution exposes detailed token information.
Max Steps
Section titled “Max Steps”The maximum number of reasoning steps allowed during execution.
Example:
Max Steps5Input Tokens
Section titled “Input Tokens”The number of tokens sent to the model.
This includes:
- User messages
- System instructions
- Retrieved context
- Tool responses
- Conversation history
- AI flow state
Example:
Input Tokens2412Output Tokens
Section titled “Output Tokens”The number of tokens generated by the model.
Example:
Output Tokens584Total Tokens
Section titled “Total Tokens”The total token consumption for the run.
Total Tokens=Input Tokens+Output TokensExample:
Total Tokens2996Run-Level Cost Metrics
Section titled “Run-Level Cost Metrics”Karyam calculates costs based on model pricing and provider rates.
Input Cost
Section titled “Input Cost”Cost associated with input tokens.
Example:
Input Cost$0.0124Output Cost
Section titled “Output Cost”Cost associated with output tokens.
Example:
Output Cost$0.0042System Cost
Section titled “System Cost”Additional costs incurred during execution.
Examples include:
- Provider processing fees
- Infrastructure overhead
- Future platform services
Example:
System Cost$0.0011Example Usage Section
Section titled “Example Usage Section”Usage──────────────Max Steps 5Input Tokens 2,412Output Tokens 584Total Tokens 2,996
Cost──────────────Input Cost $0.0124Output Cost $0.0042System Cost $0.0011Model Pricing Configuration
Section titled “Model Pricing Configuration”Each model in Karyam can define pricing and spending controls.
Navigate to:
AI Infra └── Models └── Model ConfigurationInput Token Cost ($)
Section titled “Input Token Cost ($)”Defines the cost per 1 million input tokens.
Example:
Input Token Cost ($)0.15Meaning:
$0.15 per 1,000,000 input tokensOutput Token Cost ($)
Section titled “Output Token Cost ($)”Defines the cost per 1 million output tokens.
Example:
Output Token Cost ($)0.60Meaning:
$0.60 per 1,000,000 output tokensSpending Cap ($)
Section titled “Spending Cap ($)”Defines the maximum amount that can be spent using the model.
Example:
Spending Cap ($)100This would limit total spending on the model to:
$100Unlimited Spending
Section titled “Unlimited Spending”Setting:
Spending Cap ($)0means:
Unlimited spendingThis is useful for:
- Development environments
- Internal testing
- Dedicated enterprise deployments
Why Spending Caps Matter
Section titled “Why Spending Caps Matter”Spending caps help organizations:
- Prevent unexpected costs
- Enforce budgets
- Control experimentation
- Limit runaway AI flows
- Improve governance
Dashboard Analytics
Section titled “Dashboard Analytics”Karyam provides organization-wide visibility into token consumption and model costs.
Highest Token Consuming Flows
Section titled “Highest Token Consuming Flows”Identify AI flows consuming the highest number of tokens.
Examples:
- Long-running flows
- Large context windows
- Multi-step reasoning pipelines
This helps identify optimization opportunities.
Top Expensive AI flows
Section titled “Top Expensive AI flows”Identify AI flows generating the highest costs.
This helps organizations:
- Optimize prompts
- Reduce context size
- Improve model selection
Monthly Token Consumption by Model
Section titled “Monthly Token Consumption by Model”Track token usage across all configured models.
Examples:
- GPT-4o
- Claude Sonnet
- Gemini
- Llama
- Bedrock Models
Monthly Cost Trend by Model
Section titled “Monthly Cost Trend by Model”Monitor how model costs evolve over time.
Examples:
- Spending spikes
- Usage trends
- Provider distribution
Monthly Total Token Consumption
Section titled “Monthly Total Token Consumption”Provides a consolidated workspace-wide token usage overview.
Useful for:
- Capacity planning
- Budget forecasting
- Chargeback reporting
Monthly Total Model Cost Overview
Section titled “Monthly Total Model Cost Overview”Track overall model spending across providers and teams.
Examples:
- OpenAI costs
- Anthropic costs
- Google costs
- AWS Bedrock costs
Daily Token Consumption
Section titled “Daily Token Consumption”Monitor daily token usage trends.
Useful for detecting:
- Usage spikes
- Adoption growth
- Operational anomalies
Daily Token Cost
Section titled “Daily Token Cost”Track daily model spending across the organization.
Useful for:
- Budget monitoring
- Cost control
- Forecasting
Why Token Usage Matters
Section titled “Why Token Usage Matters”Token visibility allows organizations to:
- Optimize AI costs
- Select appropriate models
- Forecast budgets
- Monitor adoption
- Understand AI utilization patterns
Production AI systems should be observable not only from an execution perspective, but also from a financial perspective.
Best Practices
Section titled “Best Practices”Monitor Expensive AI flows
Section titled “Monitor Expensive AI flows”Review AI flows with unusually high token usage.
Optimize Context Size
Section titled “Optimize Context Size”Large prompts increase both latency and costs.
Use Appropriate Models
Section titled “Use Appropriate Models”Not every workload requires the most expensive model.
Configure Spending Caps
Section titled “Configure Spending Caps”Protect against unexpected costs by setting model-level spending limits.
Track Usage Trends
Section titled “Track Usage Trends”Monitor token and cost trends over time to identify growth patterns and optimization opportunities.
Related Observability Features
Section titled “Related Observability Features”Token Usage works alongside:
- Agent Logs
- AI flow Logs
- RAG Logs
- Cost Analytics
Together they provide complete operational visibility into AI systems.
Next Steps
Section titled “Next Steps”Continue with:
➡️ Cost Analytics
Learn how Karyam provides financial visibility and spending insights across your AI infrastructure.
