Debugging
Building AI systems is only half the challenge.
Operating them in production requires the ability to quickly understand failures, identify bottlenecks, and resolve issues.
Unlike traditional AI systems that behave like black boxes, Karyam provides visibility into every stage of execution.
This allows teams to answer questions such as:
- Why did my Agent fail?
- Why didn’t the AI flow continue?
- Why wasn’t the expected document retrieved?
- Why did the MCP tool fail?
- Why is the model not responding?
- Why are approvals stuck?
The Karyam Debugging AI flow
Section titled “The Karyam Debugging AI flow”Most issues can be diagnosed using four observability surfaces:
Agent Logs ↓AI flow Logs ↓RAG Runs ↓Run InformationTogether these provide a complete picture of system execution.
Debugging Agents
Section titled “Debugging Agents”Agent-related issues can be investigated from:
Agents └── LogsCommon issues include:
- Tool execution failures
- MCP authentication failures
- Approval requests
- Agent transfers
- Model failures
Useful events include:
agent_starttool_callmcp_tool_callhitl_requestagent_transferagent_error
Example
Section titled “Example”agent_starttool_callmcp_tool_callagent_error
Reason:GitHub MCP authentication failedDebugging AI Flows
Section titled “Debugging AI Flows”AI flow execution issues can be investigated from:
OPERATE └── Runs └── LogsCommon issues include:
- Failed steps
- Missing approvals
- Authentication requirements
- Paused AI flows
- Cancelled executions
Useful events include:
REQUESTTOOL_EXECUTION_STARTTOOL_EXECUTION_ENDAPPROVALPAUSEDRESUME
Example
Section titled “Example”REQUESTTOOL_EXECUTION_STARTAPPROVALPAUSED
Status:WAITING_APPROVALThe AI flow is operating correctly and is waiting for human approval.
Debugging Knowledge Retrieval
Section titled “Debugging Knowledge Retrieval”Knowledge ingestion and retrieval issues can be investigated from:
AI Infra └── RAG RunsCommon issues include:
- Embedding failures
- Vector database failures
- Retrieval failures
- Missing context
- Parsing failures
Useful events include:
rag_embedding_startedrag_vector_store_startedrag_retrieval_startedrag_chunk_retrievedrag_context_builtrag_error
Example
Section titled “Example”rag_vector_store_startedrag_error
Reason:Qdrant connection timeoutDebugging MCP Integrations
Section titled “Debugging MCP Integrations”Common MCP issues include:
- Invalid URLs
- Expired OAuth tokens
- Missing permissions
- Authentication failures
- Missing tools
Authentication Failures
Section titled “Authentication Failures”Symptoms:
AUTH_REQUIREDPossible causes:
- Expired OAuth session
- Missing credentials
- Invalid headers
Resolution:
- Re-authenticate the MCP Server.
- Verify request headers.
- Verify OAuth permissions.
Missing Tools
Section titled “Missing Tools”Symptoms:
No tools availableResolution:
MCP Server ↓Tools ↓Sync ToolsDebugging Approvals
Section titled “Debugging Approvals”Human approvals are a common source of confusion during testing.
Symptoms:
PAUSEDStatus: WAITING_APPROVALThis does not indicate a failure.
The AI flow is waiting for an approval decision.
Possible outcomes:
APPROVEDREJECTEDRESUMEDDebugging Costs
Section titled “Debugging Costs”Unexpected costs can usually be traced to:
- Large prompts
- Large retrieval contexts
- Expensive models
- Excessive AI flow steps
Useful dashboards include:
- Highest Token Consuming Flows
- Top Expensive AI flows
- Monthly Cost Trend by Model
Debugging Slow Executions
Section titled “Debugging Slow Executions”Long execution times are often caused by:
- External API latency
- Large retrieval operations
- Slow MCP servers
- Human approvals
Useful metrics include:
- Run duration
- Tool execution times
- Retrieval duration
- Approval wait times
Debugging Checklist
Section titled “Debugging Checklist”When investigating issues, follow this order:
1. Check Run Status2. Review Agent Logs3. Review AI flow Logs4. Review RAG Runs5. Verify MCP Authentication6. Check Token Usage7. Review Cost AnalyticsThis process resolves the majority of production issues.
Common Failure Types
Section titled “Common Failure Types”| Area | Common Cause |
|---|---|
| Agents | Tool failures |
| AI Flows | Waiting approvals |
| MCP | Authentication |
| RAG | Vector database connectivity |
| Models | Provider limits |
| Retrieval | Missing embeddings |
The Karyam Philosophy
Section titled “The Karyam Philosophy”Debugging AI systems should feel similar to debugging distributed software systems.
Request ↓Logs ↓Events ↓Root Cause ↓ResolutionKaryam provides the visibility required to operate AI systems confidently in production environments.
Related Features
Section titled “Related Features”Debugging works alongside:
- Agent Logs
- AI flow Logs
- RAG Logs
- Token Usage
- Cost Analytics
Together they provide complete operational visibility into AI systems.
