File Uploads
File uploads are the entry point for enterprise knowledge in Karyam.
Unlike traditional AI platforms where documents are manually attached to chatbots, Karyam uses an event-driven ingestion architecture.
When files are uploaded, Karyam can automatically:
- Parse content
- Generate chunks
- Create embeddings
- Store vectors
- Track ingestion progress
- Expose observability through RAG Runs
This allows enterprise knowledge to continuously evolve alongside business operations.
The File Upload Pipeline
Section titled “The File Upload Pipeline”Uploading a file can trigger an entire retrieval pipeline.
File Upload ↓Embedding Listener Triggered ↓Document Parsing ↓Chunk Generation ↓Embedding Creation ↓Vector Storage ↓RAG Run CreatedNo manual synchronization is required.
Supported File Types
Section titled “Supported File Types”Karyam supports multiple content formats.
Documents
Section titled “Documents”- PDF (
.pdf) - Microsoft Word (
.docx) - Text (
.txt) - Markdown (
.md) - CSV (
.csv) - JSON (
.json)
Images
Section titled “Images”- PNG (
.png) - JPEG (
.jpg) - WebP (
.webp) - GIF (
.gif)
Image understanding depends on the configured AI model.
- MP3 (
.mp3) - WAV (
.wav) - M4A (
.m4a) - FLAC (
.flac) - WebM (
.webm)
Audio ingestion requires a configured speech-to-text integration.
Folder Organization
Section titled “Folder Organization”Files are organized using folders.
Embedding listeners can monitor specific folders and automatically trigger ingestion when new content appears.
Example:
knowledge-base/│├── hr/│ ├── employee-handbook.pdf│ ├── leave-policy.pdf│ └── benefits-guide.pdf│├── finance/│ ├── reimbursement-policy.pdf│ └── procurement-guidelines.pdf│└── it/ ├── vpn-policy.pdf └── access-sop.pdfThis enables teams to isolate retrieval pipelines by business domain.
Embedding Listeners
Section titled “Embedding Listeners”Embedding listeners monitor folders for new files.
Example configuration:
| Setting | Value |
|---|---|
| Listener Type | Embedding |
| Folder Path | /knowledge-base/it |
| Vector Database | IT Vector DB |
| Embedding Model | gemini-embedding-2 |
Whenever a file is uploaded into the configured folder:
knowledge-base/it/└── vpn-policy.pdfthe listener automatically starts the ingestion process.
Metadata
Section titled “Metadata”Karyam stores metadata alongside embeddings.
Examples include:
- File path
- File name
- Upload timestamp
- Workspace identifier
- Chunk position
- Source references
Metadata improves retrieval quality and traceability.
Sidecar Descriptions
Section titled “Sidecar Descriptions”Additional context can be provided using sidecar description files.
Example:
vpn-policy.pdfvpn-policy.pdf.txtContents:
This document contains VPN access policies,renewal procedures, and approval requirements.The description becomes additional context during retrieval.
Chunk Creation
Section titled “Chunk Creation”Large files are automatically divided into smaller chunks before embeddings are generated.
Example:
50-page employee handbook ↓175 chunks generated ↓175 embeddings createdChunking improves retrieval relevance and context quality.
Upload Example
Section titled “Upload Example”A user uploads:
vpn-policy.pdfKaryam executes:
Upload Received ↓File Parsed ↓32 Chunks Created ↓32 Embeddings Generated ↓Stored In Vector Database ↓RAG Run CreatedThe file is now available to agents and workflows.
Monitoring Upload Processing
Section titled “Monitoring Upload Processing”Navigate to:
RAG RunsYou can monitor:
- Parsing status
- Chunk generation
- Embedding progress
- Vector storage
- Failures and retries
This provides complete visibility into ingestion pipelines.
Best Practices
Section titled “Best Practices”Organize By Domain
Section titled “Organize By Domain”Example:
/hr/finance/legal/engineering/operationsUse Descriptive File Names
Section titled “Use Descriptive File Names”Good:
vpn-access-policy-2026.pdfemployee-handbook-v3.pdfAvoid:
document1.pdfpolicy-final-final.pdfAdd Context Using Sidecar Files
Section titled “Add Context Using Sidecar Files”Sidecar descriptions improve retrieval quality for:
- Images
- Logos
- Scanned documents
- Ambiguous content
Monitor RAG Runs
Section titled “Monitor RAG Runs”Always monitor ingestion status after large uploads.
This helps detect:
- Parsing failures
- Embedding errors
- Storage issues
Example Enterprise Workflow
Section titled “Example Enterprise Workflow”New Policy Uploaded ↓Embedding Listener Triggered ↓Embeddings Generated ↓Vector Database Updated ↓Agents Immediately Gain AccessNo redeployment.
No retraining.
No manual synchronization.
The Karyam Model
Section titled “The Karyam Model”Files ↓Listeners ↓Chunking ↓Embeddings ↓Vector Database ↓Retrieval ↓AgentsFile uploads are not just storage.
They are the beginning of an AI system’s understanding of your organization.
Next Steps
Section titled “Next Steps”Continue with:
➡️ Chunking
Learn how Karyam breaks large documents into retrieval-friendly pieces before generating embeddings.
