Chunking
Large Language Models do not search entire documents.
They search chunks of documents.
Chunking is the process of breaking large files into smaller pieces before embeddings are generated and stored in a vector database.
Without chunking, Retrieval-Augmented Generation (RAG) becomes slow, expensive, and inaccurate.
Why Chunking Exists
Section titled “Why Chunking Exists”Consider an employee handbook containing:
200 Pages150,000 WordsAn employee asks:
How many sick leaves do I get per year?
The answer exists in only a small section:
Employees are entitled to 12 sick leaves annually.Retrieving the entire handbook would:
- Increase token usage
- Slow down responses
- Introduce irrelevant context
- Reduce answer quality
Instead, Karyam retrieves only the relevant section.
The Problem Without Chunking
Section titled “The Problem Without Chunking”Without chunking:
Employee Handbook.pdf ↓One Large Embedding ↓Poor RetrievalThe model receives huge amounts of unrelated information.
This often leads to:
- Hallucinations
- Missing information
- Context overflow
- Increased costs
The Karyam Approach
Section titled “The Karyam Approach”Karyam automatically splits documents into smaller chunks before generating embeddings.
Employee Handbook.pdf ↓ Chunk 1 Chunk 2 Chunk 3 Chunk 4 Chunk 5 ↓Embeddings Generated ↓Stored In Vector DatabaseEach chunk receives:
- Content
- Embedding
- Metadata
- Source references
Example
Section titled “Example”Original document:
VPN credentials expire every 90 days.
VPN renewals require manager approval.
Contractors require security review before renewal.After chunking:
Chunk 1
Section titled “Chunk 1”VPN credentials expire every 90 days.Chunk 2
Section titled “Chunk 2”VPN renewals require manager approval.Chunk 3
Section titled “Chunk 3”Contractors require security review before renewal.Each chunk receives its own embedding.
Chunk Lifecycle
Section titled “Chunk Lifecycle”File Uploaded ↓Document Parsing ↓Chunk Generation ↓Embedding Creation ↓Vector Database Storage ↓RAG Run CreatedChunking occurs automatically during ingestion.
Chunk Metadata
Section titled “Chunk Metadata”Every chunk stores metadata alongside its embedding.
Examples:
| Metadata | Example |
|---|---|
| File Path | /knowledge-base/it/vpn-policy.pdf |
| Chunk Index | 12 |
| Workspace UUID | ws_12345 |
| Source File | vpn-policy.pdf |
| Upload Time | 2026-07-16 |
This metadata improves:
- Traceability
- Retrieval quality
- Observability
- Debugging
Why Smaller Chunks Improve Retrieval
Section titled “Why Smaller Chunks Improve Retrieval”Suppose an employee asks:
Can contractors renew VPN access?
The query embedding is compared against every chunk.
The retrieval system identifies:
Contractors require security review before renewal.instead of returning the entire VPN policy document.
This results in:
- Faster retrieval
- Lower token usage
- More accurate answers
Chunk Size Tradeoffs
Section titled “Chunk Size Tradeoffs”Choosing the correct chunk size is important.
Small Chunks
Section titled “Small Chunks”Advantages:
- Highly precise retrieval
- Lower token usage
- Better relevance
Disadvantages:
- Loss of context
- Fragmented information
Large Chunks
Section titled “Large Chunks”Advantages:
- More context preserved
- Better reasoning across paragraphs
Disadvantages:
- Higher token costs
- Lower retrieval precision
The Goal
Section titled “The Goal”The ideal chunk size balances:
Context Preservation +Retrieval Precision +Token EfficiencyChunk Overlap
Section titled “Chunk Overlap”Some chunking strategies introduce overlap.
Example:
Chunk 1--------VPN credentials expire every 90 days.VPN renewals require manager approval.
Chunk 2--------VPN renewals require manager approval.Contractors require security review.Overlap helps preserve context between chunks.
Chunking In Karyam
Section titled “Chunking In Karyam”Karyam automatically handles:
- Parsing
- Chunk generation
- Metadata creation
- Embedding generation
- Storage
Users do not manually create chunks.
This makes ingestion pipelines fully automated.
Observability
Section titled “Observability”Chunk generation is fully observable through:
- RAG Runs
- Chunk Metadata
- Embedding Logs
- Processing Status
Example:
vpn-policy.pdf ↓32 Chunks Generated ↓32 Embeddings Created ↓Stored In Qdrant ↓RAG Run CompletedExample Retrieval Flow
Section titled “Example Retrieval Flow”Employee:"My VPN expired."
Agent:↓Generate Query Embedding↓Search Vector Database↓Retrieve Relevant Chunks↓Generate ResponseRetrieved chunk:
VPN credentials expire every 90 days and require manager approval for renewal.Best Practices
Section titled “Best Practices”Use Structured Documents
Section titled “Use Structured Documents”Documents with clear headings and sections produce better chunks.
Keep Related Information Together
Section titled “Keep Related Information Together”Avoid splitting related information across multiple files when possible.
Monitor Retrieval Quality
Section titled “Monitor Retrieval Quality”Poor retrieval quality often originates from:
- Poor chunk boundaries
- Weak embeddings
- Missing metadata
Use RAG Runs
Section titled “Use RAG Runs”RAG Runs help diagnose:
- Parsing failures
- Missing chunks
- Embedding errors
- Retrieval issues
The Karyam Model
Section titled “The Karyam Model”Files ↓Chunking ↓Embeddings ↓Vector Database ↓Retrieval ↓AgentsChunking is the bridge between enterprise documents and AI understanding.
Next Steps
Section titled “Next Steps”Continue with:
➡️ Retrieval
Learn how Karyam performs semantic search and delivers context to AI systems in real time.
