Skip to content
Karyam

Chunking

Large Language Models do not search entire documents.

They search chunks of documents.

Chunking is the process of breaking large files into smaller pieces before embeddings are generated and stored in a vector database.

Without chunking, Retrieval-Augmented Generation (RAG) becomes slow, expensive, and inaccurate.


Consider an employee handbook containing:

200 Pages
150,000 Words

An employee asks:

How many sick leaves do I get per year?

The answer exists in only a small section:

Employees are entitled to 12 sick leaves annually.

Retrieving the entire handbook would:

  • Increase token usage
  • Slow down responses
  • Introduce irrelevant context
  • Reduce answer quality

Instead, Karyam retrieves only the relevant section.


Without chunking:

Employee Handbook.pdf
↓
One Large Embedding
↓
Poor Retrieval

The model receives huge amounts of unrelated information.

This often leads to:

  • Hallucinations
  • Missing information
  • Context overflow
  • Increased costs

Karyam automatically splits documents into smaller chunks before generating embeddings.

Employee Handbook.pdf
↓
Chunk 1
Chunk 2
Chunk 3
Chunk 4
Chunk 5
↓
Embeddings Generated
↓
Stored In Vector Database

Each chunk receives:

  • Content
  • Embedding
  • Metadata
  • Source references

Original document:

VPN credentials expire every 90 days.
VPN renewals require manager approval.
Contractors require security review before renewal.

After chunking:

VPN credentials expire every 90 days.

VPN renewals require manager approval.

Contractors require security review before renewal.

Each chunk receives its own embedding.


File Uploaded
↓
Document Parsing
↓
Chunk Generation
↓
Embedding Creation
↓
Vector Database Storage
↓
RAG Run Created

Chunking occurs automatically during ingestion.


Every chunk stores metadata alongside its embedding.

Examples:

Metadata Example
File Path /knowledge-base/it/vpn-policy.pdf
Chunk Index 12
Workspace UUID ws_12345
Source File vpn-policy.pdf
Upload Time 2026-07-16

This metadata improves:

  • Traceability
  • Retrieval quality
  • Observability
  • Debugging

Suppose an employee asks:

Can contractors renew VPN access?

The query embedding is compared against every chunk.

The retrieval system identifies:

Contractors require security review before renewal.

instead of returning the entire VPN policy document.

This results in:

  • Faster retrieval
  • Lower token usage
  • More accurate answers

Choosing the correct chunk size is important.

Advantages:

  • Highly precise retrieval
  • Lower token usage
  • Better relevance

Disadvantages:

  • Loss of context
  • Fragmented information

Advantages:

  • More context preserved
  • Better reasoning across paragraphs

Disadvantages:

  • Higher token costs
  • Lower retrieval precision

The ideal chunk size balances:

Context Preservation
+
Retrieval Precision
+
Token Efficiency

Some chunking strategies introduce overlap.

Example:

Chunk 1
--------
VPN credentials expire every 90 days.
VPN renewals require manager approval.
Chunk 2
--------
VPN renewals require manager approval.
Contractors require security review.

Overlap helps preserve context between chunks.


Karyam automatically handles:

  • Parsing
  • Chunk generation
  • Metadata creation
  • Embedding generation
  • Storage

Users do not manually create chunks.

This makes ingestion pipelines fully automated.


Chunk generation is fully observable through:

  • RAG Runs
  • Chunk Metadata
  • Embedding Logs
  • Processing Status

Example:

vpn-policy.pdf
↓
32 Chunks Generated
↓
32 Embeddings Created
↓
Stored In Qdrant
↓
RAG Run Completed

Employee:
"My VPN expired."
Agent:
↓
Generate Query Embedding
↓
Search Vector Database
↓
Retrieve Relevant Chunks
↓
Generate Response

Retrieved chunk:

VPN credentials expire every 90 days and require manager approval for renewal.

Documents with clear headings and sections produce better chunks.


Avoid splitting related information across multiple files when possible.


Poor retrieval quality often originates from:

  • Poor chunk boundaries
  • Weak embeddings
  • Missing metadata

RAG Runs help diagnose:

  • Parsing failures
  • Missing chunks
  • Embedding errors
  • Retrieval issues

Files
↓
Chunking
↓
Embeddings
↓
Vector Database
↓
Retrieval
↓
Agents

Chunking is the bridge between enterprise documents and AI understanding.


Continue with:

➡️ Retrieval

Learn how Karyam performs semantic search and delivers context to AI systems in real time.