Skip to content

How to Build a RAG Chatbot Over Your Company Documents, Step by Step

AI7 min readBy the Nexzem team

A practical, step-by-step path from a folder of PDFs and wiki pages to a chatbot that answers accurately, cites sources and respects permissions.

In this article
  1. 01What you are building
  2. 02Step 1: Define scope and build an evaluation set first
  3. 03Step 2: Collect and clean documents
  4. 04Step 3: Split documents into chunks
  5. 05Step 4: Create embeddings and an index
  6. 06Step 5: Retrieve well, because generation cannot fix bad retrieval
  7. 07Step 6: Generate grounded answers with citations
  8. 08Step 7: Evaluate and iterate
  9. 09Common failure modes and how to fix them
  10. 10Step 8: Prepare for production
  11. 11Where to go next

What you are building

A RAG chatbot answers questions using your own documents instead of only what a language model learned in training. It has two halves. The ingestion pipeline turns documents into searchable chunks and stores them in an index. The query pipeline takes a question, retrieves the most relevant chunks and asks a language model to answer using only those chunks, with citations.

This pattern, retrieval-augmented generation, is usually the right first step for company knowledge because documents change often and answers must be traceable. Fine-tuning changes a model's behavior but is a poor way to teach it facts that change weekly; our RAG vs fine-tuning comparison covers the difference.

Step 1: Define scope and build an evaluation set first

Before writing code, collect 50 to 100 real questions people ask, such as from support tickets, HR inboxes or sales calls, and write the correct answer and source document for each. This evaluation set is how you will know whether a change helped. Without it, teams tune prompts by gut feeling and quality drifts.

Also decide who can ask what. If some documents are confidential to finance or HR, permissions must be enforced in retrieval from the first version, not added later.

Step 2: Collect and clean documents

Most effort in a RAG project goes into turning messy files into clean text. PDFs, Word files, slide decks, wiki pages and spreadsheets each need their own parser. Libraries such as pypdf, Unstructured and Apache Tika handle common formats; scanned documents need OCR, for example with Tesseract or a cloud document AI service.

  • Strip repeated headers, footers, page numbers and navigation text
  • Keep tables as structured text or Markdown rather than flattened words
  • Record metadata for every document: title, source URL, owner, last updated date and access group
  • Skip or flag outdated versions so the bot does not quote old policies

Step 3: Split documents into chunks

Retrieval works on chunks, not whole documents. Split by structure first, on headings and sections, then by length so each chunk is a few hundred tokens. A small overlap between neighboring chunks keeps sentences that span a boundary retrievable. Prefix each chunk with its document title and heading path, such as "Leave Policy > Parental leave > Eligibility", so the chunk makes sense on its own.

Chunk size is a trade-off. Small chunks retrieve precisely but lose context; large chunks carry context but dilute relevance and cost more tokens. Test two or three sizes against your evaluation set rather than guessing. Our LLM token estimator helps you check how much text fits in a prompt.

Step 4: Create embeddings and an index

An embedding model turns each chunk into a vector, a list of numbers where similar meanings land close together. Hosted embedding models from OpenAI, Cohere or Google work well, and open models from the sentence-transformers ecosystem can run on your own servers. Use the same embedding model for documents and questions, and re-embed everything if you change models. See our explainer on embeddings for how this works.

Store vectors with their text and metadata in a vector database. If you already run PostgreSQL, the pgvector extension is often enough to start. Dedicated options such as Qdrant, Weaviate, Pinecone or OpenSearch add features for larger collections. Make sure the store supports metadata filters, since you will need them for permissions.

Step 5: Retrieve well, because generation cannot fix bad retrieval

Pure vector search misses exact terms like product codes, policy numbers and names. Hybrid search combines vector similarity with keyword search (BM25) and merges the results, for example with reciprocal rank fusion. Then a reranker, a model that scores each question and chunk pair directly, reorders the top 20 to 50 candidates so the best 5 or so reach the prompt.

  • Apply permission filters in the retrieval query, based on the signed-in user, never in the prompt
  • Rewrite follow-up questions into standalone queries using the chat history
  • Filter by metadata such as region or product when the question implies it
  • Measure retrieval separately: is the right chunk in the top results?

Step 6: Generate grounded answers with citations

The prompt should tell the model to answer only from the provided context, to cite the chunk IDs it used and to say it does not know when the context is insufficient. Pass each chunk with an ID and its source title, then turn the cited IDs into links in the interface. Use a low temperature for factual answers and stream the response so users see progress.

Treat retrieved text as data, not instructions. A document can contain text that tries to change the bot's behavior, a form of prompt injection, so keep system instructions separate and never give the chatbot tools with side effects based only on document content.

Step 7: Evaluate and iterate

Run the evaluation set after every meaningful change. Track retrieval hit rate (was a correct source in the top results), answer correctness and faithfulness (does the answer stick to the sources). Frameworks such as Ragas and LLM-as-judge scoring speed this up, but review a sample by hand, since automated judges make mistakes too. Add every real failure from production to the evaluation set.

Common failure modes and how to fix them

When answers go wrong, the cause is usually one of a few patterns. Diagnose from the logged trace, which shows the question, the retrieved chunks and the answer, before changing anything:

  • The right document was never ingested or was parsed badly: fix the parser, not the prompt
  • The right chunk exists but ranks low: add hybrid search, a reranker or clearer chunk titles
  • The answer blends two versions of a policy: add metadata filters for region, product or effective date
  • The model answers beyond its sources: tighten the instructions and require a citation for every claim
  • Follow-up questions fail: rewrite them into standalone queries before retrieval

Step 8: Prepare for production

A production RAG system needs more than a good demo. Plan for these from the start:

  • Incremental re-indexing when documents change, and deletion when they are removed
  • Logging of questions, retrieved chunks and answers, with personal data redacted
  • Thumbs up and down feedback linked to the logged trace
  • Rate limits and cost tracking per user or team
  • A clear hand-off to a human for questions the bot cannot answer

Where to go next

Start with one document collection and one team, prove accuracy against the evaluation set, then expand. Most quality gains come from better parsing, chunking and retrieval rather than a bigger model. If you want help taking a prototype to production, Nexzem's RAG development service covers ingestion, retrieval, evaluation and hosting.

Planning something similar?

Get a straight answer on scope, cost and timeline.

Talk to the team

Tell us what you're building.

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.