VIDEO TUTORIAL
Build a RAG Chatbot for Equipment Manuals (Video)
Watch LangChain's overview of retrieval-augmented generation, then build a manual-reading chatbot that cites the page, refuses what it cannot find, and resists injection.

This RAG chatbot tutorial builds an assistant that answers questions from PLC and drive manuals, cites the page, and refuses when the manual is silent. Watch the short overview of retrieval-augmented generation first. Then follow the written steps: chunking, embeddings, retrieval, prompting, evaluation, and the security checks that stop the bot leaking data or being hijacked.
In this lesson by LangChain, the first part of the RAG From Scratch series gives the overview: indexing, retrieval and generation. LangChain publishes notebooks for the series on GitHub. The lesson opens our free RAG and Chatbots with LangChain course, where the first task is to write a failure-by-stage table for an assistant over your own company's manuals. Follow along, then check your work against these steps.
What you are building
A language model does not know what is in your drive manual, and when asked it may invent a parameter number with complete confidence. Retrieval-augmented generation fixes that in three stages. You index the manuals once, retrieve the passages relevant to each question, and generate an answer from those passages only. Because the answer comes from retrieved pages, it can cite them, and an engineer can check the citation before acting.
Each stage fails in its own way. Bad chunking splits a table from its heading. Bad retrieval returns the wrong model's manual. Bad prompting lets the model fill gaps from memory. Knowing which stage failed is most of the debugging.

Step 1: load the manuals properly
Load each PDF with its metadata: manual title, revision and page number. Strip repeating headers and footers, or every chunk will be half boilerplate. Tables need care: a parameter table read as flowing text loses the link between a number and its meaning, so convert tables row by row. Scanned pages need OCR before any of this works.
Step 2: chunk with a reason
In LangChain, the recursive character text splitter is the usual starting point. Try chunks of about 800 characters with 120 characters of overlap, and carry the page number and section heading into each chunk's metadata. Print three chunks and read them. If a chunk cannot be understood on its own, it will not be retrieved well either.
Step 3: embeddings and a vector store
Turn each chunk into an embedding with a sentence-transformers model and store the vectors in FAISS. Both run free in Colab or on a laptop. Embeddings find passages by meaning, which is what you want for "how do I set the ramp-up time?".
They are weaker on exact strings. Fault codes, parameter numbers and part numbers are better found by keyword search. Hybrid search runs both and merges the results, and it matters more for manuals than for most documents.
Step 4: retrieve, then filter
Retrieve the top few chunks for each question. Filter by manual and revision where the user has told you the equipment, so a question about one drive family does not get answers from another. Re-ranking the retrieved chunks is a later improvement. Get the basics measured first.
Step 5: prompt for grounded answers
The prompt does three jobs. Answer only from the supplied context. Cite the manual and page for every claim. If the context does not contain the answer, say so. Also tell the model that the retrieved text is reference material, not instructions. That one line matters in the security section below.
Step 6: evaluate on a test set
Write about 30 real questions, each with the page that answers it. Measure three numbers:
- Recall at k: how often the right page is among the retrieved chunks.
- Correctness: how often the final answer is right.
- Refusal rate: how often the bot rightly says the manual does not cover it.
Change one variable at a time (chunk size, k, the embedding model, the prompt) and re-run all 30. Without this, every change is a guess.
Security: what can go wrong
A chatbot over plant documents is a new attack surface. The OWASP Top 10 for LLM Applications 2025 names the risks. Four apply directly here:
- LLM01 Prompt Injection. A document can carry hidden instructions, such as white text in a PDF, that the model follows when the chunk is retrieved. This is indirect injection, and the user never sees it.
- LLM02 Sensitive Information Disclosure. If confidential manuals and public ones share one index, the bot can quote either to anyone.
- LLM07 System Prompt Leakage. Assume the system prompt can be extracted. Never put credentials or internal notes in it.
- LLM08 Vector and Embedding Weaknesses. Access control has to apply to chunks before retrieval, not to answers after it.
Misinformation (LLM09) matters too: a wrong torque value or a missed isolation step is a safety problem, not a style problem.

Failure patterns you will meet
Three problems come up on almost every manual chatbot. The answer comes from the wrong revision, because the old and new manuals are both indexed: add revision to the metadata and filter on it. A number is quoted without its unit or its table heading, because chunking cut the table: convert tables row by row with the heading repeated. The bot answers a question the manual does not cover, because the prompt let it: tighten the refusal rule and add those questions to the test set.
Keep it private
For manuals you cannot send to a hosted model, run the whole pipeline locally: a local embedding model, FAISS on disk, and a model served by Ollama. The course's final module serves the assistant with Streamlit or FastAPI, showing its sources, and covers re-indexing when a manual is revised.
Build it in the free course
RAG and Chatbots with LangChain covers loading and chunking, embeddings and hybrid search, the RAG chain with memory, query translation, routing, re-ranking, evaluation, LangGraph agents and serving, on free models throughout. It needs Python and Generative AI and LLM Foundations first. Then take AI Security and the OWASP Top 10 for LLMs, which works through all ten risks one at a time. A free account saves your progress.
The courses are free in full. The optional EDWartens Certificate of Completion for these intermediate courses is a small one-off fee, with a code anyone can check at edwartens.com/verification. It is not a vendor certification or an accredited qualification.
Take the free course
Questions
What is a RAG chatbot?
A chatbot that retrieves passages from your own documents and gives them to a language model as context, so the answer comes from the manual rather than the model's memory, and can cite the page.
What chunk size should I use for PDF manuals?
Start around 800 characters with 120 of overlap, keep the page number in each chunk's metadata, then change one setting at a time and measure retrieval on your test set.
Can a RAG chatbot run without sending manuals to the cloud?
Yes. With a local embedding model, a local vector store such as FAISS and a model served by Ollama, the whole pipeline can run on your own machine.
What is indirect prompt injection in RAG?
Instructions hidden inside a document the chatbot retrieves, which the model may follow as if the user had typed them. OWASP lists prompt injection as LLM01 in its 2025 Top 10 for LLM Applications.
How do I know the chatbot is good enough?
Write about 30 real questions with the page that answers each, then measure how often retrieval finds that page, how often the answer is correct, and how often it rightly refuses.
Sources
Written by the EDWartens engineering team for general education. Product names are trademarks of their owners; mentioning them does not imply endorsement. Prices and terms of other providers were checked on the date shown and can change.





