Product manuals
“How do I reset it to factory settings?”
Step-by-step content retrieves particularly well because each procedure is self-contained.
Upload the PDF. CustomerBot extracts the text — including from image-heavy pages — splits it into passages, indexes it, and answers from it with the document named. A download link becomes an answer.
“How do I reset it to factory settings?”
Step-by-step content retrieves particularly well because each procedure is self-contained.
“What does the mid tier cost annually?”
Tabular data is extracted as text, so the numbers stay attached to their row labels.
“Can I cancel in the first fourteen days?”
The clause is quoted back in plain language, with the document named.
“Is it rated for outdoor use?”
Dense reference data that nobody wants to read but everybody wants to query.
Drop the file into the training screen. It joins the same knowledge base as your crawled pages, text snippets and Q&A pairs — one index, not four.
The text layer is pulled out. Pages that are mostly image — scanned tables, diagrams, screenshots of a form — go through image extraction so their content is recovered rather than silently dropped.
The document is split into overlapping passages. A 90-page manual is not one blob; it becomes hundreds of retrievable pieces, each small enough to be specific.
Each passage becomes a vector in the same index as everything else, so a question can be answered from a PDF and a web page in the same reply.
The reply names the document it drew on, so anyone can check it against the original.
PDF is a printing format, not a data format. Being honest about that saves you from wondering why one document works beautifully and another does not.
| Extracts cleanly | Expect trouble |
|---|---|
| A born-digital PDF exported from Word, InDesign or a CMS | A photo of a page taken on a phone at an angle |
| Selectable text you can highlight in a PDF reader | A flattened scan with no text layer at all |
| Tables with clear headers and consistent columns | Multi-column magazine layouts that interleave when read linearly |
| One document per subject | A 400-page compendium of everything you have ever published |
PDFs are not a separate mode. They sit in the same knowledge base as crawled URLs, pasted text and Q&A pairs, which means a single answer can draw on your terms PDF and your pricing page at once — and cite both.
Image-heavy pages are put through extraction so text inside them can be recovered rather than skipped. Quality depends on the scan: a clean, straight, high-resolution scan does well, a crooked phone photo does not. If a document matters and the scan is poor, the fastest fix is to paste the text directly as a text source.
Yes, within reason. Tables are extracted as text, which preserves the relationship between a row label and its values for most well-formed tables. Very wide or deeply nested tables can lose structure — if a fee table is critical, it is worth checking a couple of answers against the original before you rely on it.
The practical limit is your assistant’s storage: 5 MB on the free plan and 40 MB on paid plans. Because documents are stored as extracted text rather than as the original file, that covers a great deal more paper than the numbers suggest.
Yes, and you generally should. A single assistant can be trained on crawled URLs, uploaded PDFs, pasted text and Q&A pairs at once. Retrieval draws from all of them, so a question can be answered from a policy PDF and a pricing page in the same reply.
Replace it in the training screen. The old passages go with it, so there is no risk of the assistant quoting version 3 after you have published version 4 — which is precisely the failure mode of leaving old PDFs on a website.
Yes. It detects the language of the question and replies in it across 95+ languages, reading from the source document in whatever language that was written.
Then ask it the question you get asked about that document. That is the whole test, and it takes two minutes.
Start free