PDF training

The manual nobody reads, finally answering questions

Upload the PDF. CustomerBot extracts the text — including from image-heavy pages — splits it into passages, indexes it, and answers from it with the document named. A download link becomes an answer.

Four documents worth uploading today

Product manuals

“How do I reset it to factory settings?”

Step-by-step content retrieves particularly well because each procedure is self-contained.

Price and fee tables

“What does the mid tier cost annually?”

Tabular data is extracted as text, so the numbers stay attached to their row labels.

Terms and policies

“Can I cancel in the first fourteen days?”

The clause is quoted back in plain language, with the document named.

Spec sheets

“Is it rated for outdoor use?”

Dense reference data that nobody wants to read but everybody wants to query.

What happens to the file after you drop it in

  1. 01Upload

    Drop the file into the training screen. It joins the same knowledge base as your crawled pages, text snippets and Q&A pairs — one index, not four.

  2. 02Extract

    The text layer is pulled out. Pages that are mostly image — scanned tables, diagrams, screenshots of a form — go through image extraction so their content is recovered rather than silently dropped.

  3. 03Chunk

    The document is split into overlapping passages. A 90-page manual is not one blob; it becomes hundreds of retrievable pieces, each small enough to be specific.

  4. 04Embed and index

    Each passage becomes a vector in the same index as everything else, so a question can be answered from a PDF and a web page in the same reply.

  5. 05Answer with the source

    The reply names the document it drew on, so anyone can check it against the original.

What extracts well, and what does not

PDF is a printing format, not a data format. Being honest about that saves you from wondering why one document works beautifully and another does not.

Extracts cleanlyExpect trouble
A born-digital PDF exported from Word, InDesign or a CMSA photo of a page taken on a phone at an angle
Selectable text you can highlight in a PDF readerA flattened scan with no text layer at all
Tables with clear headers and consistent columnsMulti-column magazine layouts that interleave when read linearly
One document per subjectA 400-page compendium of everything you have ever published

One index, four kinds of source

PDFs are not a separate mode. They sit in the same knowledge base as crawled URLs, pasted text and Q&A pairs, which means a single answer can draw on your terms PDF and your pricing page at once — and cite both.

PDF questions

What happens to scanned PDFs?

Image-heavy pages are put through extraction so text inside them can be recovered rather than skipped. Quality depends on the scan: a clean, straight, high-resolution scan does well, a crooked phone photo does not. If a document matters and the scan is poor, the fastest fix is to paste the text directly as a text source.

Can it answer from tables?

Yes, within reason. Tables are extracted as text, which preserves the relationship between a row label and its values for most well-formed tables. Very wide or deeply nested tables can lose structure — if a fee table is critical, it is worth checking a couple of answers against the original before you rely on it.

How large a PDF can I upload?

The practical limit is your assistant’s storage: 5 MB on the free plan and 40 MB on paid plans. Because documents are stored as extracted text rather than as the original file, that covers a great deal more paper than the numbers suggest.

Can I mix PDFs with website content?

Yes, and you generally should. A single assistant can be trained on crawled URLs, uploaded PDFs, pasted text and Q&A pairs at once. Retrieval draws from all of them, so a question can be answered from a policy PDF and a pricing page in the same reply.

What if I update the document?

Replace it in the training screen. The old passages go with it, so there is no risk of the assistant quoting version 3 after you have published version 4 — which is precisely the failure mode of leaving old PDFs on a website.

Does it answer in other languages?

Yes. It detects the language of the question and replies in it across 95+ languages, reading from the source document in whatever language that was written.

Upload the document you keep emailing people

Then ask it the question you get asked about that document. That is the whole test, and it takes two minutes.

Start free