Add your website
Import a whole site by crawl, a sitemap or a list of links, and keep it in sync automatically.
Your website is usually the best first knowledge source: it is already written for customers and it is already public. PepoChat can import a single page, a list of links, a sitemap or a whole site.
Import options
| Option | What it does | Counts as |
|---|---|---|
| Single URL | Fetches one page. | 1 source |
| List of links | Up to 20 pages you paste, one per line. | 1 source per link |
| Sitemap | Every page listed in your sitemap.xml, including sitemap indexes that point to other sitemaps. | 1 source |
| Website crawl | Starts at the address you give and follows links on the same site, breadth first, up to 50 pages and 3 levels deep. Include and exclude path filters narrow it down. | 1 source |
A whole-site crawl or a sitemap import counts as one knowledge source, which matters on the Free plan's five-source limit.
Import your site
Open Knowledge Base and click Add New
Choose the website option and paste your address, for example
https://www.example.com/.Pick a category (optional)
Categories such as "Product" or "Policies" help you find sources later. They do not change how the agent answers.
Start the import
The page shows Importing with a running count, for example "Website crawl · 18 of 28 pages". You can leave the page; the import continues in the background.
Check the result
Each imported page appears as a row with its address and size. Pages that could not be fetched are skipped.
What the crawler respects
robots.txtDisallow rules are honoured.- Fetches are spaced out so they do not load your server.
- Pages larger than 5 MB are skipped.
- Private and internal addresses (localhost, private network ranges) are refused.
- Only pages on the same site as the start address are followed.
Keeping it in sync
Website pages are re-fetched every week and only pages whose content changed are re-indexed, so a price change on your site reaches the agent without you doing anything. Two buttons let you do it sooner:
- Sync website pages, at the top of the Knowledge Base, re-fetches every website page you have right now. A card shows the progress ("Website sync · 18 of 40 pages · 3 updated"), and unchanged pages cost nothing.
- Import again, on a finished crawl or sitemap card, runs that import again with the same settings. Use it when you have added new pages, since a sync only refreshes pages the agent already knows. The pages stay grouped under the original import, so the site still counts as one knowledge source. If the card is gone, simply start a new crawl of the same site from Add New: a site that is already in your knowledge base is treated as a refresh, not a second source.
Only one import or sync runs at a time per workspace, and you can cancel either from its card.
Tip
If a page is behind a login, a cookie wall or a JavaScript-only render, the crawler sees what an anonymous visitor's first request sees. Upload such content as a document instead.
Limits
Sources and text volume count against your plan: 5 sources and 20 MB of text on Free, unlimited sources with 200 MB on Starter and 1 GB on Growth. The Knowledge Base page shows your usage. See Manage sources.
Something missing or wrong on this page? Tell us and we will fix it.
