Before you start
Before you start, sign in to the dashboard and select the organization you want to work in.- You need a knowledge base, or you can start one from the same dialog. See Create a knowledge base.
- Your role must allow resource management.
Start the import
1
Open the knowledge base and select Add more
The Add a source dialog opens. Its description reads Fetch a website or upload a file.
2
Select Import from website
The website tab shows the settings on the left and a preview pane on the right.
3
Enter the Website URL
Type the address of the site. Discovery starts shortly after you stop typing.
4
Review the discovery preview
The preview shows the site identity and the discovered pages as a tree.
5
Select the pages to import
Clear Include all discovered pages to pick pages, or uncheck whole sections.
6
Select Start fetching
The button reads Starting… while the import begins. The knowledge base page then shows its progress.
The Website URL field
The Website URL field validates as you type. The submit button stays disabled while the field is empty or shows one of these messages:
If you look up too many URLs in a row, a rate-limit note appears. It asks you to wait a minute, then edit the URL to try again.
Read the discovery preview
Discovery lists candidate pages before any content is imported. The right pane starts with Enter a website to see its pages and fills as discovery runs.
Discovery on a large sitemap can be slow. After a few seconds the preview notes This site has a large sitemap. Still collecting pages…
Three special cases change the preview:
- On a site with more than 1000 discovered pages, individual page picking is unavailable. Fetch all pages, or uncheck whole sections to exclude them.
- When discovery finds no sitemap and one page at most, the preview titles itself Pages are found as we fetch. Start fetching, and Tars follows links from the start URL to discover the rest.
- When nothing is found, the preview says No pages were found here.
Fetch settings
Changing a setting that affects discovery restarts the preview, so the pages you review always match the settings you submit.
If the page cap is more than the pages left on your plan, a warning under the two fields asks you to lower the max pages value. The submit button stays disabled until the selection fits.
Open Advanced settings for the rest. It is collapsed by default.
Turn on Render JavaScript for sites whose content only appears after the page runs in a browser. With Media search on, JavaScript rendering is also needed to capture videos on those sites. Set User agent when a site blocks unknown crawlers or allows only a named one.
Watch the import
While the fetch runs, a banner at the top of the Documents tab counts the scraped pages. Failed and skipped counts appear beside it when they occur. Select Cancel on the banner to stop the import. The document list fills in as pages are scraped, each row with a live status badge. See Manage documents for the list during an import. When the run ends, a toast reports the result as new, updated, and unchanged pages. Unchanged pages are detected by content hash, so a re-fetch never duplicates identical content.Schedule re-fetches
A schedule keeps imported pages current without manual work. The Re-fetch schedule options carry usage hints: Daily for blogs and news, Weekly for support centers, Monthly for policies and pricing. Change the schedule later from Sync management.Sync management
Select Sync management above the document list to open the sheet. The sheet holds the re-crawl schedules, the source settings, and the recent crawls. A knowledge base with only uploaded files has nothing on a schedule, and the sheet says so. The Re-crawl schedule section shows one website source at a time.
Both action buttons are disabled while a retriever builds, with the tooltip Locked while a retriever is building. A source that never imported successfully shows Import this website successfully before changing its sync settings.
The Fetch settings dialog
The dialog notes Applies on next sync for . It holds Max pages in the range 1 to 5000, Max depth, User agent, Render JavaScript, Include subdomains, and Exclude paths. Select Save, and a toast confirms Source fetch settings updated.Review crawl history
Below the schedule section, each crawl run of the selected source appears as a history row.
Expand a run to see its Failed and Skipped pages grouped by reason, with the first five URLs per reason. Page-level details are kept for 14 days, after which the row notes they expired.
Retry a failed crawl
When a fetch fails, a toast names the reason, such as a timeout or an unreachable site. Open Sync management and select Sync this source now to run the source again.Review a held re-fetch
A re-fetch does not update the knowledge base on its own in two cases. The run is held as Awaiting review, and the import banner shows a warning.
Select Review on the banner to open the Review pages dialog. Deselect pages, or add exclusions under Advanced settings, then confirm the fetch. Select Cancel to keep the current documents.
The confirm button states exactly what will run: Fetch included, Fetch up to of included, Fetch up to pages, or No pages to fetch when nothing is selected.
