WEB INGESTION

Turn the web into fresh company knowledge

We map pages, render JavaScript listings static crawlers miss, and watch URLs on a schedule you control, down to hourly.

  • URL and sitemap ingest
  • AJAX page rendering for JS shells
  • Watch freshness on your schedule
  • Internal wikis, sites, YouTube, Substack
Request a demo
Web ingestion and URL watches

URL

Single pages or full trees

AJAX

Browser fallback for shells

Watch

Daily default, hourly minimum

Sources

Wikis, intranet, YouTube, Substack

Index the URLs that matter

Paste one page, a list of links, or a seed URL. We normalize duplicates, honor your include and exclude patterns, and turn each page into searchable, citable knowledge for your project.

  • Individual URLs and bulk link lists
  • Include and exclude path filters (globs)
  • Query-aware URLs when filters live in the address bar

Map a site from a seed URL

Give us a starting URL and we walk the site within depth and volume limits you set. Sitemap parsing runs alongside link extraction, so you get published maps and pages that only show up in navigation.

  • Recursive link follow with max depth and max URL caps
  • Sitemap parsing in parallel with page-link extraction
  • Optional subdomain inclusion when you need the full property

Render pages that static crawlers miss

Listing pages that load workshops, products, or docs after JavaScript often return an empty shell to a plain HTTP fetch. We detect that shell and retry through a real browser session so AJAX-populated content gets ingested.

  • Shell-page detection for filter widgets and empty listings
  • Browser rendering fallback with wait time for client-side loads
  • Built for public sites where the inventory lives behind JS, not in the first HTML response

Watch URLs on your schedule

Watch is not a monthly full-site recrawl. Attach a watch to a URL, pick Discovery, Content, or Both, and set how often we poll. Default is daily. Minimum is hourly. Content mode checks for changes before re-ingest, so unchanged pages do not burn credits.

  • Modes: Discovery (new pages), Content (seed changed), or Both
  • Poll interval you set: daily by default, hourly minimum
  • Content-hash change detection before re-ingest
  • Auto-ingest new pages when you want coverage without manual steps

Internal wikis and sites, plus channel watches

Web ingestion is not limited to public marketing pages. Point it at internal wikis, intranet sites, and company knowledge bases the same way you point it at the open web. Source-specific watches sit beside those URL watches for YouTube and Substack.

  • Internal wikis and intranet sites via URL
  • Company knowledge bases and docs sites
  • YouTube channel and playlist watches
  • Substack publication watches

How web ingestion compares to enterprise search tools

These platforms are strong on workplace connectors. Public web crawl control and freshness are where the differences show up. Ratings below are for website ingestion only.

Comparison as of July 13, 2026

CapabilityTricky WombatGleanGoSearchOnyxCoveoGuru
Index individual URLsYesYesPartialPartialYesNo
Crawl / link follow from a seedYesYesPartialPartialYesNo
Sitemap ingestYesYesPartialPartialYesNo
Include / exclude path filtersYesYesPartialPartialYesNo
JavaScript / AJAX page renderingYesPartialNoPartialPartialNo
Per-URL freshness watchesYesNoNoNoNoNo
User-set poll interval (hourly+)YesNoNoNoNoNo
Discovery + content change modesYesNoNoNoNoNo
Content-hash change detectionYesNoNoNoNoNo
YouTube / Substack source watchesYesNoNoNoNoNo
  • 1.Glean Website connector supports sitemap or seed crawl and optional Client-Side Rendering for public pages. Default web refresh is a 28-day full crawl (configurable via support), not per-URL watches.
  • 2.GoSearch, Onyx, Coveo, and Guru focus on workplace knowledge and connectors. TW-style Watch for public web crawl is not their main use case.
  • 3.JavaScript "partial" means optional CSR or limited dynamic support, not guaranteed AJAX listing recovery.

Legend: Yes = productized support · Partial = limited, optional, or workaround · No = not a documented product feature for website ingestion.

How web ingestion compares to website chatbot tools

These tools crawl a site to train a support agent. Most refresh on a weekly retrain schedule. None offer per-URL Watch with hourly polling and separate discovery and content modes.

Comparison as of July 13, 2026

CapabilityTricky WombatChatbaseSiteGPTCustomGPT.aiBotpressVoiceflow
Index individual URLsYesYesYesYesPartialPartial
Crawl / link follow from a seedYesYesYesYesPartialPartial
Sitemap ingestYesYesYesPartialPartialNo
Include / exclude path filtersYesYesPartialPartialPartialNo
JavaScript / AJAX page renderingYesNoNoPartialNoNo
Per-URL freshness watchesYesNoNoNoNoNo
Scheduled retrain / refreshYesPartialPartialPartialNoNo
User-set poll interval (hourly+)YesNoNoNoNoNo
Discovery + content change modesYesNoNoNoNoNo
YouTube / Substack source watchesYesNoNoNoNoNo
  • 1.Chatbase Auto Retrain (Standard/Pro) refreshes website sources weekly and can find new links. Hobby plans require manual retrain. No documented headless JavaScript crawl.
  • 2.SiteGPT and CustomGPT.ai are close Chatbase alternatives for website Q&A. Freshness is batch sync or retrain, not per-URL Watch.
  • 3.Botpress and Voiceflow are built for conversation design and channels. Website crawl is secondary.

Legend: Yes = productized support · Partial = limited, optional, or workaround · No = not a documented product feature for website ingestion.

Common questions before you crawl

Anything that starts from a URL you can reach: public pages, mapped site trees, sitemaps, internal wikis, intranet sites, and company knowledge bases, plus watched URLs that re-check on a schedule. YouTube and Substack watches use the same freshness model beside URL watches.

Watch attaches to a specific URL with a mode and interval you choose. Discovery mode looks for newly published pages. Content mode re-fetches the seed when its content hash changes. Default is daily. You can poll as often as hourly.

Yes, on many public sites. When a static scrape returns filter chrome without listing content, we retry through a browser session and wait for client-side requests to populate the page before extracting text and links.

Enterprise search tools index workplace apps and can crawl websites, but default web refresh often runs on a multi-week schedule and client-side rendering support is limited for advanced dynamic behavior. Website chatbot tools crawl sites to train an agent and usually refresh on a weekly retrain cadence. Tricky Wombat gives you controlled site mapping plus per-URL freshness on a schedule you set.

Index your sites. Keep them fresh.

Try URL ingest, JavaScript recovery, and scheduled watches on sites you choose, not in a slide deck.

Schedule a call