WEB INGESTION

Turn the web into fresh company knowledge

We map pages, render JavaScript listings static crawlers miss, and watch URLs on a schedule you control, down to hourly.

  • URL and sitemap ingest
  • AJAX page rendering for JS shells
  • Watch freshness on your schedule
  • Internal wikis, sites, YouTube, Substack
Request a demo
Web ingestion and URL watches

URL

Single pages or full trees

AJAX

Browser fallback for shells

Watch

Daily default, hourly minimum

Sources

Wikis, intranet, YouTube, Substack

Index the URLs that matter

Paste one page, a list of links, or a seed URL. We normalize duplicates, honor your include and exclude patterns, and turn each page into searchable, citable knowledge for your project.

  • Individual URLs and bulk link lists
  • Include and exclude path filters (globs)
  • Query-aware URLs when filters live in the address bar

Map a site from a seed URL

Give us a starting URL and we walk the site within depth and volume limits you set. Sitemap parsing runs alongside link extraction, so you get published maps and pages that only show up in navigation.

  • Recursive link follow with max depth and max URL caps
  • Sitemap parsing in parallel with page-link extraction
  • Optional subdomain inclusion when you need the full property

Render pages that static crawlers miss

Listing pages that load workshops, products, or docs after JavaScript often return an empty shell to a plain HTTP fetch. We detect that shell and retry through a real browser session so AJAX-populated content gets ingested.

  • Shell-page detection for filter widgets and empty listings
  • Browser rendering fallback with wait time for client-side loads
  • Built for public sites where the inventory lives behind JS, not in the first HTML response

Watch URLs on your schedule

Watch is not a monthly full-site recrawl. Attach a watch to a URL, pick Discovery, Content, or Both, and set how often we poll. Default is daily. Minimum is hourly. Content mode checks for changes before re-ingest, so unchanged pages do not burn credits.

  • Modes: Discovery (new pages), Content (seed changed), or Both
  • Poll interval you set: daily by default, hourly minimum
  • Content-hash change detection before re-ingest
  • Auto-ingest new pages when you want coverage without manual steps

Internal wikis and sites, plus channel watches

Web ingestion is not limited to public marketing pages. Point it at internal wikis, intranet sites, and company knowledge bases the same way you point it at the open web. Source-specific watches sit beside those URL watches for YouTube and Substack.

  • Internal wikis and intranet sites via URL
  • Company knowledge bases and docs sites
  • YouTube channel and playlist watches
  • Substack publication watches