WEB INGESTION
Turn the web into fresh company knowledge
We map pages, render JavaScript listings static crawlers miss, and watch URLs on a schedule you control, down to hourly.
- URL and sitemap ingest
- AJAX page rendering for JS shells
- Watch freshness on your schedule
- Internal wikis, sites, YouTube, Substack

URL
Single pages or full trees
AJAX
Browser fallback for shells
Watch
Daily default, hourly minimum
Sources
Wikis, intranet, YouTube, Substack
START FROM ANY PAGE
Index the URLs that matter
Paste one page, a list of links, or a seed URL. We normalize duplicates, honor your include and exclude patterns, and turn each page into searchable, citable knowledge for your project.
- Individual URLs and bulk link lists
- Include and exclude path filters (globs)
- Query-aware URLs when filters live in the address bar
PAGES YOU DID NOT PASTE YOURSELF
Map a site from a seed URL
Give us a starting URL and we walk the site within depth and volume limits you set. Sitemap parsing runs alongside link extraction, so you get published maps and pages that only show up in navigation.
- Recursive link follow with max depth and max URL caps
- Sitemap parsing in parallel with page-link extraction
- Optional subdomain inclusion when you need the full property
AJAX AND FILTER BOARDS
Render pages that static crawlers miss
Listing pages that load workshops, products, or docs after JavaScript often return an empty shell to a plain HTTP fetch. We detect that shell and retry through a real browser session so AJAX-populated content gets ingested.
- Shell-page detection for filter widgets and empty listings
- Browser rendering fallback with wait time for client-side loads
- Built for public sites where the inventory lives behind JS, not in the first HTML response
KEEP IT FRESH
Watch URLs on your schedule
Watch is not a monthly full-site recrawl. Attach a watch to a URL, pick Discovery, Content, or Both, and set how often we poll. Default is daily. Minimum is hourly. Content mode checks for changes before re-ingest, so unchanged pages do not burn credits.
- Modes: Discovery (new pages), Content (seed changed), or Both
- Poll interval you set: daily by default, hourly minimum
- Content-hash change detection before re-ingest
- Auto-ingest new pages when you want coverage without manual steps
INSIDE AND OUTSIDE THE FIREWALL
Internal wikis and sites, plus channel watches
Web ingestion is not limited to public marketing pages. Point it at internal wikis, intranet sites, and company knowledge bases the same way you point it at the open web. Source-specific watches sit beside those URL watches for YouTube and Substack.
- Internal wikis and intranet sites via URL
- Company knowledge bases and docs sites
- YouTube channel and playlist watches
- Substack publication watches