Description
WaterCrawl is a modern web crawling and content extraction platform that transforms any website into structured, LLM-ready data. Built for developers, it lets users crawl any site instantly with smart crawling controls for depth, domains, and paths, precise content extraction using custom selectors that filter out ads and footers, and JavaScript rendering to capture dynamic content and take screenshots in PDF or JPG.
Beyond crawling, WaterCrawl includes a web search engine for real-time results across the web, automatic sitemap generation to map all URLs on a site, and built-in OpenAI integration to turn raw HTML into structured, meaningful data. It offers an extensible plugin system, real-time monitoring of crawl status and performance via Server-Sent Events, and comprehensive error reporting. The project is open source under a modified MIT license, with commercial use as a competing service requiring permission.
Developers integrate WaterCrawl through a clean REST API with JWT-based authentication and webhook support, and official SDKs for Python, Go, Node.js, PHP, and Rust. The API supports starting single or batch crawl requests, configuring spider and page options, streaming live status, and downloading results as JSON or Markdown, plus dedicated search endpoints.
A free plan provides 1,000 page credits per month with datacenter proxies and a 7-day retention window. Paid Startup, Growth, and Enterprise plans add more page credits, higher crawl depth and concurrency, premium residential proxies, more team seats, and longer retention.
WaterCrawl's Core Features
Smart crawling controls for depth, domains, and paths
Precise content extraction with custom selectors
JavaScript rendering and PDF or JPG screenshots
Built-in web search engine with real-time results
Automatic sitemap generation and site mapping
AI-powered processing via OpenAI integration
Open source with an extensible plugin system
REST API, webhooks, and SDKs for multiple languages
How to use WaterCrawl?
Sign up free: Create an account and get 1,000 monthly page credits.
Get an API key: Generate a Team API key from the dashboard.
Configure a crawl: Set the start URL, depth, domains, and page options.
Run and monitor: Start the crawl and track progress in real time via SSE.
Retrieve data: Download results as JSON or Markdown, or receive them via webhook.
WaterCrawl's Use Cases
- LLM data pipelines
- Content aggregation
- Market research
- Search engines
- Website monitoring




