2.3k Downloads
Overview
High-performance web scraping, Google search, and website crawling via the AnyCrawl hosted API, returning LLM-ready structured content and optional JSON-extracted data.
Key Advantages
1.End-to-end workflow: supports single-page scrape, Google search, full-site crawling, status/results retrieval, and cancellation.
2.Multiple scraping engines (cheerio, playwright, puppeteer) to balance speed vs. JavaScript rendering needs.
3.Rich extraction controls (include/exclude tags, formats, wait_for selector, proxy, json_options with schema+prompt) for precise data shaping.
4.Built-in structured Google search with optional auto-scraping of results for one-shot research flows.
5.Asynchronous crawl jobs with domain/hostname/origin strategies, depth/limit controls, and path-based include/exclude/scrape filters for targeted site coverage.
Consistent API response envelope ({"成功"
Use Cases
- Research and RAG ingestion: search-and-scrape top results, then feed markdown/json into vector stores for QA or agents.
- Competitive/market monitoring: schedule crawls of competitor sites or documentation with include_paths/scrape_paths filters.
- Product/catalog extraction: use json_options schema to reliably capture product_name, price, specs across many product pages.
- Technical documentation indexing: crawl docs sites (same-domain, max_depth, limit) and convert to markdown for downstream search.
- SEO and content analysis: scrape blogs/news sites at scale, filtering tags to focus on headings and article body text only.
Evaluation Scores
8.6
/ 10
Reliability
8.2
Functionality
9.3
Usability
9.0
Safety
7.5
Performance
8.8
Compatibility
9.0
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.6/103/19/2026▼
OS: darwin-x64LLM: moonshotai/kimi-k2.5
**Verdict:** A powerful, production-grade web data acquisition skill that covers single-page scraping, Google search, and full-site crawling with strong configurability. Well-suited as a primary web I/O layer for agents and pipelines, assuming comfort with an external hosted API and web-scraping ToS considerations.
**Strengths**
- Broad coverage: `anycrawl_scrape`, `anycrawl_search`, async crawl (start/status/results/cancel), and `search_and_scrape` helper make it easy to build end-to-end web workflows.
- Flexible engines: `cheerio` for fast static HTML, `playwright`/`puppeteer` for JS-heavy SPAs.
- Fine-grained extraction: tag filters, formats (markdown/html/text/json/screenshot), selectors, proxy support, and `json_options` (schema + user_prompt) for structured outputs.
- Clear integration: single `ANYCRAWL_API_KEY` env var, consistent response and error format, well-documented parameters and examples.
**Risks / Limitations**
- **External dependency:** All functionality relies on the AnyCrawl SaaS; outages, rate limits, or credit exhaustion will directly impact agents.
- **Legal/ToS risk:** Automated scraping and Google search automation may conflict with some sites’ terms of service or robots policies; users must enforce their own compliance.
- **Cost & limits:** Crawl limits, rate limits, and job expiry (24h) can constrain large-scale or long-running operations; design agents to check error codes (402/429) and degrade gracefully.
**Recommended Scenarios**
- Agents that need reliable, structured ingestion of web pages or docs sites for RAG, monitoring, or summarization.
- Research/workflow tools that combine Google search with automatic scraping of top results.
- Focused crawlers for product catalogs, docs, or blogs where schema-based JSON extraction provides high value.
**Use with extra care** in contexts involving sensitive/PII data, heavy or repeated crawling of the same domains, or where strict compliance with site-specific access rules is required.
Comments (0)
No comments yet. Be the first!