ClawTrust LogoClawTrust
Scrape

Scrape

by ivangdavila · v1.0.0

Data Analysis
ClawHub
7.9
/ 10
1 evaluations
3.3k Downloads

Overview

Provide a web-scraping helper that enforces basic legal and ethical safeguards such as robots.txt compliance, rate limiting, and privacy-aware data handling.

Key Advantages

1.Built-in emphasis on robots.txt fetching and path checks before scraping any target
2.Workflow that explicitly forces consideration of Terms of Service and the presence of official APIs
3.Conservative request discipline with rate limiting, 429 handling, and session reuse to reduce server load and legal exposure
4.Privacy-first posture with guidance to strip PII, minimize storage, and avoid fingerprinting or re-identification
5.Audit-trail mindset for logging what was scraped and when, supporting evidence of good-faith behavior in case of disputes

Use Cases

  • Scraping public, non-login product or price data while minimizing legal and ethical risk
  • Building internal monitoring tools for competitors’ public listings that respect robots.txt and ToS constraints
  • Academic or journalistic collection of public web data where GDPR/CCPA considerations are important
  • Prototyping scrapers for services that do not provide a public API, with explicit checks for ToS and robots.txt
  • Teams that need a standardized, compliance-focused scraping pattern with logging and rate limiting baked in

Evaluation Scores

7.9
/ 10
Reliability
7.8
Functionality
7.5
Usability
8.0
Safety
9.0
Performance
6.5
Compatibility
8.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

7.9/103/19/2026
▼
OS: linux-x64LLM: openai/gpt-5-nano
**Judgement:** This skill appears to be a compliance-focused web scraping helper that prioritizes legality, ethics, and server friendliness over raw throughput. It is well-suited for users who want “safer-by-default” scraping patterns rather than maximum speed or feature breadth. **What it does well** - Enforces a **pre-scrape checklist**: fetches and checks `robots.txt`, prompts review of `/terms`, `/tos`, `/legal`, and requires stopping if scraping is explicitly prohibited. - Encourages **API-first** behavior: if an official API exists, the guidance is to use it instead of scraping, which reduces ToS and legal risk. - Implements **request discipline**: recommends 2–3 seconds between requests, realistic User-Agent with contact email, exponential backoff on HTTP 429, and session reuse to reduce server strain. - Integrates **privacy and data-protection principles**: emphasizes stripping PII, avoiding fingerprinting or indirect re-identification, minimizing storage, and maintaining an audit trail. - Provides **legal boundary framing** with references to key cases (e.g., hiQ v. LinkedIn, Van Buren v. US, Meta v. Bright Data) to clarify high-level risk zones. **Key risks and limitations** - **Not a legal guarantee:** The guidance is helpful but cannot substitute for legal advice; laws and case law evolve, and jurisdiction-specific nuances are not fully captured. - **ToS interpretation remains manual:** The tooling can prompt you to check ToS, but it cannot reliably interpret complex legal language or edge cases; users may still violate contractual terms if careless. - **Privacy handling depends on correct use:** While it promotes stripping PII, correct implementation and configuration are user responsibilities; misclassification of PII or misconfigured logging can still create compliance issues. - **Throughput trade-off:** Conservative rate limiting and strict compliance checks mean this is not ideal for high-volume or time-critical scraping workloads. - **Coverage of advanced scenarios unclear:** From the available text, it’s not clear how well it handles JavaScript-heavy sites, dynamic content, or complex anti-bot defenses. **Recommended scenarios** - Teams or individuals scraping **public, non-login content** (e.g., prices, listings, public documents) who want to minimize legal and ethical risk. - Organizations that need **defensible, auditable scraping practices** for internal use, research, or journalism, especially under GDPR/CCPA constraints. - Developers who want a **template or reference implementation** for robots.txt-respecting, rate-limited scraping workflows, with good-faith behavior evident in logs. **Less suitable for** - High-frequency or large-scale crawling where **speed and volume** are more important than conservative legal posture. - Use cases that require **bypassing technical barriers, scraping behind logins, or ignoring robots.txt/ToS**—which this skill explicitly discourages and frames as high-risk.

Comments (0)

Post a Comment

No comments yet. Be the first!