community-automations/autonomous-crawler

Nghiên cứu & Thông tin tình báo

PublicSubagent Claude

Trình thu thập dữ liệu tự động

Crawl thủ công cả một trang web đồng nghĩa với bỏ sót trang, bắn server quá tải, hoặc mất dấu những gì đã lấy. Crawler này quản lý toàn bộ quá trình quét.

sonnet4 ngàyPlaywrightPythonRedisPostgres
ClaudeClaude
ROI for
README.md

Tại sao chọn subagent này

Crawl thủ công cả một trang web đồng nghĩa với bỏ sót trang, bắn server quá tải, hoặc mất dấu những gì đã lấy. Crawler này quản lý toàn bộ quá trình quét.

Từ một URL gốc, nó phát hiện các liên kết, chỉ hoạt động trong domain, và điều tiết tốc độ request để không gây áp lực cho trang. Lỗi được thử lại với backoff, các URL đã truy cập được theo dõi để tránh lặp vòng, và mọi trang đều được lưu kèm metadata. Khi hàng đợi trống, bạn nhận được báo cáo mức độ bao phủ.

Cách vận hành

    • Read

      Used at step 01 to kick off the pipeline.

    • Write

      Used at step 01 to kick off the pipeline.

    • WebFetch

      Used at step 01 to kick off the pipeline.

    • WebSearch

      Used at step 01 to kick off the pipeline.

Ví dụ đầu ra

json
// Sample output
// (generated when the pipeline finishes)

Crawl a single domain from a seed URL with rate limiting, retries, and dedup, storing each page and reporting coverage.

Unlock the rest

The full agent definition, install snippet, and starter task are gated for community members.

Members get the full `.md` agent file, the npm / pnpm install one-liners, a starter prompt that we've tuned against real runs, and the open-source repo when this automation ships there. One email, magic link, done.