community-automations/autonomous-crawler

Investigación e Inteligencia

PublicSubagente de Claude

Crawler autónomo

Rastrear un sitio entero a mano significa saltarte páginas, machacar el servidor o perder la cuenta de lo que ya descargaste. Este crawler gestiona todo el barrido.

sonnet4 díasPlaywrightPythonRedisPostgres
ClaudeClaude
ROI for
README.md

Por qué este subagente

Rastrear un sitio entero a mano significa saltarte páginas, machacar el servidor o perder la cuenta de lo que ya descargaste. Este crawler gestiona todo el barrido.

Desde una URL semilla descubre enlaces, se queda dentro del dominio y marca el ritmo de las peticiones para que el sitio esté tranquilo. Los fallos se reintentan con backoff, las URLs visitadas se controlan para evitar bucles y cada página se guarda con sus metadatos. Cuando la cola se vacía, recibes un informe de cobertura.

Cómo se ejecuta

    • Read

      Used at step 01 to kick off the pipeline.

    • Write

      Used at step 01 to kick off the pipeline.

    • WebFetch

      Used at step 01 to kick off the pipeline.

    • WebSearch

      Used at step 01 to kick off the pipeline.

Salida de ejemplo

json
// Sample output
// (generated when the pipeline finishes)

Crawl a single domain from a seed URL with rate limiting, retries, and dedup, storing each page and reporting coverage.

Unlock the rest

The full agent definition, install snippet, and starter task are gated for community members.

Members get the full `.md` agent file, the npm / pnpm install one-liners, a starter prompt that we've tuned against real runs, and the open-source repo when this automation ships there. One email, magic link, done.