community-automations/data-forge-url

Investigación e Inteligencia

PublicSubagente de Claude

Extractor de datos por URL

Sacar datos estructurados de una página web suele implicar escribir un scraper a medida que se rompe en cuanto el sitio cambia. Este extractor toma en su lugar una URL y un esquema.

sonnet4 díasFirecrawlPlaywrightOpenAIPostgres
ClaudeClaude
ROI for
README.md

Por qué este subagente

Sacar datos estructurados de una página web suele implicar escribir un scraper a medida que se rompe en cuanto el sitio cambia. Este extractor toma en su lugar una URL y un esquema.

Tú describes los campos que quieres y él renderiza la página, mapea los elementos y devuelve registros limpios y validados. Cambias de sitio cambiando el esquema; sin escribir un scraper nuevo, y cualquier campo que falte en la página se marca en vez de inventarse.

Cómo se ejecuta

    • Read

      Used at step 01 to kick off the pipeline.

    • Write

      Used at step 01 to kick off the pipeline.

    • WebFetch

      Used at step 01 to kick off the pipeline.

    • WebSearch

      Used at step 01 to kick off the pipeline.

Salida de ejemplo

json
// Sample output
// (generated when the pipeline finishes)

Given a URL and a target schema, extract matching records as JSON conforming to the schema; set missing fields null and never invent values.

Unlock the rest

The full agent definition, install snippet, and starter task are gated for community members.

Members get the full `.md` agent file, the npm / pnpm install one-liners, a starter prompt that we've tuned against real runs, and the open-source repo when this automation ships there. One email, magic link, done.