Pozyskiwanie ustrukturyzowanych danych ze strony internetowej zwykle oznacza pisanie jednorazowego scrapera, który psuje się w chwili, gdy strona się zmieni. Ten ekstraktor działa inaczej — podajesz mu URL i schemat.
sonnet4 dniFirecrawlPlaywrightOpenAIPostgres
Claude
65ROI
72Scale
$2.2k93Saved
ROI for
README.md
Dlaczego ten subagent
Pozyskiwanie ustrukturyzowanych danych ze strony internetowej zwykle oznacza pisanie jednorazowego scrapera, który psuje się w chwili, gdy strona się zmieni. Ten ekstraktor działa inaczej — podajesz mu URL i schemat.
Opisujesz pola, których potrzebujesz, a on renderuje stronę, mapuje elementy i zwraca czyste, zwalidowane rekordy. Zmieniasz strony, zmieniając schemat — bez pisania nowego scrapera, a każde pole, którego na stronie brakuje, zostaje oznaczone, a nie zmyślone.
Jak działa
Used at step 01 to kick off the pipeline.
Write
Used at step 01 to kick off the pipeline.
WebFetch
Used at step 01 to kick off the pipeline.
WebSearch
Used at step 01 to kick off the pipeline.
Na podstawie zdefiniowanego schematu odwzoruj elementy strony na jego pola.
pending
Wyodrębnij dane, wymuś odpowiednie typy i zweryfikuj każdy rekord względem schematu.
pending
Zwróć czyste, ustrukturyzowane rekordy i oznacz pola, których strona nie dostarczyła.
pending
Przykładowe wyjście
json
// Sample output
// (generated when the pipeline finishes)
Given a URL and a target schema, extract matching records as JSON conforming to the schema; set missing fields null and never invent values.
Unlock the rest
The full agent definition, install snippet, and starter task are gated for community members.
Members get the full `.md` agent file, the npm / pnpm install one-liners, a starter prompt that we've tuned against real runs, and the open-source repo when this automation ships there. One email, magic link, done.