Ein Ordner gescannter PDFs ist nutzlos, bis jemand ihn in eine Tabelle tippt. Dies schickt den ganzen Stapel durch OCR und strukturierte Extraktion in eine einzige CSV.
Ein Ordner gescannter PDFs ist nutzlos, bis jemand ihn in eine Tabelle tippt. Dies schickt den ganzen Stapel durch OCR und strukturierte Extraktion in eine einzige CSV.
Es liest jede Datei, bildet sie auf Ihr Spalten-Template ab, prüft die Formate und stapelt die Ergebnisse in einer sauberen CSV. Fehlgeschlagene Dateien werden separat aufgeführt, damit nichts verloren geht.
So läuft er
Used at step 01 to kick off the pipeline.
Write
Used at step 01 to kick off the pipeline.
WebFetch
Used at step 01 to kick off the pipeline.
WebSearch
Used at step 01 to kick off the pipeline.
Wendet Ihr Feld-Template an, sodass jedes Dokument auf dieselben Spalten abgebildet wird.
pending
Prüft Typen und Formate wie Daten und Währung, bevor eine Zeile akzeptiert wird.
pending
Hängt die strukturierten Zeilen an eine einzige CSV an und listet alle Dateien, deren Extraktion fehlschlug.
pending
Beispiel-Output
json
// Sample output
// (generated when the pipeline finishes)
Extract the template fields from each document into a consistent CSV row; output null where a field is absent.
Unlock the rest
The full agent definition, install snippet, and starter task are gated for community members.
Members get the full `.md` agent file, the npm / pnpm install one-liners, a starter prompt that we've tuned against real runs, and the open-source repo when this automation ships there. One email, magic link, done.