Large chat groups turn toxic fast, and human moderators cannot read every message at 2 a.m. This detector scores each post for harassment, hate, and threats the moment it lands, then flags the ones that cross your line.
Large chat groups turn toxic fast, and human moderators cannot read every message at 2 a.m. This detector scores each post for harassment, hate, and threats the moment it lands, then flags the ones that cross your line.
Moderators get an alert with the message context and a suggested action instead of a raw firehose. Repeat offenders are logged, so a pattern of borderline messages becomes visible rather than slipping by one at a time.
How it runs
Used at step 01 to kick off the pipeline.
Write
Used at step 01 to kick off the pipeline.
WebFetch
Used at step 01 to kick off the pipeline.
WebSearch
Used at step 01 to kick off the pipeline.
Score each message for toxicity, harassment, and threats against your thresholds.
pending
Flag messages over the line and notify moderators with the context and a suggested action.
pending
Log offenders and repeat patterns so moderators can spot recurring problem accounts.
pending
Sample output
json
// Sample output
// (generated when the pipeline finishes)
Score this message for toxicity from 0 to 1 across harassment, hate, and threats; return JSON with reasons.
Unlock the rest
The full agent definition, install snippet, and starter task are gated for community members.
Members get the full `.md` agent file, the npm / pnpm install one-liners, a starter prompt that we've tuned against real runs, and the open-source repo when this automation ships there. One email, magic link, done.