Sapling Logo

Content Safety Checker

Score text for toxicity, harassment, hate speech, and other unsafe content.




Using this Tool

Content moderation is the task of identifying text that is unsafe or inappropriate for a given audience. This tool scores text on seven categories—toxicity, profanity, harassment, hate speech, self-harm, sexual content, and violence—each with a probability from 0 to 1.

Type or paste the text you would like to check in the input box above, then click the Check content button. Categories scoring at or above the flagging threshold (0.5) are highlighted in red.

Common uses include moderating user-generated content such as comments, reviews, and chat messages, and screening LLM outputs before they reach users. For word-level profanity flagging, see our profanity filter.

Want to integrate this into your own application? Check out the content safety API or contact us.