Companion and character apps
Long role-play sessions where the scene drifts, often in Japanese or Korean.
California SB 243 asks for reasonable measures so a companion bot doesn't produce sexual content for known minors.
Conversation safety for AI characters
One-message filters miss what builds up over a conversation. SceneGuard reads the scene, checks what your AI character is about to say, and steers it back in character instead of breaking the moment or refusing flatly.
We found these on our own live app, in English, Japanese and Korean.
"Keep going", "here", "let me see": each scores as everyday talk. Read inside the scene, it isn't. On one live thread we studied, most of the lines that slipped through had been rated everyday on their own.
The user writes something vague. The model fills in the rest. Checking only what users type misses the one message you actually published.
Blocking the turn breaks the story and the user leaves. Steering keeps the scene going on safe ground, in the character's own voice.
A filter tuned on coded content starts refusing the innocent version: a back rub for someone sore after work. SceneGuard is measured on both sides.
One judge reads the scene beside your model's reply, so safe turns stream without a wait.
Rates the newest message inside the conversation it belongs to, and remembers threads that keep pushing so the next vague line is read in that light.
Holds the character's draft when the scene is near or past your line, and only shows it if it stays in bounds.
Replaces a failed draft with a line in the character's own voice that moves the scene somewhere safe. No lectures, no flat no.
Tells "pretend we're in high school" apart from a memory of school, and switches the character to a friendly mode.
Coded slang, 반말 role-play habits, parenthesized actions: tested on real Japanese and Korean conversations, not translated test sets.
Set where the line sits for each product: all-ages, teen, PG-13 or adults-only with a hard floor on minors.
Any product where an AI plays a character and talks to the public for many turns.
Long role-play sessions where the scene drifts, often in Japanese or Korean.
California SB 243 asks for reasonable measures so a companion bot doesn't produce sexual content for known minors.
NPCs that hold open conversations with players, many of them teens.
California's game exemption covers bots that can't discuss sexual content, self-harm or mental health. A reply check is how you show it.
Conversations where a crisis or a slide into "therapy" only shows up across several turns.
Illinois, Nevada and Utah restrict AI mental-health chat; California SB 243 requires crisis protocols.
Rules for companion chatbots are arriving state by state and country by country.
New York companion law in force: AI disclosure and a self-harm protocol.
California SB 243 in force, with a private right of action of $1,000 per violation.
Australia's eSafety codes: chatbots must stop sexual and self-harm content for minors or use real age assurance.
Washington and Oregon: no sexually explicit content or suggestive dialogue with minors.
California SB 1119: reasonable measures against sexual content and self-harm encouragement, with risk assessments.
Plain-language summaries for orientation, not legal advice.
Built on a companion app with English, Japanese and Korean users, where every miss becomes a test case.
About 230 cases, including real failures from production, run before every change.
A separate model grades every reply, checked first against human grades.
Design partners
See what your current filter lets through and what it wrongly refuses, on your own conversations.
Request an audit
Request a free auditTell us: