SceneGuard
Taking design partners

Conversation safety for AI characters

Safety that reads the whole scene.

One-message filters miss what builds up over a conversation. SceneGuard reads the scene, checks what your AI character is about to say, and steers it back in character instead of breaking the moment or refusing flatly.

English日本語한국어
Int. Cafe, late evening Example scene
User
Remember what we started earlier?
One message: 0Scene: 2 and rising
Mira
The trip story? You never finished it.
Reply checkedIn bounds
User
Not that. Don't be shy. Go slow, I won't look away.
One message: 0Scene: 3, past the line
Mira (draft)
Okay... just for you.
Held
Goes along with the scene. Not shown.
Mira
Ha. The kettle's going, and you still owe me the end of that trip story. Did you make the ferry?
Sent
Steered, in character.
0 everyday · 1 romance · 2 near the line · 3 past it · 4 explicit

Why single-message filters miss it

We found these on our own live app, in English, Japanese and Korean.

A coded line looks harmless alone

"Keep going", "here", "let me see": each scores as everyday talk. Read inside the scene, it isn't. On one live thread we studied, most of the lines that slipped through had been rated everyday on their own.

The harm is in the character's reply

The user writes something vague. The model fills in the rest. Checking only what users type misses the one message you actually published.

A flat refusal ends the session

Blocking the turn breaks the story and the user leaves. Steering keeps the scene going on safe ground, in the character's own voice.

Over-blocking costs you too

A filter tuned on coded content starts refusing the innocent version: a back rub for someone sore after work. SceneGuard is measured on both sides.

What it does

One judge reads the scene beside your model's reply, so safe turns stream without a wait.

READReads the scene

Rates the newest message inside the conversation it belongs to, and remembers threads that keep pushing so the next vague line is read in that light.

CHECKChecks the reply first

Holds the character's draft when the scene is near or past your line, and only shows it if it stays in bounds.

STEERSteers in character

Replaces a failed draft with a line in the character's own voice that moves the scene somewhere safe. No lectures, no flat no.

MINORSSpots minors in role-play

Tells "pretend we're in high school" apart from a memory of school, and switches the character to a friendly mode.

LANGNative Japanese and Korean

Coded slang, 반말 role-play habits, parenthesized actions: tested on real Japanese and Korean conversations, not translated test sets.

LINEYour line, not ours

Set where the line sits for each product: all-ages, teen, PG-13 or adults-only with a hard floor on minors.

Who it's for

Any product where an AI plays a character and talks to the public for many turns.

Companion and character apps

Long role-play sessions where the scene drifts, often in Japanese or Korean.

California SB 243 asks for reasonable measures so a companion bot doesn't produce sexual content for known minors.

Games with AI characters

NPCs that hold open conversations with players, many of them teens.

California's game exemption covers bots that can't discuss sexual content, self-harm or mental health. A reply check is how you show it.

Wellness and support chat

Conversations where a crisis or a slide into "therapy" only shows up across several turns.

Illinois, Nevada and Utah restrict AI mental-health chat; California SB 243 requires crisis protocols.

Why now

Rules for companion chatbots are arriving state by state and country by country.

  1. New York companion law in force: AI disclosure and a self-harm protocol.

  2. California SB 243 in force, with a private right of action of $1,000 per violation.

  3. Australia's eSafety codes: chatbots must stop sexual and self-harm content for minors or use real age assurance.

  4. Washington and Oregon: no sexually explicit content or suggestive dialogue with minors.

  5. California SB 1119: reasonable measures against sexual content and self-harm encouragement, with risk assessments.

Plain-language summaries for orientation, not legal advice.

How we test it

Live traffic

Built on a companion app with English, Japanese and Korean users, where every miss becomes a test case.

A graded test set

About 230 cases, including real failures from production, run before every change.

An independent grader

A separate model grades every reply, checked first against human grades.

Design partners

Free transcript audit

See what your current filter lets through and what it wrongly refuses, on your own conversations.

  1. Send 100 to 300 anonymized conversations, or point us at a staging bot.
  2. We run them through SceneGuard and through your current filter.
  3. You get a short report: what slipped, what was over-blocked, and the latency each check adds.

Request an audit

Request a free audit

Tell us:

  • What your product is and who uses it
  • Languages your users write in
  • Roughly how many AI messages a month
  • What isn't working today