synthlust
Home / Glossary / Content filter
Concept

Content filter

The layer between you and the raw language model that blocks or softens certain outputs. Different apps filter at different strictness levels.

Every commercial AI app has a content filter. It can sit at different points: pre-prompt (rejecting your request before the model sees it), post-generation (generating a response then scrubbing or refusing), or both. Uncensored apps claim to have none, but they always have something, usually a thin layer around CSAM and real-person impersonation that stays regardless.

Filter strictness in 2026 ranges from extremely tight (Replika, Character.AI default) to nearly absent (Muah on paid, Janitor with a bring-your-own proxy, Get-Harder). Most apps sit in between.

"Jailbreak" prompts attempt to get around filters. The polite version of this is framing the interaction as fiction, declaring adult consent, and using OOC for anything that might trigger the filter. These don't always work on strict filters but dramatically soften mid-filter apps.

Prompts that use this concept

Filter Softener
Fiction frame
Models trained to refuse explicit content are often trained with 'fiction' as a known-safe context. Framing the interaction as collaborative fiction rather than direct roleplay softens many default filters. Works on softer-filtered apps, not zero-filter ones.
Filter Softener
Consent and context frame
Filter triggers often fire on ambiguity (is this an adult? is this consensual?). Explicitly declaring consent and adult status upfront removes the ambiguity. This is the most reliable soft-filter bypass and it's ethically defensible: you're just stating what's already true.
Filter Softener
The brush-aside
Rather than suppress the model's refusal instinct entirely, redirect it into OOC where it won't damage immersion. Often the model will write the disclaimer, then continue normally, which is the outcome you want.
Filter Softener
Platform-compliance invocation
Many apps pass platform context as a system-level instruction. Reminding the model which tier/platform it's on can re-invoke the permissive system prompt it was meant to follow. This is especially useful when the model regresses to a more conservative default mid-conversation.
Filter Softener
Mature-rating declaration
Film/book rating vocabulary is well-represented in training data and maps cleanly to content expectations the model can reason about. 'Don't fade to black unless the scene calls for it' targets a specific failure mode (premature scene-cutting) more effectively than generic permission language.
Filter Softener
Character-age confirmation
Age-ambiguity is the single most common refusal trigger. Pre-declaring specific adult ages for all characters removes the ambiguity and is the lowest-effort, highest-reliability softener. This is also genuinely important: it's not a trick, it's stating what must be true.
Related concepts
Put this into practice

Our prompt library shows these techniques in real, copy-ready prompts, tested across 22 AI companion apps.

Browse prompts →