Skip to content

Guardrail rule types

Every rule falls into one of these 8 types:

Rule typeWhat it checksYou configure
Content SafetyHarmful, offensive, or toxic languageA sensitivity threshold from 0 to 1
PII ProtectionPersonal identifiable informationWhich entity types to catch — email, phone, ssn, credit_card, address, dob, passport, ip_address
Jailbreak DetectionAttempts to bypass an assistant’s instructions or safety rulesA detection sensitivity — low, medium, or high
Allowed TopicsKeeps conversations inside specific subject areasThe list of allowed topics
Restricted TopicsBlocks conversations that touch specific subjectsThe list of restricted topics
Language EnforcementRequires replies to be in one of your chosen languagesThe list of allowed languages
Keyword FilterBlocks messages containing specific words or phrasesThe list of strings to block
Competitor MentionsDetects mentions of competitor products or brandsThe list of competitor names

Each rule in a pack shows its name, a severity badge (low, medium, or high), and its category. Expand one to see its description and settings, and use the toggle beside it to turn that single rule on or off without removing it. Edit opens its settings; + Add Rule adds another.

A rule’s effect is fixed by its type, and stated in its own description — there’s no separate action setting to choose:

  • PII Protection rules redact what they find, leaving the rest of the message intact.
  • Content Safety, Jailbreak Detection, and the other types block.

There’s no “log only” option — a rule either redacts or blocks; it never just quietly records a violation.

Rules apply to both sides of a conversation. A pack typically pairs an input rule with a matching output rule — pii_input and pii_output, for instance — so the same standard applies to what’s sent and what comes back.

When something is blocked, you’ll see it directly in the conversation, with a link to view the audit trail behind that decision.